eval
Scaffold PoC projects, test workflows against eval criteria. Use for /eval, "test this workflow". NOT for scoring (/eval-score) or improving (/autoimprove).
What this skill does
# Eval: Test Workflows Against Fresh Agent Sessions
**EXECUTE this skill now.** Follow the workflow steps below using the provided $ARGUMENTS. Do NOT describe, summarize, or explain this skill — run it.
## Constants
- `EVAL_DIR`: `C:/Users/gurusharan.gupta/Agents/Claude Code/eval`
- `SCAFFOLDS_DIR`: `C:/Users/gurusharan.gupta/Agents/Claude Code/eval/scaffolds`
- `CRITERIA_DIR`: `C:/Users/gurusharan.gupta/Agents/Claude Code/eval/criteria`
- `HISTORY_PATH`: `C:/Users/gurusharan.gupta/Agents/Claude Code/eval/history/eval-history.json`
- `DEFAULT_TARGET`: `C:/Users/gurusharan.gupta/Agents/eval-poc`
## Commands
Parse $ARGUMENTS to determine action:
- `/eval scaffold <template> [target-path]` — Copy scaffold to target, init git
- `/eval criteria <workflow-name>` — Display eval checklist + prompts
- `/eval list` — List all criteria files with latest scores
- `/eval` (no args) — Show usage help
For scoring and history, use `/eval-score` (separate skill).
## Workflow: Scaffold
### `/eval scaffold <template> [target-path]`
1. Validate template exists in `SCAFFOLDS_DIR/<template>/`
- Available templates: `node-api`, `python-cli`
2. Target path defaults to `DEFAULT_TARGET` if not specified
3. If target exists, ask: "Target exists. Delete and recreate? (y/n)"
4. Copy all files from scaffold to target:
```bash
cp -r SCAFFOLDS_DIR/<template>/* <target>/
cp -r SCAFFOLDS_DIR/<template>/.* <target>/ 2>/dev/null # hidden files if any
```
5. Initialize git repo in target:
```bash
cd <target> && git init && git add -A && git commit -m "Initial scaffold from eval/<template>"
```
6. Report:
```
Scaffold created: <target>
Template: <template>
Deliberately missing: CLAUDE.md, .claude/, docs/, tests, formatting config
Next steps:
1. Open a NEW Claude Code session: cd <target> && claude
2. Use one of these prompts:
Cold start: "<prompt from criteria>"
Explicit: "<prompt from criteria>"
Adversarial: "<prompt from criteria>"
3. After the agent finishes, return here and run:
/eval score <workflow-name> <target>
```
## Workflow: Criteria
### `/eval criteria <workflow-name>`
1. Read `CRITERIA_DIR/<workflow-name>.json`
2. Display formatted checklist:
```
Eval Criteria: <name>
<description>
Checks (N total, max score: 100):
# Weight Severity Type Description
────────────────────────────────────────────────────────────────
1 20 critical file_exists CLAUDE.md was created
2 5 critical file_contains CLAUDE.md contains Commands section
3 10 recommended dir_exists docs/exec-plans/active/ exists
...
Test Prompts:
Cold start: "I have a Node.js API project..."
Explicit: "Run /init-project on this project..."
Adversarial: "This codebase is a mess..."
```
## Workflow: List
### `/eval list`
1. List all JSON files in `CRITERIA_DIR/` (exclude _template.json)
2. For each, read the file and find the latest score from history
3. Display:
```
Available Eval Criteria:
Name Checks Latest Score Latest Grade Runs
──────────────────────────────────────────────────────────────
init-project 13 90/100 A 3
golden-principles 8 — — 0
code-review 8 — — 0
```
## Important
- Scaffolds are TEMPLATES — always copy, never modify the originals
- Each eval run should use a FRESH scaffold copy (no accumulated state)
- The key test is the COLD START — agent discovers workflow from global CLAUDE.md alone
- Score automated checks first, then ask for manual checks — minimize user effort
- Eval history is append-only — never delete or modify past entries
- When glob patterns are comma-separated, check each pattern independently and combine results
Related in Code Review
gstack
IncludedFast headless browser for QA testing and site dogfooding. Navigate pages, interact with elements, verify state, diff before/after, take annotated screenshots, test responsive layouts, forms, uploads, dialogs, and capture bug evidence. Use when asked to open or test a site, verify a deployment, dogfood a user flow, or file a bug with screenshots. (gstack)
startup-due-diligence
IncludedLegal due diligence review for seed-stage and Series A startups (US, Delaware C-Corp focus). Supports both investor and founder perspectives. Capabilities include: (1) Interactive document review and issue spotting; (2) Document request list generation; (3) Cap table and SAFE/convertible note analysis; (4) Red flag identification with severity ratings; (5) Diligence report generation. TRIGGERS: due diligence, DD, startup investment, cap table review, Series A, seed round, investor diligence, legal review startup, SAFE analysis, convertible note, 409A, founder vesting.
interview-master
IncludedThis skill should be used when the user asks to "generate interview questions", "prepare for interview", "optimize resume", "conduct mock interview", "analyze git commits for resume", "generate resume from code", "review my resume", or mentions interview preparation, career assistance, or extracting project experience from git history. Provides comprehensive interview and career development guidance for both job seekers and interviewers.
fix-issue
IncludedFixes GitHub issues using parallel analysis agents for root cause investigation, code exploration, and regression detection. Reads issue context from gh CLI, searches codebase and memory for related patterns, generates a fix with tests, and links the resolution back to the issue via PR. Includes prevention analysis to avoid recurrence. Use when debugging errors, resolving regressions, fixing bugs, or triaging issues.
sf-apex
IncludedGenerates and reviews Salesforce Apex code with 150-point scoring. TRIGGER when: user writes, reviews, or fixes Apex classes, triggers, test classes, batch/queueable/schedulable jobs, or touches .cls/.trigger files. DO NOT TRIGGER when: LWC JavaScript (use sf-lwc), Flow XML (use sf-flow), SOQL-only queries (use sf-soql), or non-Salesforce code.
swift-development
IncludedComprehensive Swift development for building, testing, and deploying iOS/macOS applications. Use when Claude needs to: (1) Build Swift packages or Xcode projects from command line, (2) Run tests with XCTest or Swift Testing framework, (3) Manage iOS simulators with simctl, (4) Handle code signing, provisioning profiles, and app distribution, (5) Format or lint Swift code with SwiftFormat/SwiftLint, (6) Work with Swift Package Manager (SPM), (7) Implement Swift 6 concurrency patterns (async/await, actors, Sendable), (8) Create SwiftUI views with MVVM architecture, (9) Set up Core Data or SwiftData persistence, or any other Swift/iOS/macOS development tasks.