evaluate-improve
Suggest improvements to SKILL.md content, descriptions, or tool config from eval results. Use when raising pass rates, fixing triggering, or iterating on a skill after evaluation.
What this skill does
# /evaluate:improve Analyze evaluation results and suggest concrete improvements to a skill. ## When to Use This Skill | Use this skill when... | Use alternative when... | |------------------------|------------------------| | Have eval results and want to improve the skill | Need to run evals first -> `/evaluate:skill` | | Want to improve skill description for better triggering | Want to view raw results -> `/evaluate:report` | | Iterating on a skill to increase pass rate | Want to file a bug -> `/feedback:session` | | Optimizing skill instructions after benchmarking | Need structural fixes -> `plugin-compliance-check.sh` | ## Parameters Parse these from `$ARGUMENTS`: | Parameter | Default | Description | |-----------|---------|-------------| | `<plugin/skill-name>` | required | Path as `plugin-name/skill-name` | | `--apply` | false | Apply approved changes to SKILL.md | | `--description-only` | false | Focus on description improvements only | ## Execution ### Step 1: Load eval results Read the most recent benchmark from: ``` <plugin-name>/skills/<skill-name>/eval-results/benchmark.json ``` If no results exist, suggest running `/evaluate:skill` first and stop. Also read the current SKILL.md to understand the skill. ### Step 2: Analyze results Delegate analysis to the `eval-analyzer` agent via Task: ``` Task subagent_type: eval-analyzer Prompt: Analyze these evaluation results and identify improvement opportunities. Skill: <path to SKILL.md> Benchmark: <benchmark.json contents> Mode: comparison (if baseline data exists) or benchmark (otherwise) ``` The analyzer produces categorized suggestions: - **instructions**: Execution flow improvements - **description**: Better intent-matching text - **examples**: Missing or insufficient examples - **error_handling**: Missing edge cases - **tools**: Better tool configurations - **structure**: Organizational improvements ### Step 3: Filter suggestions If `--description-only`, filter to only `description` category suggestions. Sort remaining suggestions by priority (high > medium > low). ### Step 4: Present suggestions Present the categorized suggestions to the user: ``` ## Improvement Suggestions: <plugin/skill-name> Current pass rate: 72% ### High Priority 1. **[instructions]** Add explicit error handling for missing git config Evidence: eval-003 fails because the skill doesn't check for git user.name 2. **[description]** Add "conventional commit" as trigger phrase Evidence: Skill not selected when user says "make a conventional commit" ### Medium Priority 3. **[examples]** Add breaking change example to execution steps Evidence: eval-004 inconsistently handles breaking changes ### Low Priority 4. **[structure]** Move flag reference to Quick Reference table Evidence: Flags scattered across multiple sections ``` If `--apply` is NOT set, stop here. ### Step 5: Apply changes (if --apply) Use AskUserQuestion to let the user select which suggestions to apply: ``` Which improvements should I apply? [x] Add error handling for missing git config [x] Add trigger phrases to description [ ] Add breaking change example [ ] Restructure flag reference ``` For each approved suggestion: 1. Read the current SKILL.md 2. Apply the change using Edit 3. Update the `modified` date in frontmatter After applying changes, update (or create) the history file at: ``` <plugin-name>/skills/<skill-name>/eval-results/history.json ``` Add a new iteration entry recording: - Version number (increment from previous) - Timestamp - Pass rate from current benchmark - Summary of changes made ### Step 6: Suggest re-evaluation After applying changes, suggest: ``` Changes applied. Run `/evaluate:skill <plugin/skill-name>` to measure improvement. ``` ## Agentic Optimizations | Context | Command | |---------|---------| | Read benchmark | `cat <plugin>/skills/<skill>/eval-results/benchmark.json \| jq .summary` | | Read skill | `cat <plugin>/skills/<skill>/SKILL.md` | | Read history | `cat <plugin>/skills/<skill>/eval-results/history.json \| jq '.iterations[-1]'` | | Check pass rate | `cat <plugin>/skills/<skill>/eval-results/benchmark.json \| jq '.summary.with_skill.mean_pass_rate'` | ## Quick Reference | Flag | Description | |------|-------------| | `--apply` | Apply approved changes to SKILL.md | | `--description-only` | Focus on description improvements only |
Related in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.