scenario-testing
This skill should be used when writing tests, validating features, or needing to verify code works. Triggers on "write tests", "add test coverage", "validate feature", "integration test", "end-to-end", "e2e test", "mock", "unit test". Enforces scenario-driven testing with real dependencies in .scratch/ directory.
What this skill does
# Scenario-Driven Testing for AI Code Generation ## Core Principle **The Iron Law**: "NO FEATURE IS VALIDATED UNTIL A SCENARIO PASSES WITH REAL DEPENDENCIES" Mocks create false confidence. Only scenarios exercising real systems validate that code works. ## The Truth Hierarchy 1. **Scenario tests** (real system, real data) = **truth** 2. **Unit tests** (isolated) = human comfort only 3. **Mocks** = lies hiding bugs As stated in the principle: "A test that uses mocks is not testing your system. It's testing your assumptions about how dependencies behave." ## When to Use This Skill - Validating new functionality - Before declaring work complete - When tempted to use mocks - After fixing bugs requiring verification - Any time you need to prove code works ## Required Practices ### 1. Write Scenarios in `.scratch/` - Use any language appropriate to the task - Exercise the real system end-to-end - Zero mocks allowed - Must be in `.gitignore` (never commit) ### 2. Promote Patterns to `scenarios.jsonl` - Extract recurring scenarios as documented specifications - One JSON line per scenario - Include: name, description, given/when/then, validates - This file IS committed ### 3. Use Real Dependencies External APIs must hit actual services (sandbox/test mode acceptable). Mocking any dependency invalidates the scenario. ### 4. Independence Requirement Each scenario must run standalone without depending on prior executions. This enables: - Parallel execution - Prevents hidden ordering dependencies - Reliable CI/CD integration ## What Makes a Scenario Invalid A scenario is invalid if it: - Contains any mocks whatsoever - Uses fake data instead of real storage - Depends on another scenario running first - Never actually executed to verify it passes ## Common Violations to Avoid Reject these rationalizations: - **"Just a quick unit test..."** - Unit tests don't validate features - **"Too simple for end-to-end..."** - Integration breaks simple things - **"I'll mock for speed..."** - Speed doesn't matter if tests lie - **"I don't have API credentials..."** - Ask your human partner for real ones ## Definition of Done A feature is complete only when: 1. ✅ A scenario in `.scratch/` passes with zero mocks 2. ✅ Real dependencies are exercised 3. ✅ `.scratch/` remains in `.gitignore` 4. ✅ Robust patterns extracted to `scenarios.jsonl` ## Example Workflow 1. **Write scenario** - Create `.scratch/test-user-registration.py` 2. **Use real dependencies** - Hit real database, real auth service (test mode) 3. **Run and verify** - Execute scenario, confirm it passes 4. **Extract pattern** - Document in `scenarios.jsonl` 5. **Keep .scratch ignored** - Never commit scratch scenarios ## Why This Matters - **Unit tests** verify isolated logic - **Integration tests** verify components work together - **Scenario tests** verify the system actually works Only scenario tests prove your feature delivers value to users.
Related in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.