test-writing
Guide test writing and review with focus on observable behavior over implementation. Use when the user asks to write, review, or refactor tests; debates unit vs integration scope; questions whether to mock a dependency; or asks why tests break on refactor.
What this skill does
# test-writing
Test units of behavior, not units of code. A good test describes one meaningful scenario, observes the result from the system boundary, and survives refactoring.
## Decision tree
```
Complex/important code?
├─ No → Skip (trivial getters, setters, glue)
└─ Yes → Many dependencies?
├─ No → Unit test the behavior through the public API
└─ Yes → Can extract logic?
├─ Yes → Extract behavior to domain + test thin orchestration separately
└─ No → Integration test the end-to-end behavior
```
## What to test
- **Unit**: one business scenario or domain behavior through a public API, not one class by default.
- **Integration**: orchestration, controllers, real databases, real filesystems, and observable system results.
- **Skip**: trivial accessors, private methods (extract instead), implementation details (internal structure, call order).
## Mocking rules
- **Mock unmanaged dependencies** (external systems others depend on): SMTP, message buses, third-party APIs. Mock at the system boundary.
- **Never mock managed dependencies** (resources only your app uses): database, filesystem, in-process domain collaborators. Use the real thing in integration tests.
- **Never mock time** — inject it as a dependency instead of calling `DateTime.Now` / `time.now()`.
- When using mocks, verify only edge interactions that are externally visible and compatibility-sensitive.
## Test structure (AAA)
```
// Arrange: set up inputs
// Act: execute ONE operation
// Assert: verify the observable outcome
```
- One behavior per test; multiple assertions are acceptable when they describe the same outcome.
- Keep Act to one line when possible. Multiple Act lines often signal poor API encapsulation or multiple behaviors.
- No conditionals or loops in tests.
- Name tests in plain English, with underscores for readability: `User_login_fails_with_invalid_password`.
## Style preference (best → worst)
1. **Output-based**: `result = add(2, 3); assert result == 5`
2. **State-based**: `cart.add(item); assert cart.count == 1`
3. **Communication-based** (use sparingly): `verify(service.send, once)`
## Red flags
- Test duplicates production logic → use hardcoded expectations.
- Test breaks on refactor → coupled to implementation, not behavior.
- Coverage % as the goal → use coverage only as a negative signal for untested areas.
- Mocking the database → use a real instance.
- Test maps one-to-one to a class or method → reframe around the scenario a domain expert would recognize.
- Testing private methods → extract to a class with its own public API.
## Quality bar
A good test has: **protection** (catches real bugs), **refactoring resistance** (survives implementation changes), **speed** (ms for unit), **clarity** (one obvious failure reason).
A bad test is worse than no test — it locks in the wrong shape.
## Deeper material
- [REFERENCE.md](REFERENCE.md) — humble-object refactor, three-context database pattern, worked examples.
Related in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.