Claude
Skills
Sign in
Back

output-dev-eval-testing

Included with Lifetime
$97 forever

Create offline evaluation tests for Output SDK workflows using @outputai/evals. Use when implementing test evaluators with verify(), creating dataset YAML files, building eval workflows, or running workflow tests via CLI.

Backend & APIs

What this skill does


# Offline Evaluation Testing

## Overview

The `@outputai/evals` package provides an offline evaluation framework for testing workflow quality using datasets and evaluators. This is **complementary** to the runtime `evaluator()` from `@outputai/core`:

| Aspect | Runtime Evaluators (`@outputai/core`) | Offline Eval Tests (`@outputai/evals`) |
|--------|----------------------------------------|------------------------------------------|
| **When** | During workflow execution | After execution, at test time |
| **Where** | `evaluators.ts` in workflow folder | `tests/evals/` in workflow folder |
| **Purpose** | Live quality scoring with confidence | Dataset-driven pass/fail verification |
| **Triggered by** | Workflow orchestration | `output workflow test` CLI command |
| **Returns** | `EvaluationBooleanResult`, etc. | `Verdict` helpers (pass/partial/fail) |

Use offline eval testing when you want to validate workflow behavior against known datasets, build regression test suites, or assess subjective quality with LLM judges.

## When to Use This Skill

- Creating files in `tests/evals/` or `tests/datasets/`
- Writing evaluators that use `verify()` from `@outputai/evals`
- Creating YAML dataset files for test cases
- Building eval workflows with `evalWorkflow()`
- Running `output workflow test` commands
- Setting up ground truth data for evaluators

## Directory Structure

Add a `tests/` directory inside the workflow folder:

```
src/workflows/{workflow_name}/
├── workflow.ts
├── steps.ts
├── evaluators.ts          # Runtime evaluators (optional)
├── types.ts
└── tests/
    ├── datasets/
    │   ├── happy_path.yml
    │   └── edge_case.yml
    └── evals/
        ├── evaluators.ts  # Offline eval test evaluators
        ├── workflow.ts     # Eval workflow definition
        └── [email protected]  # LLM judge prompts (optional)
```

## Creating Evaluators with `verify()`

Import `verify` and `Verdict` from `@outputai/evals` (not `@outputai/core`):

```typescript
// tests/evals/evaluators.ts
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';
```

### `verify()` Signature

```typescript
verify(options, checkFn)
```

**Options:**
- `name` — unique evaluator identifier (snake_case)
- `input` — Zod schema for the workflow input (optional, defaults to `z.any()`)
- `output` — Zod schema for the workflow output (optional, defaults to `z.any()`)

**Check function receives:**
```typescript
{
  input,    // typed workflow input
  output,   // typed workflow output
  context: {
    ground_truth: Record<string, unknown>  // from dataset YAML
  }
}
```

**Returns:** any `Verdict` helper result.

### Basic Example

```typescript
import { verify, Verdict } from '@outputai/evals';
import { z } from '@outputai/core';

export const evaluateSum = verify(
  {
    name: 'evaluate_sum',
    input: z.object({ values: z.array(z.number()) }),
    output: z.object({ result: z.number() })
  },
  ({ input, output }) =>
    Verdict.equals(output.result, input.values.reduce((a, b) => a + b, 0))
);
```

### Using Ground Truth

Ground truth values come from the dataset YAML and are available via `context.ground_truth`:

```typescript
export const lengthCheck = verify(
  { name: 'length_check', input: blogInput, output: blogOutput },
  ({ output, context }) =>
    Verdict.gte(output.blog_post.length, Number(context.ground_truth.min_length ?? 100))
);
```

## Verdict Helpers

All deterministic helpers return results with confidence `1.0`.

### Equality & Comparison

| Method | Description |
|--------|-------------|
| `Verdict.equals(actual, expected)` | Strict equality (`===`) |
| `Verdict.closeTo(actual, expected, tolerance)` | Within numeric tolerance |
| `Verdict.gt(actual, threshold)` | Greater than |
| `Verdict.gte(actual, threshold)` | Greater than or equal |
| `Verdict.lt(actual, threshold)` | Less than |
| `Verdict.lte(actual, threshold)` | Less than or equal |
| `Verdict.inRange(actual, min, max)` | Within inclusive range |

### String & Array

| Method | Description |
|--------|-------------|
| `Verdict.contains(haystack, needle)` | String includes substring |
| `Verdict.matches(value, pattern)` | Regex match |
| `Verdict.includesAll(actual, expected)` | Array contains all expected values |
| `Verdict.includesAny(actual, expected)` | Array contains at least one expected value |

### Boolean

| Method | Description |
|--------|-------------|
| `Verdict.isTrue(value)` | Value is `true` |
| `Verdict.isFalse(value)` | Value is `false` |

### Manual Verdicts

| Method | Description |
|--------|-------------|
| `Verdict.pass(reasoning?)` | Explicit pass |
| `Verdict.partial(confidence, reasoning?, feedback?)` | Partial pass with confidence |
| `Verdict.fail(reasoning, feedback?)` | Explicit fail |

## LLM Judge Evaluators

Before writing a judge prompt, identify the specific failure mode via error analysis (`output-eval-error-analysis`). Design the judge following `output-eval-judge-prompt`. After writing it, validate against human labels using `output-eval-validate-judge`.

For subjective quality assessments, use judge functions with `.prompt` files:

```typescript
import { verify, judgeVerdict, judgeScore, judgeLabel } from '@outputai/evals';

// Returns pass/partial/fail verdict from an LLM
export const evaluateTopic = verify(
  { name: 'evaluate_topic', input: blogInput, output: blogOutput },
  async ({ input, output, context }) =>
    judgeVerdict({
      prompt: 'judge_topic@v1',
      variables: {
        blog_title: output.title,
        blog_post: output.blog_post,
        required_topic: String(context.ground_truth.required_topic ?? input.topic)
      }
    })
);

// Returns a numeric score from an LLM
export const evaluateQuality = verify(
  { name: 'evaluate_quality', input: blogInput, output: blogOutput },
  async ({ input, output }) =>
    judgeScore({
      prompt: 'judge_quality@v1',
      variables: { blog_title: output.title, blog_post: output.blog_post, topic: input.topic }
    })
);

// Returns a string label from an LLM
export const evaluateTone = verify(
  { name: 'evaluate_tone', input: blogInput, output: blogOutput },
  async ({ output }) =>
    judgeLabel({
      prompt: 'judge_tone@v1',
      variables: { blog_title: output.title, blog_post: output.blog_post }
    })
);
```

### Judge `.prompt` File Format

Judge prompt files live alongside evaluators in `tests/evals/`:

```yaml
# tests/evals/[email protected]
---
provider: anthropic
# current as of 2026-05-04 — run output-dev-model-selection for the latest
model: claude-haiku-4-5-20251001
temperature: 0
maxTokens: 1000
---

<system>
You are an evaluation judge. Assess whether a blog post is faithfully about the required topic.

Return a JSON object with:
- verdict: "pass" if the blog clearly focuses on the topic, "partial" if it mentions the topic but lacks depth, "fail" if it is not about the topic
- reasoning: a brief explanation of your judgment
</system>

<user>
Required topic: {{ required_topic }}

Blog title: {{ blog_title }}

Blog post:
{{ blog_post }}

Judge whether this blog post is faithfully about the required topic.
</user>
```

## Creating Eval Workflows

The eval workflow wires evaluators together and defines how to interpret results.

```typescript
// tests/evals/workflow.ts
import { evalWorkflow } from '@outputai/evals';
import { evaluateSum } from './evaluators.js';

export default evalWorkflow({
  name: 'simple_eval',
  evals: [
    {
      evaluator: evaluateSum,
      criticality: 'required',
      interpret: { type: 'boolean' }
    }
  ]
});
```

### Eval Definition Fields

Each entry in the `evals` array has:

- **`evaluator`** — the function created by `verify()`
- **`criticality`** — `'required'` (affects pass/fail) or `'informational'` (reported but doesn't block)
- **`interpret`** — how to convert the evaluator's return value into a verdict

### Interpret Types

| Type | Evaluator Returns | Mapping |
|------|-------------------|---------|

Related in Backend & APIs