Claude
Skills
Sign in
Back

exploring-llm-evaluations

Included with Lifetime
$97 forever

Investigate AI observability evaluations of both types — `hog` (deterministic code-based) and `llm_judge` (LLM-prompt-based). Find existing evaluations, inspect their configuration, run them against specific generations, query individual pass/fail results, and generate AI-powered summaries of patterns across many runs. Use when the user asks to debug why an evaluation is failing, surface common failure modes, compare results across filters, dry-run a Hog evaluator, prototype a new LLM-judge prompt, or manage the evaluation lifecycle (create, update, enable/disable, delete).

Design

What this skill does


# Exploring LLM evaluations

PostHog evaluations score `$ai_generation` events. Each evaluation is one of two types,
both first-class:

- **`hog`** — deterministic Hog code that returns `true`/`false` (and optionally N/A).
  Best for objective rule-based checks: format validation (JSON parses, schema matches),
  length limits, keyword presence/absence, regex patterns, structural assertions, latency
  thresholds, cost guards. Cheap, fast, reproducible — no LLM call per run. Prefer this
  when the criterion can be expressed as code.
- **`llm_judge`** — an LLM scores generations against a prompt you write. Best for
  subjective or fuzzy checks: tone, helpfulness, hallucination detection, off-topic
  drift, instruction-following. Costs an LLM call per run and requires AI data
  processing approval at the org level.

Results from both types land in ClickHouse as `$ai_evaluation` events with the same
schema, so the read/query/summary workflows are identical regardless of evaluator type —
the only thing that changes is whether `$ai_evaluation_reasoning` was written by Hog
code or by an LLM.

This skill covers the full lifecycle: list/inspect/manage evaluation configs (Hog or
LLM judge), run them on specific generations, query individual results, and get an
AI-generated summary of pass/fail/N/A patterns across many runs.

## Tools

| Tool                                     | Purpose                                                        |
| ---------------------------------------- | -------------------------------------------------------------- |
| `posthog:llma-evaluation-list`           | List/search evaluation configs (filter by name, enabled flag)  |
| `posthog:llma-evaluation-get`            | Get a single evaluation config by UUID                         |
| `posthog:llma-evaluation-create`         | Create a new `llm_judge` or `hog` evaluation                   |
| `posthog:llma-evaluation-update`         | Update an existing evaluation (name, prompt, enabled, …)       |
| `posthog:llma-evaluation-delete`         | Soft-delete an evaluation                                      |
| `posthog:llma-evaluation-run`            | Run an evaluation against a specific `$ai_generation` event    |
| `posthog:llma-evaluation-test-hog`       | Dry-run Hog source against recent generations (no save)        |
| `posthog:llma-evaluation-summary-create` | AI-powered summary of pass/fail/N/A patterns across runs       |
| `posthog:execute-sql`                    | Ad-hoc HogQL over `$ai_evaluation` events                      |
| `posthog:query-llm-trace`                | Drill into the underlying generation that an evaluation scored |

All `llma-evaluation-*` tools are defined in `products/ai_observability/mcp/tools.yaml`.

## Event schema

Every run of an evaluation emits an `$ai_evaluation` event. Key properties:

| Property                    | Meaning                                                  |
| --------------------------- | -------------------------------------------------------- |
| `$ai_evaluation_id`         | UUID of the evaluation config                            |
| `$ai_evaluation_name`       | Human-readable name                                      |
| `$ai_target_event_id`       | UUID of the `$ai_generation` event being scored          |
| `$ai_trace_id`              | Parent trace ID (for jumping to the trace UI)            |
| `$ai_evaluation_result`     | `true` = pass, `false` = fail                            |
| `$ai_evaluation_reasoning`  | Free-text explanation (set by the LLM judge or Hog code) |
| `$ai_evaluation_applicable` | `false` when the evaluator decided the generation is N/A |

When `$ai_evaluation_applicable = false`, the run counts as N/A regardless of `$ai_evaluation_result`.
For evaluations that don't support N/A, this property may be `null` — treat null as "applicable".

## Workflow: investigate why an evaluation is failing

Works the same way for `llm_judge` and `hog` evaluations — the differences only matter
when you eventually go to fix the evaluator (edit the prompt vs. edit the Hog source).

### Step 1 — Find the evaluation

```json
posthog:llma-evaluation-list
{ "search": "hallucination", "enabled": true }
```

Look at the returned `id`, `name`, `evaluation_type`, and either:

- `evaluation_config.prompt` for an `llm_judge`
- `evaluation_config.source` for a `hog` evaluator

The Hog source is the ground truth for why a hog evaluator passes or fails — read it
before assuming the failure is in the generation.

### Step 2 — Get the AI-generated summary

```json
posthog:llma-evaluation-summary-create
{
  "evaluation_id": "<uuid>",
  "filter": "fail"
}
```

Returns:

- `overall_assessment` — natural-language summary
- `fail_patterns` — grouped patterns with `title`, `description`, `frequency`, and `example_generation_ids`
- `pass_patterns` and `na_patterns` — same shape, populated when `filter` includes them
- `recommendations` — actionable next steps
- `statistics` — `total_analyzed`, `pass_count`, `fail_count`, `na_count`

The endpoint analyses the most recent ~250 runs (`EVALUATION_SUMMARY_MAX_RUNS`).
Results are cached for one hour per `(evaluation_id, filter, set_of_generation_ids)`.
Pass `force_refresh: true` to recompute.

**Compare filters in two calls** to spot what's distinctive about failures vs passes:

```json
posthog:llma-evaluation-summary-create
{ "evaluation_id": "<uuid>", "filter": "pass" }
```

Then diff the `pass_patterns` against the `fail_patterns` from Step 2.

### Step 3 — Drill into example failing runs

Each pattern surfaces `example_generation_ids`. Pull the underlying trace for the most
representative example:

```json
posthog:query-llm-trace
{ "traceId": "<trace_id>", "dateRange": {"date_from": "-30d"} }
```

(If you only have a generation ID, query for it via `execute-sql` first to find the
parent trace ID — see below.)

### Step 4 — Verify the pattern with raw SQL

The summary is LLM-generated and should be verified. Use `execute-sql` to count and
spot-check:

```sql
posthog:execute-sql
SELECT
    properties.$ai_target_event_id AS generation_id,
    properties.$ai_trace_id AS trace_id,
    properties.$ai_evaluation_reasoning AS reasoning,
    timestamp
FROM events
WHERE event = '$ai_evaluation'
    AND properties.$ai_evaluation_id = '<evaluation_uuid>'
    AND properties.$ai_evaluation_result = false
    AND (
        properties.$ai_evaluation_applicable IS NULL
        OR properties.$ai_evaluation_applicable != false
    )
    AND timestamp >= now() - INTERVAL 7 DAY
ORDER BY timestamp DESC
LIMIT 25
```

The N/A guard (`IS NULL OR != false`) is important — it matches the same logic the
backend uses to bucket runs.

## Workflow: run an evaluation against a specific generation

Use this when the user pastes a trace/generation URL and asks "what would evaluation X
say about this?".

```json
posthog:llma-evaluation-run
{
  "evaluationId": "<eval_uuid>",
  "target_event_id": "<generation_event_uuid>",
  "timestamp": "2026-04-01T19:39:20Z",
  "event": "$ai_generation"
}
```

The `timestamp` is required for an efficient ClickHouse lookup of the target event.
Pass `distinct_id` if you have it — it speeds up the lookup further.

## Workflow: build and test a new evaluator

### Hog evaluator (deterministic, code-based)

Reach for this first when the criterion is rule-based — it's cheaper, faster, and
reproducible. Prototype with `llma-evaluation-test-hog` (no save):

```json
posthog:llma-evaluation-test-hog
{
  "source": "return event.properties.$ai_output_choices[1].content contains 'sorry';",
  "sample_count": 5,
  "allows_na": false
}
```

The handler returns the boolean result for each of the most recent N `$ai_generation`
events. Iterate on the source until it behaves as expected, then promote it via
`llma-evaluation-create`:

```json
posthog:llma-evaluation-create
{
  "name": "Output is valid JSON",
  "description": "Fails when the assistant message can't be parsed as JSON",
  "evaluation_
Files: 1
Size: 18.2 KB
Complexity: 27/100
Category: Design

Related in Design