Claude
Skills
Sign in
Back

autonomous-bot-e2e-testing

Included with Lifetime
$97 forever

This skill should be used when the user asks to "test a bot feature", "E2E test Nook", "test without Telegram", "simulate provider outage", "test fallback chain", "verify model switching", "API test LettaBot", "test sendToAgent", or needs to independently validate Nook/LettaBot behavior end-to-end without the user sending messages through Telegram or WhatsApp.

Backend & APIs

What this skill does


# Autonomous Bot E2E Testing

Test Nook/LettaBot features end-to-end via the API without requiring the user to send
Telegram or WhatsApp messages. This workflow covers sending test messages, simulating
provider outages, verifying fallback behavior, and restoring the system to a clean state.

## When to Use

- Any time a Nook feature needs E2E validation without the user present
- Testing fallback chains, model switching, tool execution, or response quality
- Simulating provider outages to verify graceful degradation
- Verifying that both code paths (sendToAgent and processMessage) handle a feature

## Critical Concept: Dual Code Paths

Nook has two independent message-handling paths. Features must exist in both:

| Path | Entry Point | Context |
|------|------------|---------|
| **sendToAgent** | API calls (`/api/v1/chat`) | Direct agent interaction; `msg.stopReason` lives on the result message directly |
| **processMessage** | Telegram/WhatsApp webhook | Channel-driven; error details in `lastErrorDetail`; retry logic skips when `retryConvId` is undefined (default conversation) |

API-based E2E testing exercises `sendToAgent` only. To cover `processMessage`, use the
actual messaging channel or inspect that path's code separately.

## Workflow

### Step 1 — Send a Test Message via API

Retrieve the API key from `lettabot-api.json` (field name: `apiKey`, camelCase) on ROG.

```bash
curl -X POST http://127.0.0.1:8081/api/v1/chat \
  -H "Content-Type: application/json" \
  -H "X-API-Key: <apiKey from lettabot-api.json>" \
  -d '{"message": "Hello, this is a test"}'
```

This routes through `sendToAgent`, not `processMessage`.

### Step 2 — Simulate a Provider Outage

To trigger fallback behavior, break the LLM provider API key in the environment:

1. **Save the real key** from `~/nook/docker/.env` before modifying it.
2. **Replace the API key** with an invalid value (e.g., `sk-BROKEN`).
   - Modify the API key, not the model handle. Letta validates model handles at PATCH
     time, so an invalid handle causes an immediate rejection rather than a runtime
     fallback.
3. **Recreate the container**: `docker compose up -d` (never use `restart` -- it does
   not pick up new env vars).
4. Send a test message (Step 1) and observe the fallback behavior.

### Step 3 — Verify via Logs

Inspect LettaBot service logs for fallback-related activity:

```bash
journalctl -u lettabot.service --since '5 min ago' --no-pager
```

Grep for key indicators:
- `fallback` -- fallback chain activation
- `switching` -- model switching events
- `tier` -- tier transitions
- `stopReason` -- why the response stopped
- `llm_api_error` -- provider-level failures

### Step 4 — Verify Model State

Confirm which model the agent is currently using:

```bash
curl -s http://127.0.0.1:8283/v1/agents/<agent-id>/ \
  -H "Authorization: Bearer <letta-token>" | jq '.llm_config.model'
```

Compare against the expected primary or fallback model.

### Step 5 — Restore

1. **Replace the broken key** in `~/nook/docker/.env` with the saved real key.
2. **Recreate the container**: `docker compose up -d`.
3. **PATCH the agent back** to the primary model if the fallback chain switched it:
   ```bash
   curl -X PATCH http://127.0.0.1:8283/v1/agents/<agent-id>/ \
     -H "Authorization: Bearer <letta-token>" \
     -H "Content-Type: application/json" \
     -d '{"llm_config": {"model": "<primary-model-handle>"}}'
   ```
4. Send a final test message to confirm normal operation.

## Key Gotchas

- **`docker compose restart` vs `up -d`**: `restart` does not pick up `.env` changes.
  Always use `up -d` after editing `.env`.
- **`sendToAgent` stopReason location**: `msg.stopReason` is on the result message
  directly, not inside `lastErrorDetail`.
- **`processMessage` retry skip**: Retry logic in `processMessage` skips when
  `retryConvId` is undefined (default conversation). This does not affect API tests but
  is important when debugging channel-based failures.
- **Model handle vs API key**: Break the API key to simulate an outage. Breaking the
  model handle causes Letta to reject the PATCH immediately, which does not simulate a
  runtime failure.

## Additional Resources

### Reference Files

For detailed dual-path analysis and advanced testing scenarios, consult:
- **`references/dual-path-analysis.md`** -- Detailed comparison of sendToAgent vs processMessage code paths, field locations, and retry semantics

Related in Backend & APIs