openai-agents
OpenAI Agents SDK for JavaScript/TypeScript (text + voice agents). Use for multi-agent workflows, tools, guardrails, or encountering Zod errors, MCP failures, infinite loops, tool call issues.
What this skill does
# OpenAI Agents SDK Skill
Complete skill for building AI applications with OpenAI Agents SDK (JavaScript/TypeScript), covering text agents, realtime voice agents, multi-agent workflows, and production deployment patterns.
---
## Quick Start
### Installation
```bash
bun add @openai/agents zod@3
bun add @openai/agents-realtime # For voice agents
```
Set environment variable:
```bash
export OPENAI_API_KEY="your-api-key"
```
### Basic Text Agent
```typescript
import { Agent, run, tool } from '@openai/agents';
import { z } from 'zod';
const agent = new Agent({
name: 'Assistant',
instructions: 'You are helpful.',
tools: [tool({
name: 'get_weather',
parameters: z.object({ city: z.string() }),
execute: async ({ city }) => `Weather in ${city}: sunny`,
})],
model: 'gpt-4o-mini',
});
const result = await run(agent, 'What is the weather in SF?');
```
### Voice Agent & Multi-Agent
```typescript
// Voice agent
const voiceAgent = new RealtimeAgent({
voice: 'alloy',
model: 'gpt-4o-realtime-preview',
});
// Browser session
const session = new RealtimeSession(voiceAgent, {
apiKey: sessionApiKey, // From backend!
transport: 'webrtc',
});
// Multi-agent handoffs
const triageAgent = Agent.create({
handoffs: [billingAgent, techAgent],
});
```
**17 Templates**: `templates/` directory has production-ready examples for all patterns.
---
## Top 3 Critical Errors
### 1. Zod Schema Type Errors
**Error**: Type errors with tool parameters even when structurally compatible.
**Workaround**: Define schemas inline.
```typescript
// ❌ Can cause type errors
parameters: mySchema
// ✅ Works reliably
parameters: z.object({ field: z.string() })
```
**Source**: [GitHub #188](https://github.com/openai/openai-agents-js/issues/188)
### 2. MCP Tracing Errors
**Error**: "No existing trace found" with MCP servers.
**Workaround**:
```typescript
import { initializeTracing } from '@openai/agents/tracing';
await initializeTracing();
```
**Source**: [GitHub #580](https://github.com/openai/openai-agents-js/issues/580)
### 3. MaxTurnsExceededError
**Error**: Agent loops infinitely.
**Solution**: Increase maxTurns or improve instructions:
```typescript
const result = await run(agent, input, {
maxTurns: 20,
});
// Or improve instructions
instructions: `After using tools, provide a final answer.
Do not loop endlessly.`
```
**All 9 Errors**: Load `references/common-errors.md` for complete error catalog with workarounds.
---
## When to Load References
Load reference files when working on specific aspects of agent development:
### Agent Patterns (`references/agent-patterns.md`)
Load when:
- Designing multi-agent orchestration strategies
- Choosing between LLM-based vs code-based orchestration
- Implementing parallel agent execution
- Creating agents-as-tools patterns
- Need to understand when to use which orchestration pattern
### Common Errors (`references/common-errors.md`)
Load when:
- Debugging agent issues beyond the top 3 errors above
- Implementing comprehensive error handling
- Encountering: GuardrailExecutionError, ToolCallError, Schema Mismatch, Ollama integration, webSearchTool failures, Agent Builder export bugs
- Building production error recovery patterns
### Realtime Transports (`references/realtime-transports.md`)
Load when:
- Choosing between WebRTC vs WebSocket for voice agents
- Optimizing voice agent latency
- Debugging voice connection issues
- Understanding network/firewall requirements for voice
- Implementing custom audio sources/sinks
### Cloudflare Integration (`references/cloudflare-integration.md`)
Load when:
- Deploying agents to Cloudflare Workers
- Understanding Workers limitations (CPU, memory, no voice)
- Implementing streaming in Workers
- Debugging Workers-specific issues
- Optimizing for Workers performance and costs
### Official Links (`references/official-links.md`)
Load when:
- Need official documentation links
- Looking for examples or community resources
- Checking latest SDK versions
- Finding pricing information
- Need migration guides
---
## Core Concepts Summary
**Agents**: LLMs equipped with instructions and tools.
**Tools**: Functions with Zod schemas that agents can call automatically.
**Handoffs**: Multi-agent delegation where agents route tasks to specialists.
**Guardrails**: Input/output validation for safety (content filtering, PII detection).
**Structured Outputs**: Type-safe responses using Zod schemas.
**Streaming**: Real-time event streaming for progressive responses.
**Human-in-the-Loop**: Require approval for specific tool executions (`requiresApproval: true`).
For detailed examples, see templates in `templates/text-agents/` and `templates/realtime-agents/`.
---
## Text Agents Quick Reference
```typescript
// Basic
const result = await run(agent, 'Your question');
// Streaming
const stream = await run(agent, input, { stream: true });
// Structured output
const agent = new Agent({
outputType: z.object({ sentiment: z.enum([...]), confidence: z.number() }),
});
```
**Templates**: `templates/text-agents/` (8 templates)
---
## Realtime Voice Agents Quick Reference
```typescript
const voiceAgent = new RealtimeAgent({
voice: 'alloy', // alloy, echo, fable, onyx, nova, shimmer
model: 'gpt-4o-realtime-preview',
});
const session = new RealtimeSession(voiceAgent, {
apiKey: sessionApiKey,
transport: 'webrtc', // or 'websocket'
});
```
**Voice handoff constraints**: Cannot change voice/model during handoff.
**Templates**: `templates/realtime-agents/` (3 templates) | **Details**: `references/realtime-transports.md`
---
## Framework Integration Quick Reference
### Cloudflare Workers (Experimental)
```typescript
export default {
async fetch(request: Request, env: Env) {
const { message } = await request.json();
process.env.OPENAI_API_KEY = env.OPENAI_API_KEY;
const agent = new Agent({
name: 'Assistant',
instructions: 'Be helpful and concise',
model: 'gpt-4o-mini',
});
const result = await run(agent, message, { maxTurns: 5 });
return new Response(JSON.stringify({
response: result.finalOutput,
tokens: result.usage.totalTokens,
}));
},
};
```
**Limitations**: No realtime voice, CPU time limits (30s max), memory constraints (128MB).
**Templates**: `templates/cloudflare-workers/` (2 templates)
**Details**: Load `references/cloudflare-integration.md` for complete Workers guide.
### Next.js App Router
```typescript
// app/api/agent/route.ts
import { NextRequest, NextResponse } from 'next/server';
import { Agent, run } from '@openai/agents';
export async function POST(request: NextRequest) {
const { message } = await request.json();
const agent = new Agent({ /* ... */ });
const result = await run(agent, message);
return NextResponse.json({ response: result.finalOutput });
}
```
**Templates**: `templates/nextjs/` (2 templates)
---
## Guardrails & Human-in-the-Loop
```typescript
// Input/output guardrails
const agent = new Agent({
inputGuardrails: [homeworkDetectorGuardrail],
outputGuardrails: [piiFilterGuardrail],
});
// Human approval
const tool = tool({
requiresApproval: true,
execute: async ({ amount }) => `Refunded $${amount}`,
});
// Handle approval loop
while (result.interruption?.type === 'tool_approval') {
result = (await promptUser(result.interruption))
? await result.state.approve(result.interruption)
: await result.state.reject(result.interruption);
}
```
**Templates**: `templates/text-agents/agent-guardrails-*.ts`, `agent-human-approval.ts`
---
## Orchestration Patterns Summary
**LLM-Based**: Agent decides routing autonomously. Use for adaptive workflows.
**Code-Based**: Explicit control flow. Use for predictable, deterministic workflows.
**Parallel**: Run multiple agents concurrently. Use for independent tasks.
**Agents as Tools**: Wrap agents as tools for manager LLM. Use for specialist delegation.
**Details**: Load `references/agent-patterRelated in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.