translate-video
Translate video subtitles to any language with native-quality refinement. Full pipeline: transcribe → translate → refine → embed RTL-safe subtitles. Use for: translate video, תרגם סרטון, video translation, foreign subtitles, Hebrew subtitles, translated captions.
What this skill does
# Translate Video
End-to-end video translation pipeline. Takes a video, transcribes it, translates to target language with native-speaker refinement, and burns subtitles onto the video.
## Usage
```
/translate-video /path/to/video.mp4 he
/translate-video /path/to/video.mp4 ar
/translate-video /path/to/video.mp4 es
```
**Arguments:**
- `$1` - Path to video file (required)
- `$2` - Target language code (default: `he`). Codes: he, ar, es, fr, de, ru, zh, ja, etc.
## Pipeline
### Step 1: Transcribe (via `/transcribe` skill)
Extract audio if video is too large (>25MB audio), then transcribe:
```bash
# Extract audio if needed (reduces upload size)
ffmpeg -i "$VIDEO" -vn -acodec libmp3lame -ab 128k "$AUDIO_PATH" -y
# Transcribe using the transcribe skill's script
cd ~/.claude/skills/transcribe/scripts && [ -d node_modules ] || npm install --silent
npx ts-node transcribe.ts -i "$INPUT" -o "$SRT_PATH"
```
This generates:
- `{basename}.srt` - Raw SRT file
- `{basename}.md` - Readable text
### Step 2: Translate
Read the `.md` file to understand full context, then translate the `.srt` file.
**Translation rules:**
- Translate ALL subtitle text entries, preserving SRT index numbers and timestamps exactly
- Do NOT translate proper nouns (product names, people's names, brand names)
- Keep technical terms that are commonly used untranslated in the target language (API, CLI, SaaS, etc.)
- Adapt idioms and expressions to natural equivalents in the target language
- Match the speaker's register (casual/formal) in the translation
### Step 3: Semantic Refinement
The raw transcription chunks by time, not meaning. Regroup for the target language:
1. **Read all translated text** as continuous prose
2. **Identify natural sentence/clause boundaries** in the target language
3. **Regroup words** into semantically coherent subtitle entries (max 2 lines per entry, ~40 chars per line)
4. **Adjust timestamps**: each entry starts at first word's original time, ends at last word's time
5. **Merge fragmented entries** - aim for 150-250 entries for a ~15min video (vs 500-600 raw)
6. **Native speaker test** - read each subtitle aloud. If it sounds awkward, rephrase
**Quality checklist:**
- [ ] No sentence split mid-clause
- [ ] No orphaned words (single word carrying over from previous thought)
- [ ] Punctuation at natural break points
- [ ] Reading pace is comfortable (not too much text per subtitle)
- [ ] Sounds like a native speaker wrote it, not a translation
### Step 4: Embed Subtitles (via `/embed-subtitles` skill)
**RTL is handled automatically** by the `embed-subtitles` skill - it detects RTL content and applies Unicode directional marks before embedding. No need to handle RTL here.
Burn the translated SRT onto the video:
```bash
cd ~/.claude/skills/embed-subtitles/scripts
npx ts-node embed-subtitles.ts \
-i "$VIDEO" \
-s "$TRANSLATED_SRT" \
-o "$OUTPUT" \
--font-size 24 --margin 30
```
Or directly with FFmpeg:
```bash
ffmpeg -i "$VIDEO" \
-vf "subtitles='$TRANSLATED_SRT':force_style='FontName=Arial,FontSize=24,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,Outline=2,Shadow=1,Alignment=2,MarginV=30'" \
-c:v libx264 -preset fast -crf 23 -c:a copy \
"$OUTPUT" -y
```
### Step 5: Open Result
```bash
open "$OUTPUT" # macOS
```
## Output Files
All files saved next to the original video:
| File | Description |
|------|-------------|
| `{name}.srt` | Original language SRT |
| `{name}.md` | Original readable transcript |
| `{name}_{lang}.srt` | Translated + refined SRT (with RTL marks if applicable) |
| `{name}_{lang}_subtitled.mp4` | Final video with burned-in subtitles |
## Language Codes
| Code | Language | RTL? |
|------|----------|------|
| `he` | Hebrew | Yes |
| `ar` | Arabic | Yes |
| `fa` | Farsi | Yes |
| `en` | English | No |
| `es` | Spanish | No |
| `fr` | French | No |
| `de` | German | No |
| `ru` | Russian | No |
| `zh` | Chinese | No |
| `ja` | Japanese | No |
| `pt` | Portuguese | No |
| `it` | Italian | No |
## Examples
**Hebrew translation (most common use):**
```
/translate-video ~/Desktop/tutorial.mp4 he
```
Produces: `tutorial_he.srt` + `tutorial_he_subtitled.mp4`
**Spanish translation:**
```
/translate-video ~/Desktop/talk.mp4 es
```
**Default (Hebrew):**
```
/translate-video ~/Desktop/video.mp4
```
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.