prompt-videos
Prompting techniques for AI video generation models on Replicate. Use when writing prompts for video models or building video generation features.
What this skill does
# Prompting video models on Replicate
Distilled from Replicate's blog posts on prompting video models (2025-2026). Techniques are model-agnostic and focus on transferable principles.
## Choose a model with the API, not from memory
This skill describes general prompting techniques. To choose a model, use the [find-models](../find-models/SKILL.md) skill and query the Replicate API. The video model landscape changes weekly. Don't assume specific models exist or are still state-of-the-art based on names you've seen before. Always search the API for current options, then read the schema before running anything.
For pricing and feature comparison, see the [compare-models](../compare-models/SKILL.md) skill.
## Scene description
A good video prompt is a scene description, not a caption. Write what happens, where, and how it looks.
### Layer these elements into every prompt
1. **Subject**: Who or what is in the scene (a person, animal, object, landscape).
2. **Context**: Where the subject is (indoors, a city street, a forest, a spaceship corridor).
3. **Action**: What the subject does (walks, turns, picks up a phone, runs).
4. **Style**: The visual aesthetic (cinematic, animated, stop-motion, documentary).
5. **Camera**: How the camera moves (dolly shot, tracking, static, handheld).
6. **Composition**: How the shot is framed (wide shot, close-up, over-the-shoulder).
7. **Ambiance**: Mood and lighting (warm tones, blue light, golden hour, overcast).
### Be specific, not vague
Vague: "A car chase"
Specific: "A high-speed car chase on a rain-drenched highway at night. Two muscle cars weave through heavy traffic at 140mph, headlights slicing through the downpour. One car clips a semi-truck sending sparks showering across six lanes. Tires hydroplane on standing water. Neon highway signs blur overhead."
### Overdescribe
Modern video models handle long, dense prompts well. Don't write "a man on the phone." Write "a desperate man in a weathered green trench coat picks up a rotary phone mounted on a gritty brick wall, bathed in the eerie glow of a green neon sign." Every concrete detail you add gives the model less room to improvise poorly.
### Name subjects directly
Use descriptive phrases like "the woman in the red jacket" or "the bearded man in flannel." Avoid pronouns, which are ambiguous to video models just as they are to image models.
## Camera and cinematography
Video models understand filmmaking language. Use it to direct the shot rather than hoping for good framing.
### Shot types
Use standard shot terminology to control framing:
- Wide/establishing shot: shows the full scene and environment
- Medium shot: frames the subject from roughly the waist up
- Close-up: fills the frame with the subject's face or a key object
- Extreme close-up: isolates a detail (an eye, a hand gripping a handle, a drop of water)
### Camera motion
Describe how the camera moves:
- Static/tripod: locked-off, no movement
- Pan: horizontal rotation left or right
- Tilt: vertical rotation up or down
- Dolly: camera physically moves toward or away from the subject
- Tracking: camera moves alongside the subject
- Crane: camera rises or descends vertically
- Handheld: shaky, documentary-style movement
- Drone/aerial: overhead or sweeping bird's-eye shots
- Dolly zoom (Hitchcock/vertigo effect): background stretches while subject stays locked
### Camera position
Specify the camera's height and angle:
- Eye level: neutral, natural perspective
- Low angle / worm's eye: looking up at the subject (makes subjects feel powerful or imposing)
- High angle / bird's eye: looking down (makes subjects feel small or vulnerable)
- Over-the-shoulder: frames one subject from behind another
- POV / first-person: camera is the subject's eyes
### Lens and focus language
- Shallow depth of field: subject sharp, background blurred
- Deep focus: everything sharp from foreground to background
- Macro lens: extreme close-up with shallow focus
- Wide-angle lens: exaggerated perspective, more environment visible
- Tilt-shift: miniature effect, selective focus band
### Escalation pattern
A natural progression for short clips is wide > medium > close-up > extreme close-up. This maps well onto 8-15 second clips and gives the model clear structure. For example:
- 0-3s: wide establishing shot of the location
- 3-7s: medium shot, the subject enters or acts
- 7-12s: close-up on the key moment
- 12-15s: extreme close-up on a detail (a hand, an eye, a drop of rain)
## Audio and dialogue
Many video models generate audio natively alongside the visuals. If you don't prompt for the audio you want, the model will guess, and it often guesses wrong.
### Prompt all four audio layers
1. **Dialogue**: What characters say, either exact words or described intent.
2. **Ambient sound**: The background audio of the scene (rain on metal awnings, city traffic, forest birds).
3. **Sound effects**: Specific sounds from actions (a door slamming, glass breaking, a sword being drawn).
4. **Music**: Genre, mood, and instrumentation (a tense cinematic score, soft jazz piano, no music).
If you skip ambient audio, models may hallucinate inappropriate sounds. A common failure mode is adding a "live studio audience" laughing in the background. Prevent this by describing the soundscape explicitly: "sounds of distant bands, noisy crowd, ambient background of a busy festival field."
### Dialogue prompting
There are two approaches:
- Explicit: "The man says: My name is Ben." This gives you exact control over the words.
- Implicit: "The man introduces himself." This lets the model decide the phrasing.
Explicit dialogue should be short enough to fit the clip duration. Packing too much dialogue into an 8-second clip produces unnaturally fast speech. Too little dialogue can produce awkward silence or AI gibberish.
### Syntax that avoids subtitles
Many video models were trained on videos with baked-in subtitles and will add them to outputs. To prevent this:
- Use a colon for dialogue: "She says: Hello there" rather than "She says 'Hello there'"
- Add "(no subtitles)" to the prompt
- If subtitles persist, repeat the instruction: "No subtitles. No subtitles!"
### Pronunciation
If a model mispronounces a name or word, spell it phonetically in the prompt. For example, write "foh-fur" instead of "fofr" or "Shreedar" instead of "Shridhar."
### Who says what
In multi-character scenes, the model can mix up who says what. Tie dialogue to distinctive visual descriptions: "The woman wearing pink says: ..." and "The man with glasses replies: ..."
## Multi-shot and time-coded prompting
Some models support generating multiple shots within a single clip (up to ~15 seconds). You can direct each shot individually using time codes.
### Time-coded format
Write timestamps directly into the prompt:
```
[0-4s]: Wide establishing shot, static camera, misty bamboo forest at dawn
[4-9s]: Medium shot, slow push-in, the fighter steps forward
[9-15s]: Close-up, orbit shot, the fighter strikes, slow motion
```
Each shot should specify:
- Camera position and motion
- Subject action
- Lighting or mood shifts
### Transition language
Use explicit transition instructions between shots:
- "Hard cut to..." for an abrupt switch
- "Seamless morph into..." for a fluid transition
- "Whip pan to..." for a fast, energetic cut
- "Snap cut to..." for a jarring, dramatic shift
Without explicit transitions, the model improvises, which may or may not match your intent.
### Example: multi-shot commercial
```
(0-3s) Macro shot of a luxury perfume bottle among scattered pink peonies,
shallow depth of field, petals floating in warm afternoon light,
soft ambient music.
(3-7s) Camera glides closer, a feminine hand enters frame from the right,
fingers gently touch the glass bottle, the sound of silk rustling.
(7-12s) Hard cut to slow-motion spray, golden mist diffuses through the air,
particles catching rim light against Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.