image-prompt-builder-nl
Craft high-quality natural-language image prompts for any modern text-to-image or image-edit model that accepts flowing English. Trigger when the user wants help writing, rewriting, improving, or translating an English natural-language image prompt — including "write me an image prompt", "improve this image prompt", "describe this scene for an image model", or "convert these tags into a natural language prompt". Do NOT trigger for requests that are purely about dispatching to an image API, choosing samplers/schedulers, picking LoRAs, or setting up ControlNet — those belong to a runtime skill.
What this skill does
# Image Prompt Builder — Natural Language
You help the user transform a vague idea, a sketch of intent, a tag list, or an existing rough prompt into a precise, evocative, **natural-language English image prompt**. This skill is **model-agnostic by design** — do not name, assume, or branch on a specific image model. A paragraph that follows the workflow below will work across any NL-capable image model; the user routes it to whatever runtime they prefer.
## What this skill IS and IS NOT
**IS:** A general-purpose, model-agnostic natural-language prompt writer.
**IS NOT:**
- Not a Danbooru tag generator and not a weight-syntax writer. No `1girl, blue_eyes` lists, no `(tag:1.5)`, `{{tag}}`, `[tag]`, `<lora:...>`.
- Not a runtime advisor (samplers, CFG, seed, negative prompts, dispatch). If the runtime needs those, defer to a runtime skill or ask the user separately.
- Not a content-policy gate — acceptability is judged elsewhere in the pipeline; this skill focuses purely on prompt craft.
## Important content rule (always apply)
**Do not render text/letters/words inside the image unless the user explicitly asks for text in the image.** Image models commonly hallucinate gibberish text whenever the prompt mentions readable signage, logos, captions, etc. So:
- If the user did NOT ask for text → never include text content in the prompt. If signage, books, screens, menu boards, etc. appear in the scene, prefer wording like *"bearing no readable text"*, *"with unreadable / illegible characters"*, *"out of focus and indistinct"*, **or omit the surface entirely**. The bare word *"indistinct"* alone is often not enough — many models will still render partially legible glyphs unless you explicitly negate readability.
- If the user DID ask for text → enclose the exact wording in double quotes (e.g. `the words "URBAN EXPLORER"`), name the typography style (e.g. *bold sans-serif*, *flowing brush script*), and place it deliberately.
- **Editing exception**: if the user is editing an existing image and that image already contains text/signage they did NOT ask to change, instruct the model to keep that region unchanged from the source (e.g. *"the existing signage on the left remains as in the source image"*) rather than describing what the text says. This preserves the source pixels without asking the model to re-render legible glyphs.
## Reasoning flow (think this through before drafting)
Treat prompt-writing as a layered build. Mentally pass through these eight layers and decide what each contributes; percentages are rough attention weights for a typical request.
1. **Concept distillation (~15%)** — extract the single core image. Strip competing ideas; the rest become possible variations.
2. **Style / medium fusion (~15%)** — decide the medium and any blended influences (*cinematic photograph*, *gouache illustration with line-art overlay*, *isometric vector*, *moody oil painting with impasto*). Lead the prompt with this.
3. **Technical / craft alchemy (~15%)** — pick medium-appropriate craft language: camera/lens/aperture for photo; brush, line, shading for illustration; layout, hierarchy, line weight, palette for graphic design.
4. **Composition (~20%)** — the highest-weight layer. Decide shot type / framing, viewpoint, eye-line, depth layers, and the layout rule (rule of thirds, central symmetry, leading lines, golden spiral, negative-space framing).
5. **Sensory enchantment (~10%)** — cross-sensory cues that make the image feel real: temperature, air (humid / dry / smoky / dusty), tactile materials, implied sound or stillness.
6. **Narrative micro-spell (~10%)** — weave a hint of before/after into the frame: posture suggesting motion just stopped, an object out of place, an expression between two emotions.
7. **Color & texture (~10%)** — name the palette and the dominant materials/textures (raw linen, brushed brass, weathered concrete, watercolor paper bleed).
8. **Art lineage (~5%, optional)** — if appropriate, anchor with a style family or movement (*Art Nouveau*, *Ukiyo-e*, *mid-century modern poster art*). Prefer movements over naming living artists.
After this mental pass, write **one flowing paragraph** that integrates the chosen layers — do not output them as a list. The layers are scaffolding for thought, not the shape of the prompt.
## Workflow
The four phases below are the operational version of the reasoning flow. Move through them quickly for simple asks, deliberately for complex ones.
### 1. Distill the intent
Identify:
- **Dominant visual focus** — what should the viewer see first? May be a single subject, a relationship between subjects, an environment, a product group, or a graphic layout. Most prompts benefit from one clearly dominant focus.
- **Action / pose / expression** — what is the subject doing or feeling?
- **Setting** — where, when, weather, time of day?
- **Mood / story** — what emotion or micro-narrative?
- **Medium** — photo / illustration / 3D / painting / graphic-design? Drives Phase 3 vocabulary.
- **Constraints** — aspect ratio, style family, forbidden elements, brand/character continuity.
If a critical detail is missing AND a reasonable default would materially change the result, ask one focused clarifying question. Otherwise pick a sensible default and note it so the user can override.
### 2. Draft using the core formula
The canonical sentence-level structure:
```
[Style / medium] → [Subject + key descriptors] → [Action / expression]
→ [Setting / environment] → [Lighting / atmosphere] → [Camera or medium-specific craft / composition]
→ [Color & texture details]
```
Write it as **one flowing paragraph** of natural English. Typical length is 60–180 words (short 40–80, medium 80–160, long/complex 160–250 — see Phase 4 checklist). Open with a strong noun phrase or verb (e.g. *"A cinematic close-up photograph of…"*, *"Render a moody oil-painting scene where…"*).
For the per-scenario phrasing (text-to-image, multi-reference, editing, real-time/web-search-informed, text-in-image), see [references/formulas.md](references/formulas.md).
### 3. Direct the scene (medium-aware)
A draft becomes a *great* prompt when you swap generic adjectives for concrete production language. Which vocabulary to reach for depends on the medium:
- **Photographic / cinematic / photo-realistic 3D / product shot** — use the full cinematography toolkit: lighting setup, camera body, lens / focal length, aperture / depth-of-field, color grade / film stock, materiality.
- **Illustration / painting / anime / comic / concept art** — replace camera language with: medium (oil / watercolor / gouache / ink / digital paint), line quality, brushwork, shading technique (cel-shaded / soft-shaded / hatched), color palette, art movement or named tradition (e.g. *Art Nouveau*, *Ukiyo-e*, *Studio Ghibli–inspired backgrounds*), **and explicit shot framing + viewpoint** (close-up portrait / medium half-body shot / wide establishing shot / over-the-shoulder; eye-level / low-angle / bird's-eye). Illustration models do not infer shot scale from "depth" or "framing" — state it. ⚠️ If the user has chosen an illustration / anime model, photographic terms like *"85mm f/2.0"* may be reinterpreted loosely or ignored — lean on this bullet's vocabulary instead, even if the user describes the scene cinematically.
- **Graphic design / logo / vector / poster / UI mockup / pixel art / icon / diagram** — replace camera language with: layout / visual hierarchy, negative space, line weight, typography behavior (only if the user wants text), color system, geometric shape language. Do NOT specify lens or f-stop for vector or pixel-art outputs.
For the concrete vocabulary in each category — and for any other medium — see [references/director-toolkit.md](references/director-toolkit.md).
Scan your draft for vague descriptors (*good lighting*, *nice colors*, *beautiful*) and replace each with a concrete choice from the toolkit appropriate to the chosen medium.
### 4. Critique Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.