image-sprout
Use image-sprout to create and manage image generation projects with consistent style and subject identity. Use this skill when the user asks to generate images with saved context, iteratively refine image outputs, manage image projects with references and guides, or run image-sprout CLI commands in a workflow.
What this skill does
# image-sprout Generate and iterate on images with consistent style and subject identity. Image Sprout turns reusable project context — reference images, derived guides, and persistent instructions — into repeatable outputs you can build on across runs. Use this skill when: - The user asks to generate images and wants results that stay consistent across multiple runs - The user wants to iterate on outputs from a previous generation - The user needs to set up or manage an image project with references or style context - The workflow requires structured, scriptable image generation with `--json` output ## 1. Installation Install globally from npm: ```bash npm install -g image-sprout ``` Or run without installing — prefer this in Codex sandbox environments where PATH is unreliable: ```bash npx image-sprout help ``` All examples below use `npx image-sprout`. Substitute `image-sprout` directly if you have a global install. ## 2. OpenRouter Key Setup Image Sprout stores its OpenRouter key on disk. Set it once per machine: ```bash npx image-sprout config set apiKey <your-openrouter-key> npx image-sprout config show # confirm key is set (does not reveal the raw key) ``` ## 3. The Project Model Three context layers drive every generation: - **Visual Style** — consistent look and feel across outputs - **Subject Guide** — consistent subject identity across outputs - **Instructions** — persistent generation constraints (watermarks, framing, branding) Two reference pools: - **Shared refs** — drive both guides (default, simplest) - **Split refs** — separate style and subject pools (advanced; use `--role style` or `--role subject` when adding) Understanding this model prevents the most common agent mistake: generating without saved context and wondering why outputs are inconsistent. ## 4. Core CLI Workflow ```bash # Create a project npx image-sprout project create <name> # Add references (3+ recommended; more refs = better derivation) npx image-sprout ref add --project <name> ./ref1.png ./ref2.png ./ref3.png # Optional: persistent instructions npx image-sprout project update <name> --instructions "Watermark bottom-right: subtle." # Derive guides from refs npx image-sprout project derive <name> --target both # or: style, subject # Check readiness before generating npx image-sprout project status <name> --json # Generate (--count controls images per run: 1, 2, 4, 6; default is 4) npx image-sprout project generate <name> --prompt "hero in neon rain" npx image-sprout project generate <name> --prompt "hero in neon rain" --count 1 # Inspect results npx image-sprout run latest --project <name> --json # Delete a session and all its runs/images npx image-sprout session delete --project <name> <session-id> ``` Top-level aliases: ```bash npx image-sprout generate --project <name> --prompt "hero in neon rain" # same as project generate npx image-sprout analyze --project <name> --target both # same as project derive ``` ## 5. JSON Output — the Agent Pattern Always use `--json` for structured output: ```bash npx image-sprout project show <name> --json npx image-sprout project status <name> --json npx image-sprout run latest --project <name> --json npx image-sprout run list --project <name> --json --limit 5 ``` Use `--value PATH` to pluck a single field: ```bash npx image-sprout run latest --project <name> --json --value images[0].path ``` This is how agents hand image paths to downstream tools. Run images land in image-sprout's internal app data directory — use `run latest --json --value images[0].path` to get the path and leave what to do with it to the calling workflow. ## 6. Parallel-Safe Usage `npx image-sprout project use <name>` sets a shared "current project" state on disk. In multi-agent workflows, this state can collide across concurrent processes. **Always pass `--project <name>` explicitly** — never rely on the current project shortcut. ## 7. Web UI — Agent Awareness The web app runs over the same on-disk store as the CLI. Agents won't use it directly, but should surface it to users when interactive review is appropriate. ```bash npx image-sprout web # launches local app npx image-sprout web --open # also opens in default browser npx image-sprout web --port 8080 # custom port (default: 4310) ``` Useful for: - reviewing and comparing generated images visually - setting up a project interactively before handing off to CLI/agent use - iterating on outputs via the canvas interface **Security: do not expose the web UI to the public internet.** The server has no authentication. Safe options are localhost only, or a private network like Tailscale. The risk is public internet exposure — LAN and tailnet access are fine. ## 8. Model Management ```bash npx image-sprout model list npx image-sprout model set-default google/gemini-3.1-flash-image-preview npx image-sprout model add openai/gpt-5-image npx image-sprout model restore-defaults ``` Default generation model is **Nano Banana 2** (`google/gemini-3.1-flash-image-preview`). Custom models must accept image input and produce image output via OpenRouter. Guide derivation uses a separate configurable analysis model (default: `google/gemini-3.1-flash-image-preview`): ```bash # Set a persistent analysis model npx image-sprout config set analysisModel google/gemini-2.5-flash # Override per-derive npx image-sprout project derive <name> --target both --analysis-model google/gemini-2.5-flash ```
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.