fal-text-to-image
Generate, remix, and edit images using fal.ai's AI models. Supports text-to-image generation, image-to-image remixing, and targeted inpainting/editing.
What this skill does
# fal.ai Image Generation & Editing Skill Professional AI-powered image workflows using fal.ai's state-of-the-art models including FLUX, Recraft V3, Imagen4, and more. ## Three Modes of Operation ### 1. Text-to-Image (fal-text-to-image) Generate images from scratch using text prompts ### 2. Image Remix (fal-image-remix) Transform existing images while preserving composition ### 3. Image Edit (fal-image-edit) Targeted inpainting and masked editing ## When to Use This Skill Trigger when user: - Requests image generation from text descriptions - Wants to transform/remix existing images with AI - Needs to edit specific regions of images (inpainting) - Wants to create images with specific styles (vector, realistic, typography) - Needs high-resolution professional images (up to 2K) - Wants to use a reference image for style transfer - Mentions specific models like FLUX, Recraft, or Imagen - Asks for logo, poster, or brand-style image generation - Needs object removal or targeted modifications ## Quick Start ### Text-to-Image: Generate from Scratch ```bash # Basic generation uv run python fal-text-to-image "A cyberpunk city at sunset with neon lights" # With specific model uv run python fal-text-to-image -m flux-pro/v1.1-ultra "Professional headshot" # With style reference uv run python fal-text-to-image -i reference.jpg "Mountain landscape" -m flux-2/lora/edit ``` ### Image Remix: Transform Existing Images ```bash # Transform style while preserving composition uv run python fal-image-remix input.jpg "Transform into oil painting" # With strength control (0.0=original, 1.0=full transformation) uv run python fal-image-remix photo.jpg "Anime style character" --strength 0.6 # Premium quality remix uv run python fal-image-remix -m flux-1.1-pro image.jpg "Professional portrait" ``` ### Image Edit: Targeted Modifications ```bash # Edit with mask image (white=edit area, black=preserve) uv run python fal-image-edit input.jpg mask.png "Replace with flowers" # Auto-generate mask from text uv run python fal-image-edit input.jpg --mask-prompt "sky" "Make it sunset" # Remove objects uv run python fal-image-edit photo.jpg mask.png "Remove object" --strength 1.0 # General editing (no mask) uv run python fal-image-edit photo.jpg "Enhance lighting and colors" ``` ## Model Selection Guide The script intelligently selects the best model based on task context: ### **flux-pro/v1.1-ultra** (Default for High-Res) - **Best for**: Professional photography, high-resolution outputs (up to 2K) - **Strengths**: Photo realism, professional quality - **Use when**: User needs publication-ready images - **Endpoint**: `fal-ai/flux-pro/v1.1-ultra` ### **recraft/v3/text-to-image** (SOTA Quality) - **Best for**: Typography, vector art, brand-style images, long text - **Strengths**: Industry-leading benchmark scores, precise text rendering - **Use when**: Creating logos, posters, or text-heavy designs - **Endpoint**: `fal-ai/recraft/v3/text-to-image` ### **flux-2** (Best Balance) - **Best for**: General-purpose image generation - **Strengths**: Enhanced realism, crisp text, native editing - **Use when**: Standard image generation needs - **Endpoint**: `fal-ai/flux-2` ### **flux-2/lora** (Custom Styles) - **Best for**: Domain-specific styles, fine-tuned variations - **Strengths**: Custom style adaptation - **Use when**: User wants specific artistic styles - **Endpoint**: `fal-ai/flux-2/lora` ### **flux-2/lora/edit** (Style Transfer) - **Best for**: Image-to-image editing with style references - **Strengths**: Specialized style transfer - **Use when**: User provides reference image with `-i` flag - **Endpoint**: `fal-ai/flux-2/lora/edit` ### **imagen4/preview** (Google Quality) - **Best for**: High-quality general images - **Strengths**: Google's highest quality model - **Use when**: User specifically requests Imagen or Google models - **Endpoint**: `fal-ai/imagen4/preview` ### **stable-diffusion-v35-large** (Typography & Style) - **Best for**: Complex prompts, typography, style control - **Strengths**: Advanced prompt understanding, resource efficiency - **Use when**: Complex multi-element compositions - **Endpoint**: `fal-ai/stable-diffusion-v35-large` ### **ideogram/v2** (Typography Specialist) - **Best for**: Posters, logos, text-heavy designs - **Strengths**: Exceptional typography, realistic outputs - **Use when**: Text accuracy is critical - **Endpoint**: `fal-ai/ideogram/v2` ### **bria/text-to-image/3.2** (Commercial Safe) - **Best for**: Commercial projects requiring licensed training data - **Strengths**: Safe for commercial use, excellent text rendering - **Use when**: Legal/licensing concerns matter - **Endpoint**: `fal-ai/bria/text-to-image/3.2` ## Command-Line Interface ```bash uv run python fal-text-to-image [OPTIONS] PROMPT Arguments: PROMPT Text description of the image to generate Options: -m, --model TEXT Model to use (see model list above) -i, --image TEXT Path or URL to reference image for style transfer -o, --output TEXT Output filename (default: generated_image.png) -s, --size TEXT Image size (e.g., "1024x1024", "landscape_16_9") --seed INTEGER Random seed for reproducibility --steps INTEGER Number of inference steps (model-dependent) --guidance FLOAT Guidance scale (higher = more prompt adherence) --help Show this message and exit ``` ## Authentication Setup Before first use, set your fal.ai API key: ```bash export FAL_KEY="your-api-key-here" ``` Or create a `.env` file in the skill directory: ```env FAL_KEY=your-api-key-here ``` Get your API key from: https://fal.ai/dashboard/keys ## Advanced Examples ### High-Resolution Professional Photo ```bash uv run python fal-text-to-image \ -m flux-pro/v1.1-ultra \ "Professional headshot of a business executive in modern office" \ -s 2048x2048 ``` ### Logo/Typography Design ```bash uv run python fal-text-to-image \ -m recraft/v3/text-to-image \ "Modern tech startup logo with text 'AI Labs' in minimalist style" ``` ### Style Transfer from Reference ```bash uv run python fal-text-to-image \ -m flux-2/lora/edit \ -i artistic_style.jpg \ "Portrait of a woman in a garden" ``` ### Reproducible Generation ```bash uv run python fal-text-to-image \ -m flux-2 \ --seed 42 \ "Futuristic cityscape with flying cars" ``` ## Model Selection Logic The script automatically selects the best model when `-m` is not specified: 1. **If `-i` provided**: Uses `flux-2/lora/edit` for style transfer 2. **If prompt contains typography keywords** (logo, text, poster, sign): Uses `recraft/v3/text-to-image` 3. **If prompt suggests high-res needs** (professional, portrait, headshot): Uses `flux-pro/v1.1-ultra` 4. **If prompt mentions vector/brand**: Uses `recraft/v3/text-to-image` 5. **Default**: Uses `flux-2` for general purpose ## Output Format Generated images are saved with metadata: - Filename includes timestamp and model name - EXIF data stores prompt, model, and parameters - Console displays generation time and cost estimate ## Troubleshooting | Problem | Solution | |---------|----------| | `FAL_KEY not set` | Export FAL_KEY environment variable or create .env file | | `Model not found` | Check model name against supported list | | `Image reference fails` | Ensure image path/URL is accessible | | `Generation timeout` | Some models take longer; wait or try faster model | | `Rate limit error` | Check fal.ai dashboard for usage limits | ## Cost Optimization - **Free tier**: FLUX.2 offers 100 free requests (expires Dec 25, 2025) - **Pay per use**: FLUX Pro charges per megapixel - **Budget option**: Use `flux-2` or `stable-diffusion-v35-large` for general use - **Premium**: Use `flux-pro/v1.1-ultra` only when high-res is required ## Image Remix: Model Selection Guide Available models for image-to-image remixing: ### **flux-2/dev** (Default, Free) - **Best for**: General re
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.