gemini-image
Analyze images using Gemini's vision capabilities. Use for image analysis, text extraction from screenshots, and visual content understanding.
What this skill does
# Gemini Image Analysis Analyze images using Gemini Pro's vision capabilities. ## Prerequisites ```bash pip install google-generativeai export GEMINI_API_KEY=your_api_key ``` ## CLI Reference ### Basic Image Analysis ```bash # Analyze an image gemini -m pro -f /path/to/image.png "Describe this image in detail" # With specific question gemini -m pro -f screenshot.png "What error message is shown?" # Multiple images gemini -m pro -f image1.png -f image2.png "Compare these two images" ``` ## Analysis Operations ### General Description ```bash gemini -m pro -f image.png "Describe this image comprehensively: 1. Main subject/content 2. Colors and composition 3. Text visible (if any) 4. Context and purpose 5. Notable details" ``` ### Extract Text (OCR) ```bash gemini -m pro -f screenshot.png "Extract all text from this image. Format as plain text, preserving layout where possible. Include any text in buttons, labels, or UI elements." ``` ### Code from Screenshot ```bash gemini -m pro -f code-screenshot.png "Extract the code from this screenshot. Provide as properly formatted code with correct indentation. Note any parts that are unclear or partially visible." ``` ### UI Analysis ```bash gemini -m pro -f ui-screenshot.png "Analyze this UI: 1. What application/website is this? 2. What page/screen is shown? 3. Main UI elements and their purpose 4. User flow/actions available 5. Any UX issues or suggestions" ``` ### Error Analysis ```bash gemini -m pro -f error-screenshot.png "Analyze this error: 1. What error is shown? 2. What is the likely cause? 3. How to fix it? 4. Any related information visible?" ``` ### Diagram Understanding ```bash gemini -m pro -f diagram.png "Explain this diagram: 1. What type of diagram is this? 2. Main components and their relationships 3. Data/process flow 4. Key takeaways" ``` ## Specific Use Cases ### Debug Screenshot ```bash gemini -m pro -f debug-screen.png "I'm debugging an issue. From this screenshot: 1. What is the current state? 2. What errors or warnings are visible? 3. What should I look at? 4. Suggested next steps" ``` ### Compare Before/After ```bash gemini -m pro -f before.png -f after.png "Compare these before and after images: 1. What changed? 2. Is this an improvement? 3. Any issues in the 'after' version? 4. Anything missing?" ``` ### Design Feedback ```bash gemini -m pro -f design.png "Provide design feedback: 1. Visual hierarchy 2. Color usage 3. Typography 4. Spacing and alignment 5. Accessibility concerns 6. Suggestions for improvement" ``` ### Data Extraction ```bash gemini -m pro -f chart.png "Extract data from this chart: 1. Chart type 2. Data series and values 3. Axes labels and ranges 4. Key trends or insights 5. Output as structured data if possible" ``` ### Form Analysis ```bash gemini -m pro -f form.png "Analyze this form: 1. Form purpose 2. Fields and their types 3. Required vs optional 4. Validation rules visible 5. UX suggestions" ``` ## Workflow Patterns ### Screenshot to Issue ```bash # Capture screenshot (macOS) screencapture -i /tmp/bug.png # Analyze and format as issue gemini -m pro -f /tmp/bug.png "Create a bug report from this screenshot: ## Summary [One-line description] ## Steps to Reproduce [Inferred from screenshot] ## Expected Behavior [What should happen] ## Actual Behavior [What the screenshot shows] ## Environment [Any visible system info]" ``` ### UI to Code ```bash gemini -m pro -f ui-design.png "Generate React component code that recreates this UI: - Use Tailwind CSS for styling - Make it responsive - Include proper TypeScript types - Add appropriate accessibility attributes" ``` ### Documentation ```bash gemini -m pro -f app-screen.png "Write user documentation for this screen: - What this screen is for - How to use each feature - Common tasks - Tips and notes" ``` ## Image Types Supported - PNG, JPEG, GIF, WebP - Screenshots - Photos - Diagrams and charts - UI mockups - Code snippets - Documents ## Best Practices 1. **Use clear images** - Higher quality = better analysis 2. **Crop to relevant area** - Remove unnecessary context 3. **Ask specific questions** - Vague prompts get vague answers 4. **Provide context** - Tell Gemini what you're looking for 5. **Verify extracted text** - OCR isn't perfect 6. **Multiple angles** - Use multiple images for complex subjects
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.