image-processing
Image processing and manipulation
What this skill does
# Image Processing
Image manipulation using Python and Pillow (PIL).
## Resize image
```bash
python3 -c "
from PIL import Image
img = Image.open('input.jpg')
# Resize to specific dimensions
resized = img.resize((800, 600), Image.LANCZOS)
resized.save('resized.jpg')
# Resize keeping aspect ratio
img.thumbnail((800, 800), Image.LANCZOS)
img.save('thumbnail.jpg')
print('Done')
"
```
## Convert format
```bash
python3 -c "
from PIL import Image
img = Image.open('input.png')
# PNG to JPEG (must convert RGBA to RGB)
if img.mode == 'RGBA':
img = img.convert('RGB')
img.save('output.jpg', 'JPEG', quality=90)
# JPEG to PNG
# img = Image.open('input.jpg')
# img.save('output.png', 'PNG')
# Convert to WebP
# img.save('output.webp', 'WEBP', quality=85)
print('Converted successfully')
"
```
## Crop image
```bash
python3 -c "
from PIL import Image
img = Image.open('input.jpg')
# Crop box: (left, upper, right, lower)
cropped = img.crop((100, 100, 500, 400))
cropped.save('cropped.jpg')
print(f'Cropped from {img.size} to {cropped.size}')
"
```
## Add text overlay
```bash
python3 -c "
from PIL import Image, ImageDraw, ImageFont
img = Image.open('input.jpg')
draw = ImageDraw.Draw(img)
# Use default font (or specify a TTF path)
try:
font = ImageFont.truetype('/System/Library/Fonts/Helvetica.ttc', 36)
except OSError:
font = ImageFont.load_default()
text = 'Hello World'
# Draw text with shadow for readability
draw.text((12, 12), text, fill='black', font=font)
draw.text((10, 10), text, fill='white', font=font)
img.save('with_text.jpg')
print('Text overlay added')
"
```
## Create thumbnail
```bash
python3 -c "
from PIL import Image
img = Image.open('input.jpg')
sizes = [(128, 128), (256, 256), (512, 512)]
for size in sizes:
thumb = img.copy()
thumb.thumbnail(size, Image.LANCZOS)
thumb.save(f'thumb_{size[0]}x{size[1]}.jpg')
print(f'Created thumb_{size[0]}x{size[1]}.jpg')
"
```
## Get image metadata
```bash
python3 -c "
from PIL import Image
from PIL.ExifTags import TAGS
import json
img = Image.open('input.jpg')
info = {
'format': img.format,
'mode': img.mode,
'size': {'width': img.size[0], 'height': img.size[1]},
}
# Extract EXIF data if present
exif = img.getexif()
if exif:
info['exif'] = {}
for tag_id, value in exif.items():
tag = TAGS.get(tag_id, tag_id)
try:
info['exif'][str(tag)] = str(value)
except Exception:
pass
print(json.dumps(info, indent=2))
"
```
## Optimize/compress
```bash
python3 -c "
from PIL import Image
import os
img = Image.open('input.jpg')
original_size = os.path.getsize('input.jpg')
# Optimize JPEG
if img.mode == 'RGBA':
img = img.convert('RGB')
img.save('optimized.jpg', 'JPEG', quality=80, optimize=True)
new_size = os.path.getsize('optimized.jpg')
ratio = (1 - new_size / original_size) * 100
print(f'Original: {original_size:,} bytes')
print(f'Optimized: {new_size:,} bytes')
print(f'Saved: {ratio:.1f}%')
"
```
## Batch process
```bash
python3 -c "
from PIL import Image
import os, glob
input_dir = 'images'
output_dir = 'processed'
os.makedirs(output_dir, exist_ok=True)
for filepath in glob.glob(os.path.join(input_dir, '*')):
try:
img = Image.open(filepath)
# Resize all to max 1024px wide
if img.width > 1024:
ratio = 1024 / img.width
new_size = (1024, int(img.height * ratio))
img = img.resize(new_size, Image.LANCZOS)
if img.mode == 'RGBA':
img = img.convert('RGB')
name = os.path.splitext(os.path.basename(filepath))[0]
out_path = os.path.join(output_dir, f'{name}.jpg')
img.save(out_path, 'JPEG', quality=85, optimize=True)
print(f'Processed: {filepath} -> {out_path}')
except Exception as e:
print(f'Skipped {filepath}: {e}')
"
```
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.