cloud-gpu-configs
# Cloud GPU Configurations Skill
What this skill does
# Cloud GPU Configurations Skill Platform-specific configuration templates and GPU selection guidance for Modal, Lambda Labs, and RunPod cloud platforms. --- name: cloud-gpu-configs description: Platform-specific configuration templates for Modal, Lambda Labs, and RunPod with GPU selection guides allowed-tools: Bash, Read, Write, Edit --- Use when: - Configuring cloud GPU platforms for ML training - Selecting appropriate GPU types for workloads - Setting up Modal, Lambda Labs, or RunPod environments - Optimizing cost vs performance for GPU compute - Generating platform-specific configuration files ## GPU Selection Guide ### Modal GPUs **Available GPUs**: T4, L4, A10, A100 (40GB/80GB), L40S, H100/H200, B200 **Selection Criteria**: - **T4**: Budget-friendly, light inference ($0.20-0.40/hr) - **L4**: Modern alternative to T4, better performance - **A10**: Good all-around option, up to 4 GPUs (24GB VRAM) - **L40S**: Excellent cost/performance for inference (48GB VRAM) - **A100**: Training standard, 40GB or 80GB variants - **H100/H200**: Cutting-edge performance, best software support - **B200**: Latest Blackwell architecture, premium performance **Multi-GPU Support**: - T4, L4, L40S, A100, H100, H200, B200: Up to 8 GPUs - A10: Up to 4 GPUs ### Lambda Labs GPUs **Available Instances**: - **1x A100 (40GB)**: Standard training, $1.10/hr - **1x A100 (80GB)**: Large models, $1.29/hr - **8x A100 (40GB)**: Distributed training, $8.80/hr - **8x A100 (80GB)**: Large-scale training, $10.32/hr - **1x H100**: Latest generation, $2.49/hr - **8x H100**: Maximum performance, $19.92/hr **Selection Criteria**: - **Single A100**: Most common training workload - **8x A100**: Multi-node distributed training - **H100**: When cutting-edge performance needed ### RunPod GPUs **Available GPUs**: RTX 3090, RTX 4090, A4000, A5000, A6000, A40, A100, H100 **Pricing Models**: - **Spot Instances**: 50-80% cheaper, can be interrupted - **On-Demand**: Guaranteed availability, higher cost **Selection Criteria**: - **RTX 3090/4090**: Consumer GPUs, excellent price/performance for small models - **A4000/A5000**: Professional GPUs, stable for production - **A6000**: 48GB VRAM, large model training - **A100**: Industry standard, 40GB or 80GB - **H100**: Premium performance ## Usage ### Setup Modal Environment ```bash bash scripts/setup-modal.sh ``` Prompts for: - Modal token - Default GPU type - Python version Creates: - `modal_image.py` - Configured Modal image - `.modal_config` - Environment configuration ### Setup Lambda Labs Environment ```bash bash scripts/setup-lambda.sh ``` Prompts for: - Lambda API key - SSH key path - Preferred instance type Creates: - `lambda_config.yaml` - Instance configuration - SSH configuration ### Setup RunPod Environment ```bash bash scripts/setup-runpod.sh ``` Prompts for: - RunPod API key - GPU type preference - Spot vs on-demand Creates: - `runpod_config.json` - Pod configuration - Environment setup script ## Templates ### Modal Image Template (`templates/modal_image.py`) Configurable Modal image with: - GPU selection - Python dependencies - System packages - CUDA configuration ### Lambda Config Template (`templates/lambda_config.yaml`) Instance configuration with: - Instance type selection - SSH key configuration - Startup scripts - Volume mounting ### RunPod Config Template (`templates/runpod_config.json`) Pod configuration with: - GPU type and count - Container image - Volume configuration - Network ports ## Examples - `examples/modal-t4-setup.md` - Budget-friendly Modal setup with T4 - `examples/lambda-a10-setup.md` - Standard Lambda Labs A100 configuration ## Cost Optimization Tips ### Modal - Use L40S for inference (best cost/performance) - Enable automatic upgrades (H100 → H200) for no extra cost - Use GPU fallback for faster scheduling - Avoid requesting >2 GPUs unless necessary ### Lambda Labs - Single A100 instances are most cost-effective for training - Use persistent storage to avoid re-downloading datasets - Terminate instances when not in use - Consider 8x A100 only for true distributed workloads ### RunPod - Use spot instances for interruptible workloads (50-80% savings) - RTX 4090 offers excellent value for smaller models - Use on-demand only for production/critical workloads - Enable auto-shutdown to prevent idle costs ## Performance Considerations ### Memory-Bound Operations - Consider total VRAM over compute power - L40S (48GB) often better than A100 (40GB) for inference - A100 80GB for large model training ### Compute-Bound Operations - H100/H200 for maximum throughput - B200 for latest architecture features - Multiple A100s for distributed training ### Multi-GPU Training - Use PyTorch DDP or DeepSpeed - Ensure efficient data loading (multiple workers) - Profile to avoid CPU bottlenecks - Consider network bandwidth between GPUs ## Integration with ML Training This skill integrates with other ml-training components: - **framework-templates**: Provides GPU configs for generated training scripts - **training-orchestrator**: Uses these configs for distributed training - **cost-estimator**: Uses pricing data for budget planning
Related in Cloud & DevOps
appbuilder-action-scaffolder
IncludedCreate, implement, deploy, and debug Adobe Runtime actions with consistent layout, validation, and error handling. Use this skill whenever the user needs to add actions to an App Builder project, understand action structure (params, response format, web/raw actions), configure actions in the manifest, use App Builder SDKs (State, Files, Events, database), deploy and invoke actions via CLI, debug action issues, or implement patterns such as webhook receivers, custom event providers, journaling consumers, large payload redirects, action sequence pipelines, and Asset Compute workers. Also trigger when users mention serverless functions in Adobe context, action logging, IMS authentication for actions, or cron-style scheduled actions.
orchestrating-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. Use this skill when the user needs a multi-step Data Cloud pipeline, cross-phase troubleshooting, or data space and data kit management. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase sf data360 workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching phase-specific skill), the task is STDM/session tracing/parquet telemetry (use observing-agentforce), standard CRM SOQL (use querying-soql), or Apex implementation (use generating-apex).
github-project-automation
IncludedAutomate GitHub repository setup with CI/CD workflows, issue templates, Dependabot, and CodeQL security scanning. Includes 12 production-tested workflows and prevents 18 errors: YAML syntax, action pinning, and configuration. Use when: setting up GitHub Actions CI/CD, creating issue/PR templates, enabling Dependabot or CodeQL scanning, deploying to Cloudflare Workers, implementing matrix testing, or troubleshooting YAML indentation, action version pinning, secrets syntax, runner versions, or CodeQL configuration. Keywords: github actions, github workflow, ci/cd, issue templates, pull request templates, dependabot, codeql, security scanning, yaml syntax, github automation, repository setup, workflow templates, github actions matrix, secrets management, branch protection, codeowners, github projects, continuous integration, continuous deployment, workflow syntax error, action version pinning, runner version, github context, yaml indentation error
sf-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase `sf data360` workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching sf-datacloud-* skill), the task is STDM/session tracing/parquet telemetry (use sf-ai-agentforce-observability), standard CRM SOQL (use sf-soql), or Apex implementation (use sf-apex).
fabric-cli
IncludedUse this skill for Fabric.so CLI workflows with the `fabric` terminal command: diagnose/install/login, search or browse a Fabric library, save notes/links/files, create folders, ask the Fabric AI assistant, manage tasks/workspaces, generate shell completion, check subscription usage, produce JSON output, and use Fabric as persistent agent memory. Do not use for Microsoft Fabric/Azure/Power BI `fab`, Daniel Miessler's Fabric framework, Python Fabric SSH, Fabric.js, or textile/fashion fabric.
lark
IncludedLark/Feishu CLI skills: lark-cli operations for docs, markdown, sheets, base, calendar, im, mail, task, okr, drive, wiki, slides, whiteboard, apps, approval, attendance, contact, vc, minutes, event. Use when the user needs to operate Lark/Feishu resources via lark-cli, send messages, manage documents, spreadsheets, calendars, tasks, OKRs, deploy web pages, or any Feishu/Lark workspace operations.