ml-deployment
Deploy ML models to production - APIs, containerization, monitoring, and MLOps
What this skill does
# ML Deployment Skill
> Take models from development to production.
## Quick Start
```python
from fastapi import FastAPI
from pydantic import BaseModel
import numpy as np
import joblib
app = FastAPI(title="ML Model API")
model = joblib.load('model.pkl')
class PredictRequest(BaseModel):
features: list[float]
class PredictResponse(BaseModel):
prediction: float
@app.post("/predict", response_model=PredictResponse)
async def predict(request: PredictRequest):
X = np.array([request.features])
prediction = model.predict(X)[0]
return PredictResponse(prediction=float(prediction))
@app.get("/health")
async def health():
return {"status": "healthy"}
```
## Key Topics
### 1. Model Export
```python
import torch
import torch.onnx
# Export PyTorch to ONNX
def export_to_onnx(model, sample_input, path='model.onnx'):
model.eval()
torch.onnx.export(
model,
sample_input,
path,
export_params=True,
opset_version=14,
input_names=['input'],
output_names=['output'],
dynamic_axes={'input': {0: 'batch'}, 'output': {0: 'batch'}}
)
# ONNX inference
import onnxruntime as ort
session = ort.InferenceSession('model.onnx')
input_name = session.get_inputs()[0].name
output = session.run(None, {input_name: input_data})[0]
```
### 2. Docker Containerization
```dockerfile
# Dockerfile
FROM python:3.10-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
```
```yaml
# docker-compose.yml
version: '3.8'
services:
api:
build: .
ports:
- "8000:8000"
environment:
- MODEL_PATH=/models/model.onnx
volumes:
- ./models:/models:ro
restart: unless-stopped
```
### 3. Monitoring with Prometheus
```python
from prometheus_client import Counter, Histogram, start_http_server
# Define metrics
REQUESTS = Counter('model_requests_total', 'Total requests', ['status'])
LATENCY = Histogram('model_latency_seconds', 'Latency in seconds')
@app.post("/predict")
async def predict(request: PredictRequest):
import time
start = time.time()
try:
prediction = model.predict(request.features)
REQUESTS.labels(status='success').inc()
return {"prediction": prediction}
except Exception as e:
REQUESTS.labels(status='error').inc()
raise
finally:
LATENCY.observe(time.time() - start)
```
### 4. Model Versioning
```python
import mlflow
# Log model
with mlflow.start_run():
mlflow.log_params({"n_estimators": 100, "max_depth": 10})
mlflow.log_metrics({"accuracy": 0.95, "f1": 0.93})
mlflow.sklearn.log_model(model, "model")
# Load model
model_uri = "runs:/abc123/model"
model = mlflow.sklearn.load_model(model_uri)
```
### 5. A/B Testing
```python
import random
class ABTest:
def __init__(self, variants: dict[str, float]):
self.variants = variants # {"A": 0.5, "B": 0.5}
self.results = {v: {"count": 0, "success": 0} for v in variants}
def get_variant(self, user_id: str) -> str:
random.seed(hash(user_id))
r = random.random()
cumulative = 0
for variant, weight in self.variants.items():
cumulative += weight
if r <= cumulative:
return variant
return list(self.variants.keys())[-1]
def record(self, variant: str, success: bool):
self.results[variant]["count"] += 1
if success:
self.results[variant]["success"] += 1
```
## Best Practices
### DO
- Version your models
- Implement health checks
- Use async logging
- Set up monitoring day one
- Use canary deployments
### DON'T
- Don't deploy without validation
- Don't skip latency testing
- Don't ignore drift
- Don't hard-code configs
## Exercises
### Exercise 1: FastAPI Service
```python
# TODO: Create a FastAPI service that:
# 1. Loads a model on startup
# 2. Has /predict and /health endpoints
# 3. Validates input with Pydantic
```
### Exercise 2: Docker Deployment
```python
# TODO: Containerize your ML service
# Create Dockerfile and docker-compose.yml
```
## Unit Test Template
```python
import pytest
from fastapi.testclient import TestClient
def test_health_endpoint():
"""Test health check."""
client = TestClient(app)
response = client.get("/health")
assert response.status_code == 200
assert response.json()["status"] == "healthy"
def test_predict_endpoint():
"""Test prediction."""
client = TestClient(app)
response = client.post("/predict", json={"features": [1.0, 2.0, 3.0]})
assert response.status_code == 200
assert "prediction" in response.json()
```
## Troubleshooting
| Problem | Cause | Solution |
|---------|-------|----------|
| High latency | Model too large | Quantize, use ONNX |
| Memory leaks | Poor cleanup | Implement proper lifecycle |
| API errors | Input validation | Add Pydantic schemas |
| Scaling issues | Blocking I/O | Use async, add workers |
## Related Resources
- **Agent**: `07-model-deployment`
- **Previous**: `computer-vision`
- **Docs**: [FastAPI](https://fastapi.tiangolo.com/)
---
**Version**: 1.4.0 | **Status**: Production Ready
Related in Cloud & DevOps
appbuilder-action-scaffolder
IncludedCreate, implement, deploy, and debug Adobe Runtime actions with consistent layout, validation, and error handling. Use this skill whenever the user needs to add actions to an App Builder project, understand action structure (params, response format, web/raw actions), configure actions in the manifest, use App Builder SDKs (State, Files, Events, database), deploy and invoke actions via CLI, debug action issues, or implement patterns such as webhook receivers, custom event providers, journaling consumers, large payload redirects, action sequence pipelines, and Asset Compute workers. Also trigger when users mention serverless functions in Adobe context, action logging, IMS authentication for actions, or cron-style scheduled actions.
orchestrating-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. Use this skill when the user needs a multi-step Data Cloud pipeline, cross-phase troubleshooting, or data space and data kit management. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase sf data360 workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching phase-specific skill), the task is STDM/session tracing/parquet telemetry (use observing-agentforce), standard CRM SOQL (use querying-soql), or Apex implementation (use generating-apex).
github-project-automation
IncludedAutomate GitHub repository setup with CI/CD workflows, issue templates, Dependabot, and CodeQL security scanning. Includes 12 production-tested workflows and prevents 18 errors: YAML syntax, action pinning, and configuration. Use when: setting up GitHub Actions CI/CD, creating issue/PR templates, enabling Dependabot or CodeQL scanning, deploying to Cloudflare Workers, implementing matrix testing, or troubleshooting YAML indentation, action version pinning, secrets syntax, runner versions, or CodeQL configuration. Keywords: github actions, github workflow, ci/cd, issue templates, pull request templates, dependabot, codeql, security scanning, yaml syntax, github automation, repository setup, workflow templates, github actions matrix, secrets management, branch protection, codeowners, github projects, continuous integration, continuous deployment, workflow syntax error, action version pinning, runner version, github context, yaml indentation error
sf-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase `sf data360` workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching sf-datacloud-* skill), the task is STDM/session tracing/parquet telemetry (use sf-ai-agentforce-observability), standard CRM SOQL (use sf-soql), or Apex implementation (use sf-apex).
fabric-cli
IncludedUse this skill for Fabric.so CLI workflows with the `fabric` terminal command: diagnose/install/login, search or browse a Fabric library, save notes/links/files, create folders, ask the Fabric AI assistant, manage tasks/workspaces, generate shell completion, check subscription usage, produce JSON output, and use Fabric as persistent agent memory. Do not use for Microsoft Fabric/Azure/Power BI `fab`, Daniel Miessler's Fabric framework, Python Fabric SSH, Fabric.js, or textile/fashion fabric.
lark
IncludedLark/Feishu CLI skills: lark-cli operations for docs, markdown, sheets, base, calendar, im, mail, task, okr, drive, wiki, slides, whiteboard, apps, approval, attendance, contact, vc, minutes, event. Use when the user needs to operate Lark/Feishu resources via lark-cli, send messages, manage documents, spreadsheets, calendars, tasks, OKRs, deploy web pages, or any Feishu/Lark workspace operations.