mlops-pipelines
Model deployment strategies, monitoring and drift detection, CI/CD for ML models, feature store concepts, and model versioning
What this skill does
# MLOps Pipelines ## Model Deployment Strategies ### Batch Deployment - **Description**: Run model on fixed schedule on accumulated data - **Use Cases**: Credit scoring, churn prediction, recommendations - **Advantages**: Simple, cost-effective, handles large volumes - **Challenges**: Latency, stale predictions - **Tools**: Apache Airflow, dbt, cron jobs, cloud batch services ### Real-time Deployment - **Description**: Serve model as API for immediate predictions - **Use Cases**: Fraud detection, dynamic pricing, personalization - **Advantages**: Low latency, fresh predictions - **Challenges**: Scalability, infrastructure complexity - **Tools**: Flask, FastAPI, TensorFlow Serving, TorchServe, KServe ### Edge Deployment - **Description**: Deploy model on edge devices (IoT, mobile, embedded) - **Use Cases**: Computer vision, speech recognition, offline scenarios - **Advantages**: Low latency, privacy, no internet required - **Challenges**: Limited compute, model size constraints - **Tools**: TensorFlow Lite, ONNX, Core ML, ML Kit ### Streaming Deployment - **Description**: Process data streams with real-time predictions - **Use Cases**: Real-time analytics, monitoring, anomaly detection - **Advantages**: Continuous processing, low latency - **Challenges**: State management, exactly-once semantics - **Tools**: Apache Kafka, Apache Flink, Apache Spark Streaming ## Model Monitoring and Drift Detection ### Performance Monitoring - **Prediction Metrics**: Track model outputs and distributions - **Accuracy Metrics**: Monitor precision, recall, F1, MAE, RMSE - **Business Metrics**: Connect predictions to business KPIs - **Latency**: Track prediction response times - **Throughput**: Monitor predictions per second ### Data Drift Detection - **Covariate Drift**: Changes in input feature distribution - **Prior Probability Drift**: Changes in target class distribution - **Concept Drift**: Changes in relationship between features and target - **Detection Methods**: Statistical tests, KL divergence, PSI - **Visualization**: Feature distribution plots over time ### Drift Mitigation - **Retraining Triggers**: Automatic retraining on drift detection - **Ensemble Methods**: Combine multiple models for robustness - **Online Learning**: Update model continuously with new data - **Feature Monitoring**: Track feature distributions and correlations ### Alerting - **Threshold-based Alerts**: Alert when metrics exceed thresholds - **Anomaly Detection**: Detect unusual patterns automatically - **Dashboard Monitoring**: Real-time dashboards for visibility - **Incident Response**: Procedures for handling model failures ## CI/CD for ML Models ### ML Pipeline Stages - **Data Ingestion**: Collect and validate training data - **Feature Engineering**: Create and validate features - **Model Training**: Train and validate models - **Model Evaluation**: Evaluate model performance - **Model Deployment**: Deploy model to production - **Monitoring**: Monitor model performance and data drift ### Continuous Integration - **Code Testing**: Unit tests, integration tests - **Data Validation**: Validate data quality and schema - **Model Testing**: Test model performance and behavior - **Artifact Storage**: Store models, features, and metadata - **Automated Builds**: Build and test on every commit ### Continuous Deployment - **Automated Deployment**: Deploy models automatically after validation - **Canary Releases**: Gradual rollout to subset of users - **A/B Testing**: Compare model versions in production - **Rollback**: Quick rollback to previous version if issues occur - **Blue-Green Deployment**: Switch between production environments ### MLOps Platforms - **MLflow**: Open-source ML lifecycle platform - **Kubeflow**: Kubernetes-native ML platform - **Vertex AI**: Google Cloud ML platform - **SageMaker**: AWS ML platform - **Azure ML**: Microsoft Azure ML platform ## Feature Store Concepts ### Feature Store Benefits - **Feature Reusability**: Share features across models and teams - **Consistency**: Ensure consistent feature computation - **Latency**: Low-latency feature serving for real-time predictions - **Versioning**: Track feature versions and lineage - **Governance**: Control feature access and permissions ### Feature Types - **Batch Features**: Computed from batch data (e.g., daily aggregates) - **Streaming Features**: Computed from streaming data (e.g., real-time counts) - **On-demand Features**: Computed at request time (e.g., time since last event) - **Derived Features**: Combinations of other features ### Feature Store Architecture - **Offline Store**: Store historical features for training - **Online Store**: Low-latency serving for inference - **Feature Registry**: Catalog of available features - **Feature Monitoring**: Track feature quality and drift ### Feature Store Tools - **Feast**: Open-source feature store - **Tecton**: Enterprise feature store platform - **Hopsworks**: Open-source feature store - **AWS Feature Store**: AWS feature store service - **Azure Feature Store**: Azure feature store service ## Model Versioning and Registry ### Model Versioning - **Version Numbers**: Semantic versioning for models - **Metadata**: Track training data, hyperparameters, metrics - **Artifacts**: Store model files, weights, configurations - **Lineage**: Track model provenance and dependencies - **Tags**: Label models for easy identification ### Model Registry - **Central Repository**: Store all model versions - **Model Promotion**: Promote models through stages (dev, staging, prod) - **Access Control**: Control who can deploy models - **Model Search**: Find models by metadata or tags - **Model Documentation**: Document model purpose and behavior ### Model Artifacts - **Model Files**: Saved model weights and architecture - **Configuration Files**: Model hyperparameters and settings - **Training Code**: Code used to train the model - **Evaluation Results**: Model performance metrics - **Deployment Artifacts**: Docker images, serving configurations ### Model Lifecycle - **Development**: Initial model development and experimentation - **Staging**: Test model in staging environment - **Production**: Deploy model to production - **Retired**: Decommission model when no longer needed - **Archived**: Store model for historical reference ### Best Practices - **Reproducibility**: Ensure models can be reproduced - **Documentation**: Document model purpose, behavior, and limitations - **Testing**: Test models thoroughly before deployment - **Monitoring**: Monitor model performance in production - **Governance**: Establish approval processes for model deployment
Related in Cloud & DevOps
appbuilder-action-scaffolder
IncludedCreate, implement, deploy, and debug Adobe Runtime actions with consistent layout, validation, and error handling. Use this skill whenever the user needs to add actions to an App Builder project, understand action structure (params, response format, web/raw actions), configure actions in the manifest, use App Builder SDKs (State, Files, Events, database), deploy and invoke actions via CLI, debug action issues, or implement patterns such as webhook receivers, custom event providers, journaling consumers, large payload redirects, action sequence pipelines, and Asset Compute workers. Also trigger when users mention serverless functions in Adobe context, action logging, IMS authentication for actions, or cron-style scheduled actions.
orchestrating-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. Use this skill when the user needs a multi-step Data Cloud pipeline, cross-phase troubleshooting, or data space and data kit management. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase sf data360 workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching phase-specific skill), the task is STDM/session tracing/parquet telemetry (use observing-agentforce), standard CRM SOQL (use querying-soql), or Apex implementation (use generating-apex).
github-project-automation
IncludedAutomate GitHub repository setup with CI/CD workflows, issue templates, Dependabot, and CodeQL security scanning. Includes 12 production-tested workflows and prevents 18 errors: YAML syntax, action pinning, and configuration. Use when: setting up GitHub Actions CI/CD, creating issue/PR templates, enabling Dependabot or CodeQL scanning, deploying to Cloudflare Workers, implementing matrix testing, or troubleshooting YAML indentation, action version pinning, secrets syntax, runner versions, or CodeQL configuration. Keywords: github actions, github workflow, ci/cd, issue templates, pull request templates, dependabot, codeql, security scanning, yaml syntax, github automation, repository setup, workflow templates, github actions matrix, secrets management, branch protection, codeowners, github projects, continuous integration, continuous deployment, workflow syntax error, action version pinning, runner version, github context, yaml indentation error
sf-datacloud
IncludedSalesforce Data Cloud product orchestrator for connect→prepare→harmonize→segment→act workflows. TRIGGER when: user needs a multi-step Data Cloud pipeline, asks to set up or troubleshoot Data Cloud across phases, manages data spaces or data kits, or wants a cross-phase `sf data360` workflow. DO NOT TRIGGER when: work is isolated to a single phase (use the matching sf-datacloud-* skill), the task is STDM/session tracing/parquet telemetry (use sf-ai-agentforce-observability), standard CRM SOQL (use sf-soql), or Apex implementation (use sf-apex).
fabric-cli
IncludedUse this skill for Fabric.so CLI workflows with the `fabric` terminal command: diagnose/install/login, search or browse a Fabric library, save notes/links/files, create folders, ask the Fabric AI assistant, manage tasks/workspaces, generate shell completion, check subscription usage, produce JSON output, and use Fabric as persistent agent memory. Do not use for Microsoft Fabric/Azure/Power BI `fab`, Daniel Miessler's Fabric framework, Python Fabric SSH, Fabric.js, or textile/fashion fabric.
lark
IncludedLark/Feishu CLI skills: lark-cli operations for docs, markdown, sheets, base, calendar, im, mail, task, okr, drive, wiki, slides, whiteboard, apps, approval, attendance, contact, vc, minutes, event. Use when the user needs to operate Lark/Feishu resources via lark-cli, send messages, manage documents, spreadsheets, calendars, tasks, OKRs, deploy web pages, or any Feishu/Lark workspace operations.