promotion-eval-mission-control
Evaluates a Mission Control environment's platform health for release or promotion readiness. Checks health check pipelines, config scrapers, background jobs, notifications, event queues, and MC infrastructure. Use for pre-release checks, environment promotion, or environment status. Triggers: "check environment health", "is it ready for release", "pre-release health check", "evaluate environment", "promotion readiness", "environment status"
What this skill does
# Mission Control Promotion Evaluation Skill
## Core Purpose
Systematically evaluate the health of a Mission Control environment **as a platform** (not demo workloads) to support release and promotion readiness decisions. This skill queries the live environment using MCP tools and produces a structured diagnostic report.
## Important Distinction
This evaluation focuses on **Mission Control platform health** — the components that make MC work (canary-checker, config-db, mission-control deployments, core health checks, scrapers, jobs). It does NOT evaluate demo workloads or user-created resources unless they indicate a platform problem.
**Known expected-fail checks**: Some health checks have the label `Expected-Fail=true`. These are intentional test checks and should be excluded from failure counts and findings.
## Parameters
When invoked, check if the user specified:
- **time_window**: Lookback period (default: `24h`)
- **target**: Environment to evaluate (ask user if not specified)
## Evaluation Procedure
Execute these phases sequentially. After each phase, record component status and findings.
Initialize a running JSON result conforming to @skills/promotion-eval-mission-control/schema.json with:
```json
{
"verdict": "READY",
"evaluated_at": "<current ISO timestamp>",
"time_window": "<window>",
"target": "<target environment>",
"components": {},
"findings": [],
"recommendations": []
}
```
## Catalog Type Reference
These are the confirmed MissionControl catalog types:
- `MissionControl::ScrapeConfig` — config scrapers
- `MissionControl::Playbook` — playbook definitions
- `MissionControl::Notification` — notification rules
- `MissionControl::Job` — background jobs
- `MissionControl::Canary` — canary check definitions
- `MissionControl::Connection` — external connections
- `MissionControl::Topology` — topology definitions
---
### Phase 1: Health Check Pipelines
**Goal**: Determine if health checks are running and passing.
1. **Get failing checks directly**: `view_failing-health-checks_mission-control` with `withRows=true` and `select=["id","name","type","status","severity","last_transition_time","description"]`
2. **Get total check count**: `list_all_checks` for baseline metrics
3. **Filter out expected failures**: Exclude checks with label `Expected-Fail=true` from failure counts
4. **Drill into real failures**: For each genuinely unhealthy check (not expected-fail), call `get_check_status(id, limit=10)` to retrieve recent execution history. Classify as:
- **Transient**: Occasional failures mixed with passes
- **Persistent**: Consistently failing across recent executions
5. **Assess staleness**: From the check list, identify checks where `updated_at` is older than the time window
**Metrics to record**:
- `total_checks`: Total number of health checks
- `healthy_count`: Number currently healthy
- `unhealthy_count`: Number currently unhealthy (excluding expected-fail)
- `expected_fail_count`: Checks labeled Expected-Fail
- `persistent_failures`: Number failing consistently
- `stale_count`: Number not updated within time window
- `health_rate`: Percentage healthy (excluding expected-fail from denominator)
**Verdict logic**:
- PASS: No persistent failures, stale_count == 0, health_rate > 95%
- WARN: Some transient failures OR 1-2 stale checks OR health_rate 80-95%
- FAIL: Any persistent failures OR stale_count > 2 OR health_rate < 80%
---
### Phase 2: Config Scrapers
**Goal**: Verify config scrapers are active and producing fresh data.
1. **Find scraper configs**: `search_catalog` with `type=MissionControl::ScrapeConfig` and `select=["id","name","health","status","updated_at"]`
2. **Check freshness**: For each scraper, check `updated_at` timestamp. Flag any not updated within expected schedule (typically 1h)
3. **Check scraper errors**: `view_mission-control-system_mission-control` with `withPanels=true` — this returns scraper error counts and a list of scrapers with errors
4. **Review recent changes**: `search_catalog_changes` with `type=MissionControl::ScrapeConfig created_at>now-{window}` for config changes
**Metrics to record**:
- `total_scrapers`: Number of scraper configs found
- `active_count`: Scrapers updated within expected window
- `stale_count`: Scrapers not recently updated
- `error_count`: From system view scraper errors panel
**Verdict logic**:
- PASS: All scrapers active, no errors
- WARN: 1-2 scrapers slightly stale OR minor errors
- FAIL: Any scraper missing updates for > 2h OR significant errors
---
### Phase 3: Background Jobs & Playbooks
**Goal**: Check for failed playbook runs and job errors.
1. **Get failed job history**: `view_jobhistory_mission-control` with `withRows=true` and `select=["name","status","duration","error","timestamp"]` limit=20
2. **Get failed playbook runs**: `get_playbook_failed_runs(limit=10)` for recent failures
3. **Get recent playbook runs**: `get_playbook_recent_runs(limit=20)` to calculate success rate
4. **Drill into failures**: For any failed playbook runs, call `get_playbook_run_steps(run_id)` to understand the failure cause
5. **Check playbook catalog health**: `search_catalog` with `type=MissionControl::Playbook health=unhealthy` to find unhealthy playbook definitions
**Metrics to record**:
- `total_recent_runs`: Total playbook runs in window
- `failed_runs`: Number of failed runs
- `success_rate`: Percentage of successful runs
- `job_errors`: Count of job errors from job history view
**Verdict logic**:
- PASS: success_rate > 95%, no recurring job errors
- WARN: success_rate 80-95% OR some job errors
- FAIL: success_rate < 80% OR critical/recurring job failures
---
### Phase 4: Notification Delivery
**Goal**: Verify the notification pipeline is functioning.
1. **Get notification send history**: `view_notification-send-history_mission-control` with `withRows=true` and `select=["id","age","resource_name","resource_current_health","title","notification"]` limit=20
2. **Get notification stats from system view**: The `view_mission-control-system_mission-control` panel (already fetched in Phase 2) includes notification counts by status (SENT, SILENCED, REPEAT-INTERVAL, etc.)
3. **Find notification configs**: `search_catalog` with `type=MissionControl::Notification` and `select=["id","name","health","status"]`
4. **Check for error notifications**: For each notification config, call `get_notifications_for_resource(resource_id, status=error, since=now-{window})`
**Metrics to record**:
- `total_notification_configs`: Number of notification rules
- `sent_count`: From system view
- `silenced_count`: From system view
- `error_count`: Notifications with error status
- `delivery_rate`: sent / (sent + error) percentage
**Verdict logic**:
- PASS: No delivery errors, system view shows sends happening
- WARN: Some errors but delivery_rate > 95%
- FAIL: delivery_rate < 95% or notification system appears down
---
### Phase 5: System & Event Queue
**Goal**: Check overall system health indicators, database, and event queue.
1. **System overview**: `view_mission-control-system_mission-control` with `withPanels=true` (reuse from Phase 2 if already fetched)
- Check scraper errors, notification stats, agent resource counts
2. **Database health**: `view_mission-control-database_mission-control` with `withPanels=true`
- Check DB size, active users, DB connections
3. **Connection health**: `list_connections` to verify external integrations are configured
**Metrics to record**:
- `db_size_bytes`: Database size
- `db_connections`: Active connections
- `active_users`: User count
- `total_connections`: Number of configured connections
**Verdict logic**:
- PASS: DB healthy, connections configured, no concerning metrics
- WARN: High DB connections or large DB size growth
- FAIL: Database unreachable or critical system errors
---
### Phase 6: MC Infrastructure Health
**Goal**: Verify Mission Control's own Kubernetes resources are healthy.
1. **Get MC pods directly**:Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.