monitoring-guidelines
Monitoring guidelines for applications and infrastructure including metrics collection, alerting strategies, and SLO-based monitoring
What this skill does
# Monitoring Guidelines Apply these monitoring principles to ensure system reliability, performance visibility, and proactive issue detection. ## Core Monitoring Principles - Monitor the four golden signals: latency, traffic, errors, and saturation - Implement monitoring as code for reproducibility - Design monitoring around user experience and business impact - Use SLOs (Service Level Objectives) to guide alerting decisions - Balance comprehensive coverage with actionable insights ## Key Metrics to Monitor ### Application Metrics - Request rate (requests per second) - Error rate (percentage of failed requests) - Response time (p50, p90, p95, p99 latencies) - Active connections and concurrent users - Queue depths and processing times ### Infrastructure Metrics - CPU utilization and load average - Memory usage and available memory - Disk I/O and available storage - Network throughput and error rates - Container and pod health (for Kubernetes) ### Business Metrics - Transaction volumes and values - User signups and conversions - Feature usage and adoption rates - Revenue-impacting events - Customer satisfaction indicators ## Alerting Strategy ### Alert Design Principles - Alert on symptoms, not causes - Make alerts actionable with clear remediation steps - Set appropriate severity levels (critical, warning, info) - Avoid alert fatigue through proper threshold tuning - Include runbook links in alert notifications ### SLO-Based Alerting - Define SLOs for critical user journeys - Calculate error budgets and burn rates - Alert when error budget consumption is high - Use multi-window, multi-burn-rate alerts - Review and adjust SLOs quarterly ### Alert Configuration - Set meaningful thresholds based on baseline data - Use hysteresis to prevent flapping alerts - Implement alert dependencies to reduce noise - Route alerts to appropriate teams - Configure escalation policies ## Dashboard Design ### Effective Dashboards - Create overview dashboards for service health - Build detailed dashboards for debugging - Use consistent layouts and naming conventions - Include time range selectors and drill-down capabilities - Display SLO status prominently ### Dashboard Content - Show current state and recent trends - Include comparison to baseline or previous periods - Display deployment markers for correlation - Add annotations for significant events - Include links to related dashboards and logs ## Monitoring Tools Integration ### Data Collection - Use agents or sidecars for metric collection - Implement service discovery for dynamic environments - Configure appropriate scrape intervals - Use push vs pull based on use case - Ensure metric cardinality is manageable ### Data Storage and Retention - Set retention periods based on use case - Implement downsampling for long-term storage - Use appropriate storage backends for scale - Plan for disaster recovery of monitoring data - Monitor your monitoring infrastructure ## Health Checks and Probes - Implement liveness probes for crash detection - Use readiness probes for traffic management - Create deep health checks that verify dependencies - Expose health endpoints in a standard format - Monitor health check latency as a metric ## Incident Response - Use monitoring data to detect incidents early - Correlate metrics, logs, and traces during investigation - Document findings and update monitoring post-incident - Track MTTR (Mean Time to Recovery) metrics - Conduct regular monitoring reviews and improvements ## Capacity Planning - Track resource utilization trends - Set alerts for approaching capacity limits - Use forecasting for proactive scaling - Document capacity requirements and headroom - Review capacity quarterly
Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.