monitor
Use this skill when the user needs to set up production monitoring, track app health, configure error alerts, or respond to incidents. Also use when the user says 'my app went down,' 'how do I know if something breaks,' 'set up alerts,' 'is my app healthy,' or 'I found out from a user that my site was down.' Covers error tracking, uptime monitoring, performance metrics, and incident response for SaaS applications.
What this skill does
# Monitor **This skill is for production monitoring and incident response.** For debugging specific bugs, use **debug**. For pre-launch readiness checks, use **go-live**. For security-specific monitoring (auth events, API abuse), use **secure**. For analytics and user behavior tracking, use **analytics**. ### Don't Do Yet - **Don't pay for monitoring tools** until you've outgrown the free tiers. UptimeRobot + Sentry free handles most early-stage apps. - **Don't set up DataDog, New Relic, or Grafana.** These are enterprise tools. You don't need them with < 1,000 users. - **Don't build custom dashboards.** Your hosting platform (Vercel, Railway) has built-in metrics. Use those first. - **Don't monitor everything.** Three things matter at launch: is it up, are there errors, is it slow. That's it. ## Monitoring Checklist ``` Basic Monitoring: - [ ] Uptime monitoring (is site up?) - [ ] Error tracking (are errors happening?) - [ ] Performance monitoring (is it slow?) - [ ] User activity (are people using it?) - [ ] Critical alerts configured - [ ] Check dashboard daily ``` See [MONITORING-SETUP.md](MONITORING-SETUP.md) for implementation. --- ## Why Monitor? **Without monitoring:** - Users hit errors, you don't know - Site goes down, you find out from Twitter - Slow performance, users leave silently - Security issues, no alert **With monitoring:** - Errors show in dashboard immediately - Get text when site goes down - See performance degradation - Catch issues before users complain **Goal: Know about problems before users tell you.** --- ## Three Essential Monitors ### 1. Is It Up? **Uptime monitoring** - Pings your app every minute **Free tools:** - UptimeRobot (free, 50 monitors) - Pingdom (limited free tier) - Vercel/Netlify (built-in for deployed apps) **Setup:** ``` 1. Sign up for UptimeRobot 2. Add monitor for https://yourapp.com 3. Add your email for alerts 4. Get texted if site is down ``` ### 2. Are There Errors? **Error tracking** - Captures JavaScript errors and API failures **Free tools:** - Sentry (free tier: 5k errors/month) - LogRocket (limited free) - Vercel/Netlify logs (for deployed apps) **Claude Code:** ``` Add Sentry error tracking to my app: - Install @sentry/nextjs (or appropriate package) - Capture all frontend errors and API errors - Include user context (email, ID) - Configure source maps for readable stack traces - Set up Sentry.init in both client and server entry points ``` **Lovable / Replit** (paste into chat): ``` Add error tracking to my app. I want to be notified when errors happen. Use Sentry (free tier). Show me how to: 1. Create a Sentry account and project 2. Add the tracking code to my app 3. Test that errors are being captured ``` ### 3. Is It Slow? **Performance monitoring** - Tracks page load times **Free tools:** - Vercel Analytics (built-in) - Google PageSpeed Insights (free) - Cloudflare Analytics (free tier) **Setup:** - Usually automatic with hosting platform - Check dashboard weekly --- ## What to Monitor ### Critical Metrics **Must monitor:** - Site uptime (99%+) - Error rate (< 1% of requests) - API response time (< 500ms) - Page load time (< 3s) **Nice to have:** - Active users - Feature usage - Conversion rates - User paths **For MVP:** Focus on the "must monitor" only. --- ## Setting Up Alerts **Configure alerts for:** **Critical (text me immediately):** - Site is down - Error rate spike (10x normal) - Database connection lost - Payment processing failing **Important (email within hour):** - API slow (>2 seconds) - Error rate elevated (2x normal) - Disk space low (>80%) **Informational (daily digest):** - New errors discovered - Performance trending down - Traffic patterns **Tell AI:** ``` Configure monitoring alerts: - Critical: Text to [phone] - Important: Email to [email] - Send summary: Daily at 9am ``` --- ## Daily Monitoring Routine **5-minute morning check:** ``` Daily Check: 1. Open monitoring dashboard 2. Check uptime (should be 100% yesterday) 3. Check error count (any spikes?) 4. Check performance (slower than usual?) 5. Review any alerts from overnight ``` **If all green:** You're done, 5 minutes. **If red:** Investigate using debug skill. --- ## Reading Monitoring Dashboards ### Uptime Dashboard **Green:** Site responding **Red:** Site down or slow to respond **What to check:** - Uptime percentage (target: 99%+) - Response time (target: <500ms) - Recent downtime incidents ### Error Dashboard **Look for:** - Error count spikes (sudden jump) - New error types (didn't see before) - Affected users (how many hit this?) - Error frequency (happening a lot?) **Priority:** - Affecting many users → High priority - Blocking key features → High priority - Edge case error → Lower priority ### Performance Dashboard **Look for:** - Load time trending up (getting slower) - Slow endpoints (which API calls) - Slow pages (which routes) - Geographic differences (slow in specific regions) --- ## Error Investigation **When errors spike:** ``` 1. Open error tracking dashboard (Sentry) 2. Find the most frequent error 3. Read error message and stack trace 4. Note: How many users affected? 5. Note: Started when? 6. Check: Did we deploy recently? ``` **Give to AI:** ``` Error in production: [Paste error message and stack trace] Affected: [X] users in last [Y] hours Started: [timestamp] Recent deploys: [any?] Please: 1. Explain what's wrong 2. Propose hotfix 3. How to test before deploying ``` --- ## User-Reported Issues **When user reports problem:** ``` User Report Investigation: 1. Can you reproduce it? 2. Check monitoring for errors at that time 3. Check logs for that user 4. Check if others affected 5. Determine severity Then use debug skill to fix. ``` **Tell AI:** ``` User reported: [issue description] User: [email or ID] Timestamp: [when it happened] Check monitoring and logs for this user at this time. What errors or issues do you see? ``` --- ## Proactive Monitoring **Catch issues before users:** **Weekly checks:** ``` Weekly Review: - [ ] Error trends (going up or down?) - [ ] Performance trends (slower?) - [ ] New error types introduced - [ ] Uptime issues resolved - [ ] Alert noise (too many false alerts?) ``` **Monthly checks:** ``` Monthly Health: - [ ] Compare to last month - [ ] Any degradation? - [ ] Any improvements? - [ ] Monitoring gaps (what's not tracked?) ``` --- ## Free Monitoring Stack **Recommended for MVP:** **Uptime:** - UptimeRobot (free) - 50 monitors **Errors:** - Sentry (free) - 5k errors/month **Performance:** - Vercel Analytics (free on Vercel) - Cloudflare Analytics (free) **Logs:** - Platform logs (Vercel, Netlify, Railway) **Cost: $0/month until you need more.** --- ## When to Upgrade Monitoring **Upgrade when:** - Hitting free tier limits - Need more detailed analytics - Need faster alert response - Need advanced features (session replay, etc.) **Paid tiers (typically $20-50/mo):** - Sentry Pro ($26/mo) - LogRocket ($99/mo - session replay) - DataDog ($15/host/mo) **For < 1000 users:** Free tiers sufficient. --- ## Common Monitoring Mistakes | Mistake | Fix | |---------|-----| | No monitoring set up | Set up before launch | | Alert fatigue (too many alerts) | Only alert on critical issues | | Checking once a month | Check daily (5 minutes) | | Ignoring trends | Watch for degradation over time | | No alerts configured | Set up text alerts for downtime | | Monitoring but not acting | Use monitoring to find and fix issues | --- ## Interpreting Trends **Good trends:** - Errors decreasing - Performance improving - Uptime stable at 99.9%+ **Warning trends:** - Errors slowly increasing - Performance slowly degrading - Uptime dipping below 99% **Critical trends:** - Sudden error spike - Sudden performance drop - Multiple downtime incidents **Action:** Address warning trends before they become critical. --- ## Logging vs Monitoring **Logging:** - Records what happened - For debugging sp
Related in Data & Analytics
clawarr-suite
IncludedComprehensive management for self-hosted media stacks (Sonarr, Radarr, Lidarr, Readarr, Prowlarr, Bazarr, Overseerr, Plex, Tautulli, SABnzbd, Recyclarr, Unpackerr, Notifiarr, Maintainerr, Kometa, FlareSolverr). Deep library exploration, analytics, dashboard generation, content management, request handling, subtitle management, indexer control, download monitoring, quality profile sync, library cleanup automation, notification routing, collection/overlay management, and media tracker integration (Trakt, Letterboxd, Simkl).
querying-soql
IncludedSOQL query generation, optimization, and analysis with 100-point scoring. Use this skill when the user needs SOQL/SOSL authoring or optimization: natural-language-to-query generation, relationship queries, aggregates, query-plan analysis, and performance or safety improvements for Salesforce queries. TRIGGER when: user writes, optimizes, or debugs SOQL/SOSL queries, touches .soql files, or asks about relationship queries, aggregates, or query performance. DO NOT TRIGGER when: bulk data operations (use handling-sf-data), Apex DML logic (use generating-apex), or report/dashboard queries.
app-store-optimization
IncludedApp Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklists, and tracking ranking changes.
habit-flow
IncludedAI-powered atomic habit tracker with natural language logging, streak tracking, smart reminders, and coaching. Use for creating habits, logging completions naturally ("I meditated today"), viewing progress, and getting personalized coaching.
app-store-optimization
IncludedApp Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklists, and tracking ranking changes.
visualizing-data
IncludedBuilds dashboards, reports, and data-driven interfaces requiring charts, graphs, or visual analytics. Provides systematic framework for selecting appropriate visualizations based on data characteristics and analytical purpose. Includes 24+ visualization types organized by purpose (trends, comparisons, distributions, relationships, flows, hierarchies, geospatial), accessibility patterns (WCAG 2.1 AA compliance), colorblind-safe palettes, and performance optimization strategies. Use when creating visualizations, choosing chart types, displaying data graphically, or designing data interfaces.