google-cloud-waf-operational-excellence
Generates operations-focused guidance for Google Cloud workloads based on the design principles and recommendations in the Operational Excellence pillar of the Google Cloud Well-Architected Framework (WAF). Use this skill to evaluate a workload, identify operational requirements, and provide actionable recommendations for deployment, monitoring, and incident management.
What this skill does
# Google Cloud Well-Architected Framework skill for the Operational Excellence pillar
## Overview
The operational excellence pillar in the Google Cloud Well-Architected Framework
provides recommendations to operate workloads efficiently on Google Cloud.
Operational excellence in the cloud involves designing, implementing, and
managing cloud solutions that provide value, performance, security, and
reliability. The recommendations in this pillar help you to continuously improve
and adapt workloads to meet the dynamic and ever-evolving needs in the cloud.
## Core principles
The recommendations in the operational excellence pillar of the Well-Architected
Framework are aligned with the following core principles:
- **Ensure operational readiness**: Define and measure criteria for a workload
to be considered ready for production, including staffing, processes, and
governance. Grounding document:
https://docs.cloud.google.com/architecture/framework/operational-excellence/operational-readiness-and-performance-using-cloudops
- **Manage incidents and problems**: Establish structured processes for
incident response, communication, and root cause analysis to minimize impact
and prevent recurrence. Grounding document:
https://docs.cloud.google.com/architecture/framework/operational-excellence/manage-incidents-and-problems
- **Manage and optimize cloud resources**: Monitor resource utilization and
right-size environments to maintain performance while ensuring operational
efficiency. Grounding document:
https://docs.cloud.google.com/architecture/framework/operational-excellence/manage-and-optimize-cloud-resources
- **Automate and manage change**: Use Infrastructure as Code (IaC) and CI/CD
pipelines to ensure consistent, repeatable, and low-risk deployments and
configuration changes. Grounding document:
https://docs.cloud.google.com/architecture/framework/operational-excellence/automate-and-manage-change
- **Continuously improve and innovate**: Regularly review architectures,
monitor industry trends, and adapt operations to meet evolving business
needs. Grounding document:
https://docs.cloud.google.com/architecture/framework/operational-excellence/continuously-improve-and-innovate
## Relevant Google Cloud products
The following are _examples_ of Google Cloud products and features that are
relevant to operational excellence:
- **Observability and monitoring**
- **Cloud Monitoring**: Full-stack observability for Google Cloud and
hybrid environments.
- **Cloud Logging**: Real-time log management and analysis at scale.
- **Error Reporting**: Aggregates and displays errors for running cloud
services.
- **Service Monitoring**: Tools for defining and tracking Service Level
Objectives (SLOs).
- **Automation and CI/CD**
- **Cloud Build**: Serverless platform for building, testing, and
deploying software.
- **Cloud Deploy**: Managed continuous delivery service for GKE, Cloud
Run, and GCE.
- **Terraform / Infrastructure Manager**: Managed service for
Infrastructure as Code (IaC) automation.
- **Artifact Registry**: Central repository for managing build artifacts
and container images.
- **Resource management and optimization**
- **Recommender (Active Assist)**: Automatically identifies idle resources
and right-sizing opportunities.
- **Resource Manager**: Hierarchical management of resources across
organizations, folders, and projects.
- **Incident response**
- **Incident response & management (IRM)**: Structured tools and processes
for managing operational disruptions.
## Workload assessment questions
Ask appropriate questions to understand operations-related requirements and
constraints of the workload and the user's organization. Choose questions from
the following list:
- **Operational readiness and performance**
- How do you define and measure operational readiness for your cloud
workloads and what specific criteria or metrics do you use?
- Describe your process for defining, tracking, and achieving SLOs for
your critical workloads.
- **Incident and problem management**
- Describe your incident management process, including roles,
responsibilities, and communication channels.
- How do you conduct post-incident reviews (PIRs) to identify root causes
and implement preventive measures?
- **Resource management and optimization**
- How do you ensure that your cloud resources are right-sized for your
workloads, and what tools or techniques do you use?
- **Change automation**
- Describe your change management process, including approval workflows,
testing procedures, and deployment strategies.
- How do you automate deployments, ensure their consistency and manage
configuration?
- **Continuous improvement**
- How do you ensure that your cloud operations are continuously adapting
to meet evolving business needs and technological advancements?
## Validation checklist
Use the following checklist to evaluate the architecture's alignment with
operational excellence recommendations:
- **Operational readiness**
- [ ] A formal framework or set of criteria exists to assess operational
readiness before production deployment.
- [ ] Service Level Objectives (SLOs) are explicitly defined and monitored
using automated tools.
- **Incident management**
- [ ] Incident response roles and communication channels are clearly
defined and documented.
- [ ] A structured, blameless post-mortem process is followed for all
major incidents.
- **Change automation**
- [ ] All infrastructure changes are performed using Infrastructure as
Code (IaC) to ensure consistency.
- [ ] CI/CD pipelines are integrated with automated testing for all
deployment changes.
- **Resource optimization**
- [ ] Resource utilization is regularly reviewed using recommendations
from Active Assist or performance data.
- **Culture of improvement**
- [ ] A documented strategy is in place for regularly reviewing and
adapting cloud operations to industry advancements.
Related in Design
contribute
IncludedLocal-only OSS contribution command center. Auto-refreshes the user's in-flight PR and issue state on invoke so conversations start with full context — no need to brief Claude on what's in flight. Helps the user find issues to contribute to on GitHub, builds per-repo dossiers of what each upstream expects (CLA, DCO, branch convention, AI policy, draft-first, review bots, issue templates), runs deterministic gates before any external action so AI-assisted contributions don't reach maintainers as slop. State is markdown-only: candidate files at ~/.contribute-system/candidates/, repo dossiers at ~/.contribute-system/research/, append-only event log at ~/.contribute-system/log.jsonl. No database, no cloud calls. Use when the user asks about their PRs / issues / contributions, wants to find new work to take on, claim an issue, build/refresh a repo's dossier, or draft a Design Issue or PR. Trigger with "/contribute", "what's my PR status", "find a contribution", "claim issue X", "draft a Design Issue for Y", "refresh dossier for Z".
architectural-analysis
IncludedUser-triggered deep architectural analysis of a codebase or scoped subtree across eight modes — information architecture, data flow, integration points, UI surfaces, interaction patterns, data model, control flow, and failure modes. This skill should be used when the user asks to "diagram this codebase," "map the architecture," "show the data flow," "give me an ERD," "trace control flow," "find the integration points," "verify the layout pattern," "audit the UX architecture," or any similar request whose primary deliverable is mermaid diagrams plus cited reports under docs/architecture/. Dispatches haiku/sonnet sub-agents in parallel for per-mode exploration, then verifies every citation mechanically before any node lands in a diagram. Not for one-off prose explanations of code (use code-explanation) or for high-level system design from scratch (use system-design).
mcp
IncludedModel Context Protocol (MCP) server development and tool management. Languages: Python, TypeScript. Capabilities: build MCP servers, integrate external APIs, discover/execute MCP tools, manage multi-server configs, design agent-centric tools. Actions: create, build, integrate, discover, execute, configure MCP servers/tools. Keywords: MCP, Model Context Protocol, MCP server, MCP tool, stdio transport, SSE transport, tool discovery, resource provider, prompt template, external API integration, Gemini CLI MCP, Claude MCP, agent tools, tool execution, server config. Use when: building MCP servers, integrating external APIs as MCP tools, discovering available MCP tools, executing MCP capabilities, configuring multi-server setups, designing tools for AI agents.
react-native-skia
IncludedDesign, build, debug, and optimise high-polish animated graphics in React Native or Expo using @shopify/react-native-skia, Reanimated, and Gesture Handler. Use when the user wants canvas-driven UI, shaders, paths, rich text, image filters, sprite fields, Skottie, video frames, snapshots, web CanvasKit setup, or performance tuning for custom motion-heavy elements such as loaders, hero art, cards, charts, progress indicators, particle systems, or gesture-driven surfaces. Also use when the user asks for fluid, glow, glass, blob, parallax, 60fps/120fps, or GPU-friendly animated effects in React Native, even if they do not explicitly say "Skia". Do not use for ordinary form/layout work with standard views.
plaid
IncludedProduct Led AI Development — guides founders from idea to launched product. Six capabilities: Idea (discover a product idea), Validate (pressure-test the idea against fatal flaws, problem reality, competition, and 2-week MVP feasibility), Plan (vision intake + document generation), Design (translate image references into a design.md spec), Launch (go-to-market strategy), and Build (roadmap execution). Use when someone says "PLAID", "plaid idea", "help me find an idea", "product idea", "idea from my business", "idea from my expertise", "plaid validate", "validate my idea", "pressure-test", "is this idea good", "find fatal flaws", "validate the problem", "plan a product", "define my vision", "generate a PRD", "product strategy", "plaid design", "design from image", "translate image to design", "create design.md", "extract design tokens", "plaid launch", "go-to-market", "launch plan", "GTM strategy", "launch playbook", "plaid build", "build the app", "start building", or "execute the roadmap".
nextjs-framer-motion-animations
IncludedAdds production-safe Motion for React or Framer Motion animations to Next.js apps, including reveal, hover and tap micro-interactions, whileInView, stagger, AnimatePresence, layout and layoutId transitions, reorder, scroll-linked UI, and lightweight route-content transitions. Use when the user asks to add, refactor, or debug Motion or Framer Motion in App Router or Pages Router codebases, especially around server/client boundaries, reduced motion, LazyMotion, bundle size, hydration, or route transitions. Avoid for GSAP-style timelines, WebGL or 3D scenes, heavy scroll storytelling, or CSS-only effects unless Motion is explicitly requested.