azure-cloud-architect
Design, review, and validate Azure cloud architectures. Use when picking the right Azure compute (AKS / App Service / Container Apps / Functions / VMs / VMSS), data store (SQL DB / Cosmos / Postgres Flexible / Storage), networking (VNet / Private Endpoint / Front Door / Application Gateway), identity (Entra ID / Managed Identity / RBAC scopes), or applying the Azure Well-Architected Framework (Reliability, Security, Cost Optimization, Operational Excellence, Performance Efficiency) to a workload. Pairs with our existing senior-cloud-architect (multi-cloud, abstract patterns) by going deep on Azure-specific services, pricing, and operational defaults.
What this skill does
# Azure Cloud Architect
End-to-end Azure-specific architecture: service selection, Well-Architected Framework assessment, identity and networking patterns, cost optimization, and operational defaults. Provider-specific complement to our generic `senior-cloud-architect` skill — that one covers cross-cloud patterns; this one knows AKS pricing tiers, when to pick Cosmos over SQL DB, and how Front Door differs from Application Gateway.
---
## When to use this skill
| Situation | Skill applies |
|-----------|---------------|
| Designing an Azure architecture from scratch | Yes — start with the **service selection decision tree** |
| Reviewing an existing Azure architecture | Yes — run **WAF assessment** via `scripts/azure_waf_scorer.py` |
| Validating an ARM/Bicep/Terraform plan | Yes — `scripts/azure_architecture_validator.py` |
| Estimating Azure cost for a workload | Yes — `scripts/azure_cost_estimator.py` |
| Picking between AKS / App Service / Container Apps / Functions | Yes — see **compute decision tree** |
| Setting up identity / RBAC / Managed Identity correctly | Yes — see **identity reference** |
| Designing a multi-region active-active or DR posture | Yes — see **reliability reference** |
| Picking SQL DB vs Cosmos vs Postgres Flexible vs Storage | Yes — see **data store decision tree** |
| Going to production without WAF review | Don't — run the WAF scorer first |
---
## Compute decision tree
Azure offers many compute options; picking the wrong one wastes money and adds operational burden.
```
Stateless HTTP service?
├── Need full control over OS / sidecar / custom runtime?
│ └── → AKS (Kubernetes; serious operational investment)
├── Standard web app (Node / Python / .NET / Java)?
│ ├── < 5 services, low traffic, want zero infra?
│ │ └── → App Service (Linux/Windows; managed PaaS)
│ ├── Need autoscale to zero + container runtime?
│ │ └── → Container Apps (KEDA-based; sweet spot for microservices)
│ └── Need true serverless / event-driven?
│ └── → Functions (Premium for VNet/always-warm; Consumption for cheap+bursty)
├── Batch / job processing?
│ └── → Container Apps Jobs OR Batch (HPC-scale)
└── Long-running stateful processes / legacy?
└── → VMs / VMSS
Stateful service (DBs you self-manage)?
├── → Generally prefer managed: SQL DB, Cosmos, PG Flexible
├── Or VM + your own DB (rarely the right call)
ML inference?
├── Realtime, GPU?
│ └── → AKS (GPU node pools) OR Azure ML managed endpoints
└── Batch?
└── → Azure ML pipelines OR Container Apps Jobs
Static frontend?
└── → Static Web Apps (auto SSL, GitHub/Bitbucket deploy)
API gateway?
├── In-VNet, internal-only?
│ └── → Application Gateway
├── Global edge, custom routing, WAF?
│ └── → Front Door
├── API management (rate limit, dev portal, monetization)?
│ └── → API Management
```
See [references/azure-services-reference.md](references/azure-services-reference.md) for service-by-service depth: pricing tiers, SLAs, limits, when to upgrade.
---
## Data store decision tree
```
Relational?
├── Standard OLTP, low-to-medium scale?
│ └── → Azure SQL Database (Single DB; serverless tier for spiky)
├── Want PostgreSQL specifically?
│ └── → Azure Database for PostgreSQL Flexible Server
├── Want MySQL specifically?
│ └── → Azure Database for MySQL Flexible Server
├── Multi-region with high concurrency, global reads?
│ └── → Cosmos DB (PostgreSQL or NoSQL API)
Document / NoSQL?
├── Global distribution, multiple consistency models, multi-region writes?
│ └── → Cosmos DB
├── Single-region key-value at scale?
│ └── → Azure Table Storage (cheap) OR Cosmos DB (Table API)
Key-value cache?
└── → Azure Cache for Redis (Standard for HA, Premium for VNet/persistence)
Blob / object storage?
└── → Storage Account (Blob); pick Hot / Cool / Cold / Archive tier per access pattern
Time-series / metrics?
├── Operational (low cost, queryable)?
│ └── → Azure Monitor Logs (Log Analytics workspace)
├── Application time series?
│ └── → Azure Data Explorer (Kusto) — purpose-built
Search?
└── → Azure AI Search (formerly Cognitive Search) — managed Lucene-based
Vector / embeddings?
├── Same workload as existing Cosmos / SQL DB?
│ └── → Cosmos DB (vector search support) OR Azure SQL DB (vector type)
└── Dedicated vector search?
└── → Azure AI Search (vector indexes)
Data warehouse?
├── < 10TB, occasional queries?
│ └── → Azure SQL Database (DTU/vCore high tier) OR Synapse Serverless
└── > 10TB, analytical workloads?
└── → Azure Synapse Analytics (Dedicated Pool) OR Microsoft Fabric
```
---
## Networking patterns
### Three core building blocks
| Component | What it does | When |
|-----------|--------------|------|
| **Virtual Network (VNet)** | L3 isolation; private IP space; subnets | Every non-trivial Azure deployment |
| **Private Endpoint** | Brings Azure PaaS services into your VNet via private IP | Default for production access to PaaS (Storage, SQL, etc.) |
| **Service Endpoint** | Lower-cost predecessor to Private Endpoint; restricts access to your VNet | Cost-sensitive, lower-security workloads only |
### Gateway services
| Service | Use when |
|---------|----------|
| **Application Gateway (v2)** | L7 LB inside VNet; WAF; OWASP rules; private + public modes; not global |
| **Front Door** | Global L7; multi-region routing; WAF at edge; CDN; URL-based routing |
| **Azure Firewall** | L3-L4-L7 stateful; egress filtering with FQDN allowlists |
| **NAT Gateway** | Predictable outbound IP for VNet egress; replaces SNAT exhaustion concerns |
| **VPN Gateway / ExpressRoute** | On-prem connectivity (S2S VPN cheaper; ER private + faster) |
### Common networking patterns
| Pattern | What | When |
|---------|------|------|
| **Hub-and-spoke** | Central hub VNet with shared services (firewall, ER, AD DS), peered with workload-specific spoke VNets | Multi-team / multi-workload orgs |
| **VWAN** | Microsoft-managed hub network simplifying multi-region + on-prem | When hub-spoke complexity gets painful |
| **Private Link for everything** | Every PaaS access via Private Endpoint; no public endpoints | Default for production / regulated workloads |
| **Azure Front Door + AppGw** | Global edge (AFD) → regional WAF/L7 (AGW) → backend | High-traffic global apps |
---
## Identity patterns
### Entra ID, Managed Identity, RBAC
| Concept | Use |
|---------|-----|
| **Microsoft Entra ID** (formerly Azure AD) | Identity provider for users, groups, apps |
| **Service Principal** | Identity for an app; has a secret or cert |
| **Managed Identity** | Service Principal whose lifecycle is tied to an Azure resource; no secrets |
| **System-assigned Managed Identity** | One per resource; deleted when resource is deleted |
| **User-assigned Managed Identity** | Independent lifecycle; can attach to multiple resources |
| **Workload Identity (AKS)** | Federated identity for K8s pods; no secret mounting |
### Choosing identity
```
Service running in Azure that calls other Azure services?
├── Single Azure resource → System-assigned MI
├── Multiple resources sharing identity (e.g., a pool of VMSS instances) → User-assigned MI
├── K8s pod → AKS Workload Identity
└── External app (CI/CD, on-prem) → Service Principal with cert (NOT secret)
Service calling external API (non-Azure)?
└── Service Principal OR app secret in Key Vault, accessed via MI
User-facing auth?
└── Entra ID with OIDC; B2C if customer-facing identity
```
### RBAC scopes (least privilege)
Assign RBAC at the smallest necessary scope:
| Scope | When |
|-------|------|
| Resource | Most specific; preferred default |
| Resource group | When the role applies to a logical group |
| Subscription | Only for subscription-wide admins |
| Management group | Cross-subscription enterprise governance |
Use built-in roles when they exist (Reader, Contributor, Storage Blob Data Reader, etc.). Custom roles only when truly needed.
---
## Azure Well-Architected Framework (WAF)
Microsoft's WAF has five pRelated in Design
contribute
IncludedLocal-only OSS contribution command center. Auto-refreshes the user's in-flight PR and issue state on invoke so conversations start with full context — no need to brief Claude on what's in flight. Helps the user find issues to contribute to on GitHub, builds per-repo dossiers of what each upstream expects (CLA, DCO, branch convention, AI policy, draft-first, review bots, issue templates), runs deterministic gates before any external action so AI-assisted contributions don't reach maintainers as slop. State is markdown-only: candidate files at ~/.contribute-system/candidates/, repo dossiers at ~/.contribute-system/research/, append-only event log at ~/.contribute-system/log.jsonl. No database, no cloud calls. Use when the user asks about their PRs / issues / contributions, wants to find new work to take on, claim an issue, build/refresh a repo's dossier, or draft a Design Issue or PR. Trigger with "/contribute", "what's my PR status", "find a contribution", "claim issue X", "draft a Design Issue for Y", "refresh dossier for Z".
architectural-analysis
IncludedUser-triggered deep architectural analysis of a codebase or scoped subtree across eight modes — information architecture, data flow, integration points, UI surfaces, interaction patterns, data model, control flow, and failure modes. This skill should be used when the user asks to "diagram this codebase," "map the architecture," "show the data flow," "give me an ERD," "trace control flow," "find the integration points," "verify the layout pattern," "audit the UX architecture," or any similar request whose primary deliverable is mermaid diagrams plus cited reports under docs/architecture/. Dispatches haiku/sonnet sub-agents in parallel for per-mode exploration, then verifies every citation mechanically before any node lands in a diagram. Not for one-off prose explanations of code (use code-explanation) or for high-level system design from scratch (use system-design).
mcp
IncludedModel Context Protocol (MCP) server development and tool management. Languages: Python, TypeScript. Capabilities: build MCP servers, integrate external APIs, discover/execute MCP tools, manage multi-server configs, design agent-centric tools. Actions: create, build, integrate, discover, execute, configure MCP servers/tools. Keywords: MCP, Model Context Protocol, MCP server, MCP tool, stdio transport, SSE transport, tool discovery, resource provider, prompt template, external API integration, Gemini CLI MCP, Claude MCP, agent tools, tool execution, server config. Use when: building MCP servers, integrating external APIs as MCP tools, discovering available MCP tools, executing MCP capabilities, configuring multi-server setups, designing tools for AI agents.
react-native-skia
IncludedDesign, build, debug, and optimise high-polish animated graphics in React Native or Expo using @shopify/react-native-skia, Reanimated, and Gesture Handler. Use when the user wants canvas-driven UI, shaders, paths, rich text, image filters, sprite fields, Skottie, video frames, snapshots, web CanvasKit setup, or performance tuning for custom motion-heavy elements such as loaders, hero art, cards, charts, progress indicators, particle systems, or gesture-driven surfaces. Also use when the user asks for fluid, glow, glass, blob, parallax, 60fps/120fps, or GPU-friendly animated effects in React Native, even if they do not explicitly say "Skia". Do not use for ordinary form/layout work with standard views.
plaid
IncludedProduct Led AI Development — guides founders from idea to launched product. Six capabilities: Idea (discover a product idea), Validate (pressure-test the idea against fatal flaws, problem reality, competition, and 2-week MVP feasibility), Plan (vision intake + document generation), Design (translate image references into a design.md spec), Launch (go-to-market strategy), and Build (roadmap execution). Use when someone says "PLAID", "plaid idea", "help me find an idea", "product idea", "idea from my business", "idea from my expertise", "plaid validate", "validate my idea", "pressure-test", "is this idea good", "find fatal flaws", "validate the problem", "plan a product", "define my vision", "generate a PRD", "product strategy", "plaid design", "design from image", "translate image to design", "create design.md", "extract design tokens", "plaid launch", "go-to-market", "launch plan", "GTM strategy", "launch playbook", "plaid build", "build the app", "start building", or "execute the roadmap".
nextjs-framer-motion-animations
IncludedAdds production-safe Motion for React or Framer Motion animations to Next.js apps, including reveal, hover and tap micro-interactions, whileInView, stagger, AnimatePresence, layout and layoutId transitions, reorder, scroll-linked UI, and lightweight route-content transitions. Use when the user asks to add, refactor, or debug Motion or Framer Motion in App Router or Pages Router codebases, especially around server/client boundaries, reduced motion, LazyMotion, bundle size, hydration, or route transitions. Avoid for GSAP-style timelines, WebGL or 3D scenes, heavy scroll storytelling, or CSS-only effects unless Motion is explicitly requested.