python-advanced-profiling-and-optimization
Advanced Python profiling with cProfile, py-spy, and optimization techniques including Python 3.13 features. PROACTIVELY activate for: (1) CPU profiling with cProfile, (2) Flame graph generation with py-spy, (3) Memory optimization, (4) Async performance patterns, (5) Python 3.13 JIT and free-threaded mode. Triggers: "python performance", "cProfile", "py-spy", "flame graph", "memory optimization", "async performance", "JIT", "GIL"
What this skill does
# Python Advanced Profiling and Optimization
This skill provides expert-level knowledge on profiling Python applications to identify performance bottlenecks in CPU and memory usage, and guides evaluation and application of experimental Python 3.13 features.
## CPU Profiling with cProfile
cProfile is Python's built-in deterministic profiler, measuring execution time of all functions.
### Basic Usage
```python
import cProfile
import pstats
# Profile a specific function
def main():
# Your application code
result = expensive_operation()
return result
if __name__ == '__main__':
profiler = cProfile.Profile()
profiler.enable()
main()
profiler.disable()
# Print stats sorted by cumulative time
stats = pstats.Stats(profiler)
stats.sort_stats('cumulative')
stats.print_stats(20) # Top 20 functions
```
### Command-Line Profiling
```bash
# Profile entire script
python -m cProfile -o output.prof script.py
# View results
python -m pstats output.prof
# Then in pstats shell:
# sort cumtime
# stats 20
```
### Interpreting Output
```
ncalls tottime percall cumtime percall filename:lineno(function)
10000 0.156 0.000 0.234 0.000 utils.py:45(process_data)
5000 0.089 0.000 0.089 0.000 {built-in method json.loads}
```
**Columns**:
- **ncalls**: Number of calls
- **tottime**: Total time in function (excluding subcalls)
- **percall**: tottime / ncalls
- **cumtime**: Cumulative time (including subcalls) - **Most important**
- **percall**: cumtime / ncalls
**What to look for**:
- High cumtime functions are bottlenecks
- High ncalls with moderate cumtime suggests optimization opportunity
- Built-in functions with high time may indicate data structure issues
### Context Manager for Profiling
```python
import cProfile
import pstats
from contextlib import contextmanager
@contextmanager
def profile(output_file=None):
"""Context manager for profiling code blocks"""
profiler = cProfile.Profile()
profiler.enable()
try:
yield profiler
finally:
profiler.disable()
if output_file:
profiler.dump_stats(output_file)
else:
stats = pstats.Stats(profiler)
stats.sort_stats('cumulative')
stats.print_stats(20)
# Usage
with profile('endpoint.prof'):
process_api_request(data)
```
## Flame Graph Generation with py-spy
py-spy is a sampling profiler that can attach to running processes without code modification.
### Installation
```bash
pip install py-spy
```
### Basic Usage
```bash
# Profile running process by PID
py-spy record -o profile.svg --pid 12345
# Profile script directly
py-spy record -o profile.svg -- python script.py
# Live top-like view
py-spy top --pid 12345
```
### Reading Flame Graphs
**Flame graph structure**:
- **X-axis**: Alphabetical order (NOT time)
- **Y-axis**: Stack depth (bottom = entry point, top = deepest call)
- **Width**: Percentage of total time
**What to look for**:
- **Wide blocks**: Functions consuming most time
- **Tall stacks**: Deep call chains (may indicate recursion issues)
- **Flat plateaus**: Hot paths through code
### Profiling Production Services
```bash
# Attach to running Uvicorn/Gunicorn worker
# Find PID: ps aux | grep uvicorn
py-spy record -o api-profile.svg --pid $(pgrep -f uvicorn) --duration 30
# Profile with native extensions
py-spy record --native -o profile.svg --pid 12345
```
### Comparing Before/After
```bash
# Before optimization
py-spy record -o before.svg -- python script.py
# After optimization
py-spy record -o after.svg -- python script.py
# Compare visually or use speedscope.app
```
## Memory Optimization Techniques
Reducing memory footprint improves performance and allows scaling to more users.
### Using __slots__
For classes with fixed attributes, `__slots__` significantly reduces memory usage.
```python
# BAD: Uses __dict__ (high memory overhead)
class User:
def __init__(self, name, email):
self.name = name
self.email = email
# GOOD: Uses __slots__ (lower memory)
class User:
__slots__ = ('name', 'email')
def __init__(self, name, email):
self.name = name
self.email = email
# Memory savings: ~40-50% per instance
```
**When to use**: Classes with many instances (models, data structures)
**When NOT to use**: Classes needing dynamic attributes
### Generators vs Lists
Generators process data lazily, reducing memory footprint.
```python
# BAD: Loads entire dataset into memory
def get_all_users():
users = []
for row in db.query("SELECT * FROM users"):
users.append(process_user(row))
return users
total = sum(user.score for user in get_all_users())
# GOOD: Generator yields one at a time
def get_all_users():
for row in db.query("SELECT * FROM users"):
yield process_user(row)
total = sum(user.score for user in get_all_users())
```
**Memory savings**: Proportional to dataset size (can be orders of magnitude)
### String Concatenation
```python
# BAD: Creates new string object each iteration
result = ""
for item in items:
result += str(item) + "," # O(n^2) time and memory
# GOOD: Join at the end
result = ",".join(str(item) for item in items) # O(n)
```
### Using array Module for Numeric Data
```python
# BAD: List of integers (high memory)
numbers = [1, 2, 3, 4, 5] * 100000 # ~800 KB
# GOOD: Array of integers (low memory)
import array
numbers = array.array('i', [1, 2, 3, 4, 5] * 100000) # ~400 KB
# EVEN BETTER: NumPy for numerical computation
import numpy as np
numbers = np.array([1, 2, 3, 4, 5] * 100000) # ~400 KB + vectorized operations
```
### Memory Profiling with memory_profiler
```bash
pip install memory-profiler
```
```python
from memory_profiler import profile
@profile
def load_data():
data = [i for i in range(1000000)] # Line-by-line memory usage
processed = [x * 2 for x in data]
return processed
if __name__ == '__main__':
load_data()
```
Run with:
```bash
python -m memory_profiler script.py
```
### Python 3.13 Docstring Stripping
Python 3.13 supports docstring stripping to reduce memory:
```bash
# Build Python with docstring stripping
./configure --with-pydoc=no
# Or use PYTHONOPTIMIZE=2
python -OO script.py # Strips docstrings
```
**Memory savings**: 5-10% in typical applications
## Evaluating Experimental Python 3.13 Features
Python 3.13 introduces experimental performance features that require careful evaluation.
### JIT Compiler
Python 3.13 includes an experimental JIT compiler for potential performance gains.
#### Enabling JIT
```bash
# Build Python with JIT enabled
./configure --enable-experimental-jit
make
```
Or use official Python 3.13+ with:
```bash
PYTHON_JIT=1 python script.py
```
#### When to Use JIT
**Good candidates**:
- CPU-bound workloads
- Tight loops with numeric computation
- Long-running processes
- Pure Python code (not C extensions)
**Poor candidates**:
- I/O-bound workloads
- Short-lived scripts
- Code dominated by C extensions (NumPy, etc.)
#### Benchmarking JIT
```python
import time
def cpu_intensive_task(n):
"""Pure Python computation - good JIT candidate"""
result = 0
for i in range(n):
result += i * i
return result
# Benchmark without JIT
start = time.perf_counter()
result = cpu_intensive_task(10_000_000)
print(f"Without JIT: {time.perf_counter() - start:.3f}s")
# Run again with PYTHON_JIT=1
# Expected: 5-15% improvement for CPU-bound code
```
#### Current Limitations
- **Modest gains**: 5-15% for CPU-bound code (as of Python 3.13)
- **Warm-up time**: Initial runs may be slower
- **Memory overhead**: JIT compilation uses extra memory
- **Experimental**: May have bugs, not production-ready yet
### Free-Threaded Mode (No-GIL)
Python 3.13 includes experimental support for running without the Global Interpreter Lock (GIL).
#### Enabling Free-Threaded Mode
```bash
# Build Python without GIL
./configure --disable-gil
mRelated in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.