python-memory-safe-scripts
Memory-safe Python script patterns for long-running processes under systemd MemoryMax constraints. Covers allocator purge (mimalloc/glibc malloc_trim), HTTP response lifecycle, DataFrame cleanup, thread-local connection reuse, and periodic GC cadence. Battle-tested through 5 OOM optimization cycles on production GPU workstations. Use this skill proactively whenever writing or reviewing Python scripts that: run under systemd with MemoryMax, process data in loops (downloads, ETL, backfill), use ThreadPoolExecutor, or make repeated HTTP requests. Also use when diagnosing OOM kills, RSS creep, or fd exhaustion in Python services. TRIGGERS - memory optimization, OOM prevention, RSS reduction, malloc_trim, systemd MemoryMax, memory leak, allocator purge, memory-safe script, RSS creep, fd exhaustion, SIGKILL status 9, MemoryHigh, glibc arena, mimalloc purge, requests memory leak, ThreadPoolExecutor cleanup.
What this skill does
# Memory-Safe Python Script Patterns
Battle-tested patterns for keeping Python scripts alive under systemd `MemoryMax` constraints. Extracted from `repair_direct_parquet.py` (24-worker parallel repair) and `exness_tick_cache_seeder.py` (10-symbol daily seeder) after 5 OOM optimization cycles on a 62 GB GPU workstation.
**Core insight**: Python's garbage collector frees objects, but the C allocator (glibc ptmalloc2) does NOT return freed pages to the OS. Without explicit `malloc_trim(0)`, RSS only grows — even after `del` and `gc.collect()`. mimalloc with `MIMALLOC_PURGE_DELAY` helps but explicit purge is faster.
> **Self-Evolving Skill**: This skill improves through use. If instructions are wrong, parameters drifted, or a workaround was needed — fix this file immediately, don't defer. Only update for real, reproducible issues.
## The 7 Patterns
### 1. Cached Allocator Purge
The most important pattern. Cache the ctypes library handle on first call so subsequent purges are a single FFI invocation with zero allocation overhead.
```python
import ctypes
import gc
import sys
_purge_lib = None
_purge_method = None # "mimalloc" | "glibc" | "none"
def _force_allocator_purge():
"""Force mimalloc/glibc to return freed pages to the OS."""
global _purge_lib, _purge_method
if sys.platform != "linux":
return
if _purge_method is None:
try:
_purge_lib = ctypes.CDLL("libmimalloc.so.2")
_purge_method = "mimalloc"
except OSError:
try:
_purge_lib = ctypes.CDLL("libc.so.6")
_purge_method = "glibc"
except OSError:
_purge_method = "none"
if _purge_method == "mimalloc":
_purge_lib.mi_collect(ctypes.c_bool(True))
elif _purge_method == "glibc":
_purge_lib.malloc_trim(0)
def _force_gc():
"""Python GC + allocator purge. Call every 50 iterations + between work units."""
gc.collect()
_force_allocator_purge()
```
**Why cached handle matters**: `ctypes.CDLL("libc.so.6")` calls `dlopen()` which itself allocates memory. Calling it 1400 times in a loop is counterproductive. Cache it once.
**Why prefer mimalloc**: When `LD_PRELOAD=libmimalloc.so.2` is active, glibc's `malloc_trim` is a no-op because mimalloc intercepted all allocations. `mi_collect(True)` is the correct purge for mimalloc.
### 2. HTTP Response Lifecycle
Close responses immediately after extracting the content you need. The `requests` library holds the response body, connection pool references, and urllib3 internal state.
```python
# CORRECT: extract content, close, delete
resp = requests.get(url, timeout=60)
if resp.status_code != 200:
resp.close()
return None
content = resp.content # Extract what you need
resp.close() # Release connection pool reference
del resp # Drop the Python object
# Process content...
del content # Release after processing
```
```python
# WRONG: response lives until end of function scope
resp = requests.get(url, timeout=60)
data = parse(resp.content) # resp still alive, holding ~18 MB
return data # resp GC'd eventually... maybe
```
**Why this matters**: Each `requests.Response` holds `content` (the full body), a reference to the `urllib3.HTTPResponse`, and the connection pool's `PoolManager`. With 4 concurrent workers processing 1400 URLs, unclosed responses accumulate hundreds of MB.
### 3. Explicit Object Deletion
Don't rely on Python's GC for large objects. Use `del` immediately after the object is no longer needed.
```python
# After writing a DataFrame to Parquet
_atomic_write_parquet(df, path)
del df # Only reference gone → immediate refcount GC
# After extracting data from a ZIP
with zipfile.ZipFile(io.BytesIO(zip_bytes)) as zf:
df = pl.read_csv(zf.open(zf.namelist()[0]), ...)
del zip_bytes # Release raw ZIP content after parsing
# After processing a list of work items
results = process_all(missing_days)
del missing_days # Release the 1400-element date list
```
**When to `del`**: any object larger than ~1 MB that you're done with. DataFrames, byte strings from HTTP responses, ZIP contents, large lists.
### 4. Periodic GC Cadence
Call `_force_gc()` at two levels:
```python
# Level 1: Every 50 iterations within a work unit
for i, item in enumerate(items):
process(item)
if (i + 1) % 50 == 0:
_force_gc()
# Level 2: Between major work units
for symbol in symbols:
seed_symbol(symbol)
_force_gc() # Release all per-symbol state before next symbol
```
**Why 50**: Empirically validated on a 32-core workstation. At 100, RSS drifts too high before purge. At 25, the purge overhead is measurable (~2% throughput loss). 50 is the sweet spot from `repair_direct_parquet.py`.
### 5. ThreadPoolExecutor Cleanup
After the executor exits, explicitly clean up residual state.
```python
with concurrent.futures.ThreadPoolExecutor(max_workers=4) as ex:
pending = {}
# ... bounded future submission pattern ...
# After pool exits:
del pending # Future objects hold references to results
del missing # Work item list
_force_gc() # Release worker thread memory + allocator pages
```
For advanced cases (DB connections in workers), close thread-local resources explicitly:
```python
pool.shutdown(wait=False, cancel_futures=True)
for t in threading.enumerate():
if t.name.startswith("ThreadPoolExecutor"):
_close_worker_cache() # Close DB connections
gc.collect()
_force_allocator_purge()
```
### 6. Thread-Local Connection Reuse
Never create database connections or HTTP sessions inside a loop. Use `threading.local()` to get one connection per worker thread.
```python
import threading
_thread_local = threading.local()
def _get_worker_cache():
"""One DB connection per worker thread, reused across all iterations."""
cache = getattr(_thread_local, "cache", None)
if cache is None:
cache = DatabaseClient()
_thread_local.cache = cache
return cache
def _close_worker_cache():
"""Explicit cleanup at shutdown."""
cache = getattr(_thread_local, "cache", None)
if cache is not None:
cache.close()
_thread_local.cache = None
```
**Why this prevents fd exhaustion**: Each `urllib3.PoolManager(maxsize=20)` holds up to 20 file descriptors. Creating a new one per iteration in a 24-worker pool exhausts `ulimit -n 1024` within minutes. Thread-local reuse keeps fd count at ~4N+50 for N workers.
### 7. systemd Service Configuration
```ini
[Service]
# Memory limits — hard kill prevents runaway RSS
MemoryHigh=2G # Soft limit: triggers reclaim pressure
MemoryMax=4G # Hard limit: SIGKILL on breach
MemorySwapMax=0 # No swap escape — fail fast, don't thrash
# mimalloc: replaces glibc ptmalloc2, returns freed pages faster
Environment=LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libmimalloc.so.2
Environment=MIMALLOC_PURGE_DELAY=1000
# OOM priority (lower = more likely to survive)
OOMScoreAdjust=-200
ManagedOOMMemoryPressure=kill
```
**MemoryHigh vs MemoryMax**: `MemoryHigh` triggers kernel memory reclaim (cgroup pressure) — the process slows but survives. `MemoryMax` is a hard SIGKILL. Set MemoryHigh at 50-66% of MemoryMax so the kernel gets a chance to reclaim before killing.
---
## Anti-Patterns
| Anti-Pattern | Why It Fails | Fix |
| ---------------------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------- |
| `ctypes.CDLL("libc.so.6")` inside a loop | `dlopen()` allocates memory; 1000 calls wastes ~50 MB | Cache the handle in a module global |
| `requests.get()` without `resp.close()` | Response body + connection pool held until GC | `resp.close()` + `del resp` immeRelated in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.