Python Performance
SkillDocs & knowledgeGuides your agent through profiling and speeding up slow Python programs using the python claude skill.
Use Python Performance in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Python Performance and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Python Performance skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Profile and optimize Python CPU, memory, I/O, concurrency, and numerical performance.
What this skill tells your AI
The instructions your AI receives, as published by caarlos0/dotfiles in skills/python-performance/SKILL.md and read by Ahel’s review.
Name the metric first: wall time, CPU time, allocation count, retained heap, peak RSS, or I/O wait. Pin Python, dependencies, input, and environment; then change one profiled cause.
Measurement
- Use
pyperffor repeatable benchmarks with calibration, worker processes, metadata, and statistical comparison. - Use
timeitonly for small fragments. It disables cyclic GC during timing unless explicitly re-enabled, which can make allocation-heavy code look unlike production. - Use
cProfilefor call counts and cumulative development profiles; use a sampling profiler such aspy-spyfor lower-overhead process observation. - Use
tracemallocfor Python-managed allocations. If RSS grows while its traces remain stable, inspect native allocations or fragmentation with Memray or an OS profiler. sys.getsizeofis shallow; it does not measure referenced objects.- Use
python -X importtimebefore changing startup imports.
Data structures and Python operations
Choose from access patterns:
| Need | Prefer |
|---|---|
| Membership or deduplication | set or dict, not repeated list scans |
| Queue operations at both ends | collections.deque, not list.pop(0) |
| Priority queue | heapq |
| Search in maintained sorted data | bisect |
| Mutable binary accumulation | bytearray, then bytes(buffer) |
| Many string fragments | collect fragments and "".join(parts) |
These choices change semantics and memory. Do not replace a list when callers need indexing, slicing, or compact iteration.
Generators avoid eager materialization but add iteration overhead and cannot be reused. Built-ins and comprehensions often move work into optimized C loops, but they are not automatically faster for every workload.
lru_cache trades CPU for retained memory and invalidation. On an instance
method, cache keys retain self; avoid it when instances must be collected.
@dataclass(slots=True) or __slots__ can reduce memory for many instances but
affects dynamic attributes, inheritance, weak references, serialization, and
framework integration.
Memory and GC
CPython uses reference counting plus cyclic GC. Distinguish:
- growing Python allocation traces;
- retained reachable objects;
- native allocations;
- allocator fragmentation;
- peak RSS;
- allocation churn that increases CPU without retaining memory.
High RSS alone is not a leak. Tune GC thresholds, call gc.freeze(), or change
allocators only after pause, allocation, or copy-on-write measurements identify
the collector or allocator as the cause. GC defaults differ by Python version
and free-threaded build.
Threads, asyncio, and processes
- Threads overlap many blocking I/O operations because those calls release the GIL. Pure-Python CPU threads do not execute bytecode in parallel under the normal GIL; native extensions may release it.
asynciois cooperative concurrency. Any blocking call or long CPU loop in a coroutine stalls the event loop. Use bounded queues when producers can outrun consumers, and preserve cancellation and shutdown.- Processes provide CPU parallelism but add startup, pickling, IPC, memory, and failure handling. Include all of those in the benchmark.
- Start methods vary by platform and Python version. Libraries should not force a global method without owning application lifecycle.
- Free-threaded CPython enables parallel Python threads but adds evolving overhead, synchronization requirements, and extension compatibility. Test the exact interpreter and dependency set.
I/O and services
- Use buffering for repeated small reads and writes.
readinto()can reuse a buffer in measured binary pipelines but adds ownership complexity. - Batch database and network operations to reduce round trips. Oversized batches increase memory, lock duration, tail latency, and retry scope.
- Avoid constructing expensive log messages when the level is disabled. Queue handlers move slow output off latency-sensitive threads but require bounded capacity, ordering, loss, and shutdown decisions.
- Preserve flush, EOF, error, retry, ordering, cancellation, and protocol behavior when optimizing I/O.
NumPy and native acceleration
- Chained NumPy operations can allocate full-size temporaries. Use
out=, in-place operations, chunking, or fused kernels only after CPU and memory profiles show the temporary matters. - Check contiguity and strides when native kernels copy or traverse arrays poorly. Normalize layout once at a boundary, not repeatedly in a loop.
- BLAS, process pools, application threads, and runtimes can each create worker
pools. Measure oversubscription before limiting them with
threadpoolctlor environment settings. - Warm Numba before benchmarking. It helps supported Python-loop work, not code already dominated by optimized NumPy kernels.
- Cython, PyO3/Rust, GPU code, and alternative runtimes add compilation, transfer, ABI, packaging, debugging, and maintenance. Batch enough work per boundary crossing to justify them and keep a tested Python path when useful.
CPython specialization
Use dis.dis(fn, adaptive=True) after warm-up as supporting evidence for a hot
loop. Do not redesign APIs to preserve one specialized opcode; specialization
rules change between versions. Re-measure after Python upgrades.
Regression guards
Use narrow allocation, output-size, startup, or memory guards when the toolchain and platform are pinned. Wall-time gates require dedicated hardware or enough margin to avoid flaking; keep shared-runner timing advisory. Never compare runs with different GC modes, profilers, hooks, or calibration.
Related skills
code-reviewchecks a completed diff. When invoked fromcode-review, do not invoke it again.code-simplifierruns after the gain is proven.change-impact-auditortraces environment, serialization, imports, logging, and concurrency changes.runtime-process-debuggingowns subprocess, pipe, lifecycle, and shutdown failures.
Correctness overrides performance.
Signals
- GitHub stars
- 220
- Forks
- 11
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
python-performance- Source
- github.com/caarlos0/dotfiles
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonhandoff
Skill · mattpocock
More in Docs & knowledgecanvas-design
Skill · anthropics
More in Docs & knowledgedoc-coauthoring
Skill · anthropics
More in Docs & knowledgepopups
Skill · coreyhaines31
More in Docs & knowledge