Skip to main content

PerfX

Universal performance engineering for Python.

PerfX is a measurement-first performance platform for Python applications and AI/ML workloads. It provides timing, CPU and memory telemetry, optional GPU/TPU/NPU backends, statistically valid benchmarking, empirical complexity analysis, evidence-based bottleneck classification, and regression detection — all behind a single, hardware-agnostic API.

PerfX reports measurements, not claims. Every metric it cannot verify is reported as unavailable rather than estimated or fabricated.


Table of Contents


Design Principles

Principle Description
Measurement over inference Metrics are only reported when they are directly observable.
No fabricated hardware data Unsupported metrics return unavailable, never a guessed value.
Dependency-light core The core package has zero mandatory third-party dependencies.
Scientific honesty Complexity results are labeled empirical estimates, not proofs.
Exception safety Instrumentation never alters or swallows application exceptions.
Extensibility New hardware vendors and frameworks integrate via a stable plugin protocol.

Requirements

  • Python 3.10, 3.11, or 3.12
  • No mandatory third-party packages

Optional extras enable additional telemetry:

Extra Enables
psutil Process CPU percentage, RSS, context switches, core counts
cuda NVIDIA NVML device discovery, utilization, power
pytorch CUDA event kernel timing, VRAM, transfer profiling, MPS discovery
tensorflow TensorFlow graph execution integration
jax TPU device discovery, XLA compilation/execution separation
numpy NumPy array input-size resolution
pandas DataFrame / Series input-size resolution
dev pytest, coverage, ruff, mypy, build
full psutil, pynvml, torch, numpy, pandas

Installation

Install the core package:

pip install perfx

Install with optional extras:

pip install perfx[psutil]
pip install perfx[cuda]
pip install perfx[pytorch]
pip install perfx[full]

For local development from source:

git clone https://github.com/SyntaxilitY/PerfX.git
cd PerfX
python -m venv .venv
.venv\Scripts\activate        # Windows
source .venv/bin/activate     # Linux / macOS
pip install -e .[dev]

Quick Start

from perfx import performance

@performance(cpu=True, memory=True, gpu="auto")
def process(items):
    return sorted(items)

process([3, 1, 2])

Print a console report for the most recent recorded call:

from perfx.core.results_store import get_results
from perfx.reporting.console import render_console

result = get_results("__main__.process")[-1]
print(render_console(result))

Core API

Function Instrumentation

from perfx import performance

@performance(cpu=True, memory=True, gpu="auto", accelerator="auto")
def train_step(batch):
    ...

Function metadata, signature, return values, and exceptions are preserved unchanged. Instrumentation failures never mask the original exception raised by the wrapped function.

Async Instrumentation

from perfx import aperformance

@aperformance()
async def fetch(url):
    ...

Measures actual awaited execution time, not coroutine construction time.

Code Block Instrumentation

from perfx import performance_block

with performance_block("serialization") as block:
    serialize(payload)

print(block.result.timing.wall_time_ns)

Class Instrumentation

from perfx import performance_class

@performance_class(include=["process", "transform"])
class Pipeline:
    def process(self, data): ...
    def transform(self, data): ...

print(Pipeline.performance_summary())

Benchmarking

A single execution is never treated as a valid benchmark. benchmark() performs warmup iterations, runs a fixed number of timed repetitions, and computes descriptive statistics including percentiles and outlier detection.

from perfx import benchmark

result = benchmark(sorted, args=([3, 1, 2],), warmup=5, iterations=30)

print(result.mean_ns)
print(result.p95_ns)
print(result.outliers_ns)

Empirical Complexity Analysis

Complexity analysis requires an explicit, safe workload generator. Functions marked repeatable=False are refused, preventing accidental repeated execution of non-idempotent operations such as database writes or payments.

from perfx import complexity

@complexity(workload=lambda n: list(range(n)), sizes=[100, 1000, 10000, 100000])
def sort_data(data):
    return sorted(data)

report = sort_data.analyze()

print(report.best_model)
print(report.confidence)
print(report.note)

All results are explicitly labeled as an Empirical Complexity Estimate, not a formal algorithmic proof.


Bottleneck Classification

from perfx import classify_bottleneck, recommend
from perfx.core.results_store import get_results

result = get_results("train_step")[-1]
bottleneck = classify_bottleneck(result)
recommendations = recommend(result, bottleneck)

print(bottleneck.classification)
print(bottleneck.evidence)
for r in recommendations:
    print(r)

Classification is derived from multiple measured signals and is never inferred from a single metric in isolation.


Regression Detection

perfx baseline mypackage.train_step --output baseline.json
perfx baseline mypackage.train_step --output current.json
perfx compare baseline.json current.json

Exit codes:

Code Meaning
0 No regression
1 Performance regression detected
2 Configuration error
3 Execution error

Command Line Interface

perfx devices                          # Discover CPU / GPU / TPU / NPU hardware
perfx benchmark module:function        # Run a statistical benchmark
perfx baseline module.function         # Record a performance baseline
perfx compare baseline.json current.json
perfx report module.function           # Print recorded results as JSON
perfx overhead module:function         # Measure PerfX's own instrumentation cost
perfx --version
perfx --help

Configuration

PerfX reads configuration from pyproject.toml:

[tool.perfx]
cpu = true
memory = true
gpu = "auto"
accelerator = "auto"

[tool.perfx.benchmark]
warmup = 5
iterations = 30

[tool.perfx.regression]
runtime_threshold = 10
memory_threshold = 15
gpu_threshold = 10

Reporting

Format Module
Console perfx.reporting.console
JSON perfx.reporting.json_reporter
Markdown perfx.reporting.markdown
HTML perfx.reporting.html

Example:

from perfx.reporting.html import render_performance_html
from perfx.core.results_store import get_results

result = get_results("train_step")[-1]
html = render_performance_html(result)

with open("report.html", "w", encoding="utf-8") as f:
    f.write(html)

Architecture

Application
    |
Public API (performance, benchmark, complexity)
    |
Instrumentation Layer
    |
Measurement Engine
    |
Hardware Abstraction Layer
    |-- CPU Backend
    |-- Memory Backend
    |-- NVIDIA CUDA Backend
    |-- AMD ROCm Backend (plugin extension point)
    |-- Apple Metal / MPS Backend
    |-- Google TPU Backend
    |-- NPU Backend (plugin extension point)
    |
Analysis Engine
    |-- Statistics
    |-- Complexity Analysis
    |-- Bottleneck Classification
    |-- Regression Detection
    |
Reporting Engine
    |-- Console / JSON / Markdown / HTML
    |
CLI / Pytest / CI Integrations

The core engine has no direct dependency on any accelerator SDK or ML framework. All hardware-specific and framework-specific behavior is implemented behind the AcceleratorBackend protocol and discovered through the perfx.plugins entry-point group.


Testing

pip install -e .[dev]
pytest -v
pytest -m "not slow"
pytest --cov=perfx --cov-report=term-missing

Browser-based reports:

pip install pytest-html
pytest --html=report.html --self-contained-html
pytest --cov=perfx --cov-report=html

Limitations

Measurements are affected by CPU frequency scaling, thermal throttling, OS scheduling, cache and branch prediction behavior, garbage collection, GPU driver and allocator behavior, and general system load. Empirical complexity results are statistical inferences from measured samples, not formal mathematical proofs. Results are only meaningfully comparable across runs captured on identical or explicitly documented hardware and software environments.

Metrics that cannot be verified through an available hardware or framework API are never estimated. They are reported as unavailable.


License

Apache License 2.0


Author

Tariq Mehmood

Release files for perfx 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for perfx 1.0.1
File Size Uploaded
perfx-1.0.1.tar.gz 111.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for perfx 1.0.1
File Interpreter ABI Platform
perfx-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 273.7 kB

Release files / perfx-1.0.1.tar.gz

Download URL perfx-1.0.1.tar.gz
Size 111.6 kB
Tags Source
SHA-256 checksum
How to use checksums
072120483feae6e9ba31df51a093abd40871c4f72ed007d739134b04c7797e47
BLAKE2b-256 checksum
How to use checksums
e3762ac59e38460f1975f62f3c2c40cff5bffe83f93c33a0871857aa07523f95
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release files / perfx-1.0.1-py3-none-any.whl

Download URL perfx-1.0.1-py3-none-any.whl
Size 162.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c9d59d0e4e09717d1c851e4a873e104f842a2d740992254ead1c2146936d26a2
BLAKE2b-256 checksum
How to use checksums
6f101e45aecca79dd64f2f687295b01153a55066942e0dd0b384a14499688323
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page