Skip to main content

gitm-labs

image

Behavioral compiler + intervention runtime for GPU-intensive workloads. Given a workload and a time budget, gitm-labs autonomously profiles, attributes, and applies kernel-level interventions to hit a target performance improvement — and produces a provenance report showing exactly what it changed and why.

Install

pip install gitm-labs

NVIDIA GPU support:

pip install "gitm-labs[nvidia]"

Optional extras: bench (HFT/biotech/edge benchmark harness), prometheus, otlp, s3.

Requires: Python 3.10+, NVIDIA (NVML + CUPTI) or AMD (ROCm SMI + rocprof) GPU.

Quick start

export GITM_S3_ROOT="s3://your-bucket/gitm"   # durable store for datasets + run outputs
export GITM_SCRATCH="/mnt/nvme/gitm"           # local ephemeral run dir (defaults to ~/.cache/gitm)

gitm run --workload hft-lob --budget 24h --target 15%

Workloads: hft-lob (HFT order-book), af2 (AlphaFold2 protein inference), kitti (3D LiDAR detection).

--budget is the wall-clock time limit. --target is the performance improvement fraction gitm-labs commits to delivering, or issues a diagnostic explaining why the floor could not be met.

Verify your environment first:

gitm doctor

Embedded API

from gitm import optimize

result = optimize(engine, budget="24h", target=0.15)

The 24-hour loop

gitm-labs runs a five-phase autonomous loop within the allotted budget:

Phase Hours What happens
1. Profile 0–2 Capture event + state telemetry; fingerprint workload; build predicted execution graph
2. Attribute 2–6 Compute residuals against predicted graph; run causal attribution
3. Rank 6–12 Query intervention library; rank candidates via counterfactual replay
4. Apply 12–20 Apply top-N interventions with rollback gates
5. Report 20–24 Stabilize; write provenance report (claim → evidence → intervention → delta)

Architecture

gitm-labs separates the empirical half (what happened) from the predicted half (what should have happened). Everything downstream operates on residuals — the difference between the two.

Two telemetry planes

State telemetry (gitm.telemetry)

Point-in-time samples of GPU state at ~1 Hz: utilization, memory, power, clocks, temperature, throttle reasons, NVLink throughput, ECC counters.

Source: NVML (NVIDIA) / ROCm SMI (AMD). Cost: ~microseconds per sample.

Event telemetry (gitm.tracer)

Per-kernel activity records with start/end timestamps, stream IDs, and memory transfer events.

Source: CUPTI (NVIDIA) / rocprof (AMD). Required for kernel-time invariant checks.

Deviation invariants

The monitor checks observed-vs-predicted against three invariants:

  1. Kernel-time — per-kernel duration must lie within roofline bounds.
  2. Memory-traffic — per-kernel bytes-moved must match predicted.
  3. Stream-concurrency — predicted-concurrent kernels must overlap.

See docs/invariants.md.

Module responsibilities

Module Responsibility
gitm.telemetry Vendor-backend autodiscovery, NVML/ROCm SMI samples, pluggable sinks
gitm.tracer Event-telemetry capture (CUPTI/rocprof), trace schema, context manager
gitm.planner Behavioral Compiler — roofline-based predicted execution graph
gitm.optimizer.monitor Deviation monitor — residuals against 3 invariants
gitm.optimizer.attribution Granger + doubly-robust on residual subgraph
gitm.optimizer.replay Counterfactual replay for predicted intervention delta
gitm.optimizer.qualification Workload fingerprint gate (commit / diagnose)
gitm.optimizer.report Provenance chain renderer (claim → evidence → intervention → delta)
gitm.kernels Curated intervention library — 15–20 levers with applicability + safety
gitm.agents Autonomous policy — selects interventions, drives rollback
gitm.scheduler 24-hour loop phase orchestration

Data layout

Two environment variables control where data lives:

export GITM_S3_ROOT="s3://your-bucket/gitm"   # canonical store (datasets + run outputs)
export GITM_SCRATCH="/mnt/nvme/gitm"           # local ephemeral dir (defaults to ~/.cache/gitm)

Layout under $GITM_S3_ROOT:

datasets/{hft,biotech,edge}/    # benchmark inputs (immutable, sha256-pinned)
runs/                            # baseline + run outputs
traces/                          # captured event-telemetry traces
telemetry/                       # state-telemetry samples

Local scratch is ephemeral and synced to S3 after each run.

Primary interfaces

# tracer
with gitm.tracer.capture(out_path: Path) -> ContextManager[Trace]: ...

# planner
graph = gitm.planner.predict_graph(model: ModelSpec, hw: HardwareSpec, batch: BatchConfig) -> Graph

# monitor
residuals = gitm.optimizer.monitor.residuals(trace: Trace, graph: Graph) -> Residuals
violations = gitm.optimizer.monitor.check_invariants(residuals, invariants) -> list[Violation]

# attribution
hypotheses = gitm.optimizer.attribution.attribute(residuals: Residuals, graph: Graph) -> RankedHypotheses

# report
report_md = gitm.optimizer.report.write(claims: list[Claim], provenance: Provenance) -> str

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gitm_labs-0.0.4.tar.gz (306.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gitm_labs-0.0.4-py3-none-any.whl (129.5 kB view details)

Uploaded Python 3

File details

Details for the file gitm_labs-0.0.4.tar.gz.

File metadata

  • Download URL: gitm_labs-0.0.4.tar.gz
  • Upload date:
  • Size: 306.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for gitm_labs-0.0.4.tar.gz
Algorithm Hash digest
SHA256 443b292f83e4aae2e9ca0621a1633cd3abdd17a7a8e3163e359fba1e33b7d1f1
MD5 ba0e1a1a850c8b370e85d9932a575a6d
BLAKE2b-256 ee2d5edb0d7b81427e04925369e0f6437ef09b857d283fb282fb83fc5969d764

See more details on using hashes here.

Provenance

The following attestation bundles were made for gitm_labs-0.0.4.tar.gz:

Publisher: workflow.yml on GitM-Labs/runtime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gitm_labs-0.0.4-py3-none-any.whl.

File metadata

  • Download URL: gitm_labs-0.0.4-py3-none-any.whl
  • Upload date:
  • Size: 129.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for gitm_labs-0.0.4-py3-none-any.whl
Algorithm Hash digest
SHA256 835cafd9f5bc88d34f0cfde852d0042e306076358936ac5a9b2564cfabf339f5
MD5 c8b4410e92e504c23a5559dae87f3faf
BLAKE2b-256 3e2442b1992b1d1f743275a5bee94c97afa78a57e9174ef21d24c0a665b5e482

See more details on using hashes here.

Provenance

The following attestation bundles were made for gitm_labs-0.0.4-py3-none-any.whl:

Publisher: workflow.yml on GitM-Labs/runtime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page