Skip to main content

gitm-labs

image

Behavioral compiler + intervention runtime for GPU-intensive workloads. Given a workload and a time budget, gitm-labs autonomously profiles, attributes, and applies kernel-level interventions to hit a target performance improvement — and produces a provenance report showing exactly what it changed and why.

Install

pip install gitm-labs

NVIDIA GPU support:

pip install "gitm-labs[nvidia]"

Optional extras: bench (HFT/biotech/edge benchmark harness), prometheus, otlp, s3.

Requires: Python 3.10+, NVIDIA (NVML + CUPTI) or AMD (ROCm SMI + rocprof) GPU.

Quick start

export GITM_S3_ROOT="s3://your-bucket/gitm"   # durable store for datasets + run outputs
export GITM_SCRATCH="/mnt/nvme/gitm"           # local ephemeral run dir (defaults to ~/.cache/gitm)

gitm run --workload hft-lob --budget 24h --target 15%

Workloads: hft-lob (HFT order-book), af2 (AlphaFold2 protein inference), kitti (3D LiDAR detection).

--budget is the wall-clock time limit. --target is the performance improvement fraction gitm-labs commits to delivering, or issues a diagnostic explaining why the floor could not be met.

Verify your environment first:

gitm doctor

Embedded API

from gitm import optimize

result = optimize(engine, budget="24h", target=0.15)

The 24-hour loop

gitm-labs runs a five-phase autonomous loop within the allotted budget:

Phase Hours What happens
1. Profile 0–2 Capture event + state telemetry; fingerprint workload; build predicted execution graph
2. Attribute 2–6 Compute residuals against predicted graph; run causal attribution
3. Rank 6–12 Query intervention library; rank candidates via counterfactual replay
4. Apply 12–20 Apply top-N interventions with rollback gates
5. Report 20–24 Stabilize; write provenance report (claim → evidence → intervention → delta)

Architecture

gitm-labs separates the empirical half (what happened) from the predicted half (what should have happened). Everything downstream operates on residuals — the difference between the two.

Two telemetry planes

State telemetry (gitm.telemetry)

Point-in-time samples of GPU state at ~1 Hz: utilization, memory, power, clocks, temperature, throttle reasons, NVLink throughput, ECC counters.

Source: NVML (NVIDIA) / ROCm SMI (AMD). Cost: ~microseconds per sample.

Event telemetry (gitm.tracer)

Per-kernel activity records with start/end timestamps, stream IDs, and memory transfer events.

Source: CUPTI (NVIDIA) / rocprof (AMD). Required for kernel-time invariant checks.

Deviation invariants

The monitor checks observed-vs-predicted against three invariants:

  1. Kernel-time — per-kernel duration must lie within roofline bounds.
  2. Memory-traffic — per-kernel bytes-moved must match predicted.
  3. Stream-concurrency — predicted-concurrent kernels must overlap.

See docs/invariants.md.

Module responsibilities

Module Responsibility
gitm.telemetry Vendor-backend autodiscovery, NVML/ROCm SMI samples, pluggable sinks
gitm.tracer Event-telemetry capture (CUPTI/rocprof), trace schema, context manager
gitm.planner Behavioral Compiler — roofline-based predicted execution graph
gitm.optimizer.monitor Deviation monitor — residuals against 3 invariants
gitm.optimizer.attribution Granger + doubly-robust on residual subgraph
gitm.optimizer.replay Counterfactual replay for predicted intervention delta
gitm.optimizer.qualification Workload fingerprint gate (commit / diagnose)
gitm.optimizer.report Provenance chain renderer (claim → evidence → intervention → delta)
gitm.kernels Curated intervention library — 15–20 levers with applicability + safety
gitm.agents Autonomous policy — selects interventions, drives rollback
gitm.scheduler 24-hour loop phase orchestration

Data layout

Two environment variables control where data lives:

export GITM_S3_ROOT="s3://your-bucket/gitm"   # canonical store (datasets + run outputs)
export GITM_SCRATCH="/mnt/nvme/gitm"           # local ephemeral dir (defaults to ~/.cache/gitm)

Layout under $GITM_S3_ROOT:

datasets/{hft,biotech,edge}/    # benchmark inputs (immutable, sha256-pinned)
runs/                            # baseline + run outputs
traces/                          # captured event-telemetry traces
telemetry/                       # state-telemetry samples

Local scratch is ephemeral and synced to S3 after each run.

Primary interfaces

# tracer
with gitm.tracer.capture(out_path: Path) -> ContextManager[Trace]: ...

# planner
graph = gitm.planner.predict_graph(model: ModelSpec, hw: HardwareSpec, batch: BatchConfig) -> Graph

# monitor
residuals = gitm.optimizer.monitor.residuals(trace: Trace, graph: Graph) -> Residuals
violations = gitm.optimizer.monitor.check_invariants(residuals, invariants) -> list[Violation]

# attribution
hypotheses = gitm.optimizer.attribution.attribute(residuals: Residuals, graph: Graph) -> RankedHypotheses

# report
report_md = gitm.optimizer.report.write(claims: list[Claim], provenance: Provenance) -> str

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gitm_labs-0.0.5.tar.gz (313.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gitm_labs-0.0.5-py3-none-any.whl (135.7 kB view details)

Uploaded Python 3

File details

Details for the file gitm_labs-0.0.5.tar.gz.

File metadata

  • Download URL: gitm_labs-0.0.5.tar.gz
  • Upload date:
  • Size: 313.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for gitm_labs-0.0.5.tar.gz
Algorithm Hash digest
SHA256 0510d99fb110d2cd78c56e0ad240769fd20dd5206eee6a2669c9304ff72b206a
MD5 1d9cb78bb2ff35ec88973fd8e1d9ffba
BLAKE2b-256 4490d882c1d74f7dcad870249dbe3bd31ddb1fdb98a998e09bcac56fe04cedac

See more details on using hashes here.

Provenance

The following attestation bundles were made for gitm_labs-0.0.5.tar.gz:

Publisher: workflow.yml on GitM-Labs/runtime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gitm_labs-0.0.5-py3-none-any.whl.

File metadata

  • Download URL: gitm_labs-0.0.5-py3-none-any.whl
  • Upload date:
  • Size: 135.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for gitm_labs-0.0.5-py3-none-any.whl
Algorithm Hash digest
SHA256 d7572726dd6c6538749692f769259c75bb1e57cf306162f800f089d45da5b206
MD5 bcf60a1a8d1a106cc2b946bd420bf7bb
BLAKE2b-256 d453f4911b350f3279882b93846ee1e3740bbb3e4f5171d81a12a92cb74f4c68

See more details on using hashes here.

Provenance

The following attestation bundles were made for gitm_labs-0.0.5-py3-none-any.whl:

Publisher: workflow.yml on GitM-Labs/runtime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page