Performance Report Analysis Tool
This tool analyzes performance traces from Metal operations, providing insights into throughput, bottlenecks, and optimization opportunities.
Installation
This tool can be installed from PyPI:
pipx install tt-perf-report
Installing with pipx will automatically create a virtual environment and make the tt-perf-report command available.
Generating Performance Traces
- Build Metal with performance tracing (enabled in default build):
./build_metal
- Run your test in TT-Metal with the tracy module to capture traces:
python -m tracy -r -p -v -m pytest path/to/test.py
This generates a CSV file containing operation timing data.
Using Tracy Signposts
Tracy signposts mark specific sections of code for analysis. Add signposts to your Python code:
import tracy
# Mark different sections of your code
tracy.signpost("Compilation pass")
model(input_data)
tracy.signpost("Performance pass")
for _ in range(10):
model(input_data)
The tool uses the last signpost by default, which is typically the most relevant section for a performance test(e.g., the final iteration after compilation / warmup).
Common signpost usage:
--start-signpost NAME: Analyze ops after the specified signpost--end-signpost NAME: Analyze ops before the specified signpost--ignore-signposts: Analyze the entire trace--print-signposts: Prints any signposts within the window defined when using the start/end signpost arguments
Filtering Operations
The output of the performance report is a table of operations. Each operation is assigned a unique ID starting from 1. You can re-run the tool with different IDs to focus on specific sections of the trace.
Use --id-range to analyze specific sections:
# Analyze ops 5 through 10
tt-perf-report trace.csv --id-range 5-10
# Analyze from op 31 onwards
tt-perf-report trace.csv --id-range 31-
# Analyze up to op 12
tt-perf-report trace.csv --id-range -12
This is particularly useful for:
- Isolating decode pass in prefill+decode LLM inference
- Analyzing single transformer layers without embeddings/projections
- Focusing on specific model components
Output Options
--min-percentage value: Hide ops below specified % of total time (default: 0.5)--color/--no-color: Force colored/plain output--csv FILENAME: Output the table to CSV format for further analysis or inclusion into automated reporting pipelines--no-advice: Show only performance table, skip optimization advice--active-experts K: Use K active experts per input batch group forttnn.sparse_matmulrows whose CSV attributes do not include numericnnz--arch ARCH: Override architecture/SKU detection. Usep100for Blackhole P100 traces because profiler CSVs identify the chip family but not the card SKU.
Understanding the Performance Report
The performance report provides several key metrics for analyzing operation performance:
Core Metrics
- Device Time: Time spent executing the operation on device (in microseconds)
- Op-to-op Gap: Time between operations, including host overhead and kernel dispatch (in microseconds)
- Total %: Percentage of total execution time spent on this operation
- Cores: Number of compute cores used by the operation. DRAM-sharded matmuls use the architecture's DRAM-interface workers: 12 on Wormhole, 8 on Blackhole P150, and 7 on Blackhole P100.
- Available Cores: Worker cores the operation could have used, read per operation from newer profiler CSVs. On a run that partitions the chip into subdevices this is that subdevice's own budget; otherwise it is the full worker grid the profiler reports. When the whole column is absent it falls back to the architecture's registered grid (e.g. 64 on Wormhole, 110 on Blackhole, 20 on
bh20andn1); when only an individual cell is blank or malformed, it falls back to the largest budget the file does report, which on a partitioned run may be another subdevice's. Grid-size advice and the Cores coloring are measured against this value rather than against the whole chip; FLOPs % is unaffected, since utilization is based on the cores the operation actually used - Sub Device ID: Subdevice the operation ran on. Blank means the full worker grid only when the input carries the
SUB DEVICE IDcolumn. On a capture that predates that column every cell is blank because no id was recorded, so a blank there means unknown rather than full-grid — including on a run whose differing Available Cores budgets show the chip was partitioned. The terminal table hides the column in that case, but--csvalways emits it, so downstream consumers should read an entirely blank column as absent data, not as confirmed full-grid operation
Sub Device ID appears in the terminal table only when the run reports subdevices, and Available Cores only when subdevices or differing core budgets are reported — otherwise they would be columns of blanks or of one repeated value. Both are always present in --csv output, whose column set and order do not vary with the input.
Upgrading from 1.2.x: the two new columns are appended after Global Call Count, which shifts Advice and Raw OP Code two positions to the right. Read
--csvoutput by header name rather than by column index. A cell whose text would otherwise be evaluated as a spreadsheet formula (one opening with=,+,-or@) is written with a leading apostrophe.
Performance Metrics
- DRAM: Memory bandwidth achieved (in GB/s)
- DRAM %: Percentage of theoretical peak DRAM bandwidth (288 GB/s on Wormhole, 512 GB/s on Blackhole P150, or 448 GB/s on Blackhole P100)
- Overall DRAM roofline: The total row reports modeled DRAM bandwidth and DRAM % across the visible report window
- FLOPs: Compute throughput achieved (in TFLOPs)
- FLOPs %: Percentage of theoretical peak compute for the given math fidelity
- Bound: Performance classification of the operation:
DRAM: Memory bandwidth bound (>65% of peak DRAM)FLOP: Compute bound (>65% of peak FLOPs)BOTH: Both memory and compute boundSLOW: Neither memory nor compute boundHOST: Operation running on host CPU
Additional Fields
- Math Fidelity: Precision configuration used for matrix operations. Utilization is based on the operation's actual core count. Blackhole-family per-core peaks use phase divisors (HiFi4=/4, HiFi3=/3, HiFi2=/2, LoFi=/1). Wormhole uses published chip peaks; HiFi3 is HiFi4×4/3 (LoFi is empirical). Full-chip reference peaks are:
HiFi4: Highest precision — Wormhole 74 TFLOPs, Blackhole ~166 TFLOPsHiFi3: High precision — Wormhole ~98.7 TFLOPs, Blackhole ~221 TFLOPsHiFi2: Medium precision — Wormhole 148 TFLOPs, Blackhole ~332 TFLOPsLoFi: Lowest precision — Wormhole 262 TFLOPs, Blackhole ~664 TFLOPs
The tool automatically highlights potential optimization opportunities:
- Red op-to-op times indicate high host or kernel launch overhead (>6.5μs)
- Red core counts indicate underutilization (fewer than 10 cores, and less than half of the cores the operation was given), excluding DRAM-sharded matmuls
- Green core counts indicate either all the cores the operation was given — the subdevice's budget on a partitioned run — or a DRAM-sharded matmul, which runs on a fixed set of DRAM-interface workers rather than on a grid it could grow into
- Green DRAM % and FLOPs % indicate good utilization of available resources
- Yellow metrics indicate room for optimization
Examples
Note:
trace.csvin the examples below refers to your input CSV file (the performance trace you want to analyze).
Typical use:
tt-perf-report trace.csv
Merge traces captured on multiple machines from the same workload run:
tt-perf-report trace_host0.csv trace_host1.csv trace_host2.csv
Build a table of all ops with no advice:
tt-perf-report trace.csv --no-advice
View ops 100-200 with advice:
tt-perf-report trace.csv --id-range 100-200
Export the table of ops and columns as a CSV file:
tt-perf-report trace.csv --csv my_report.csv
Metadata
Release files for tt-perf-report 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tt_perf_report-1.3.0.tar.gz | 71.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tt_perf_report-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 120.2 kB
Release files / tt_perf_report-1.3.0.tar.gz
| Download URL | tt_perf_report-1.3.0.tar.gz |
|---|---|
| Size | 71.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
364aaed88a022e6e17161ce764815bc16cb4e6b88da303c125c19ff8b86edb68
|
|
BLAKE2b-256 checksum How to use checksums |
a4b4409327d7170aabd3d8433899b821d35c096c32c77f8a9d1056b732931a01
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.
Transparency logRelease files / tt_perf_report-1.3.0-py3-none-any.whl
| Download URL | tt_perf_report-1.3.0-py3-none-any.whl |
|---|---|
| Size | 48.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
16f0974e6c3106fcf40dacc5e93bd8380bb0ccb94114cd5845f67103bc2d0957
|
|
BLAKE2b-256 checksum How to use checksums |
64294133671bc85924596cb6c1b18fec61348ed760fd2d069afc64601d251ee1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.
Transparency log