Skip to main content

LLM memory capacity planner for the NVIDIA DGX Spark (unified-memory & bandwidth aware).

Project description

sparkfit

CI License: MIT Python 3.8+ Linting: Ruff Zero dependencies

LLM memory capacity planner for the NVIDIA DGX Spark (GB10).

Catalog tools answer "does this model fit on my hardware, yes or no?" on generic GPUs. sparkfit treats the DGX Spark as what it is: a 128 GB unified-memory machine that is, above all, memory-bandwidth bound (273 GB/s shared between CPU and GPU). So it answers three questions a fit/no-fit catalog does not:

  1. Unified-memory budgeting. How the 128 GB is really split between weights, KV-cache, activations, the CUDA/serving framework, and the slice you must leave for the Grace CPU side and the OS.
  2. Will it actually be fast? A memory-bandwidth roofline estimate of decode throughput (tokens/s). On Spark, "it fits" and "it is usable" are two different things, and bandwidth, not capacity, is the limit.
  3. How to make it fit or scale. A quantization advisor (the highest-quality format that fits with a safety margin) and a concurrency planner (how many parallel streams or model replicas fit at once).

Pure Python standard library. No dependencies.

Install

sparkfit is a CLI, so pipx is the cleanest way to install it (isolated environment, global sparkfit command):

pipx install sparkfit        # or: pip install sparkfit

To get the latest unreleased code, install from source instead:

pipx install git+https://github.com/engineering87/sparkfit.git

On a fresh DGX Spark (DGX OS) you may need pipx first. Note that a plain system-wide pip install is blocked by PEP 668 on these systems:

sudo apt install -y pipx && pipx ensurepath   # then restart the shell

Plain pip works inside a virtual environment:

python3 -m venv .venv && . .venv/bin/activate
pip install git+https://github.com/engineering87/sparkfit.git

Or clone and run the single file, no install needed (handy for pasting onto a fresh DGX Spark):

git clone https://github.com/engineering87/sparkfit.git
cd sparkfit
python src/sparkfit.py llama3.1-70b  # or: ./sparkfit llama3.1-70b

Quick start

Pass just a model name and sparkfit does the rest: it auto-selects the highest-quality quantization that fits with margin, a typical context (8192), and shows the budget, the verdict, estimated tok/s, and the alternatives.

sparkfit llama3.1-70b          # built-in catalog id
sparkfit "llama 70"            # partial match -> llama3.1-70b
sparkfit 72b                   # even a fragment works
sparkfit Qwen/Qwen2.5-7B       # Hugging Face id -> reads params online
sparkfit /models/my-model/config.json   # a local config.json (file or dir)

sparkfit planning llama3.1-70b on a DGX Spark

Override any default: sparkfit qwen2.5-72b -c 16384 -n 4 -q q5_k_m.

Co-serving several models

On a unified-memory box you often run more than one model at once (say a generator plus an embedding and a reranker). serve sums their footprints against the 128 GB and splits the shared 273 GB/s across the models decoding at the same time, so you see both what co-resides and how much each slows the others down:

sparkfit serve qwen2.5-7b:fp8:8192 gemma2-9b:fp8:4096 llama3.2-3b:fp8:2048

Each token is NAME[:quant[:context]]; quant and context fall back to -q and -c. As a worst case, N models decoding at once each get about their solo speed divided by N; idle or bursty models free their share.

--live: plan against the memory free right now

Run on the Spark, --live budgets against currently-free memory (read from /proc/meminfo and nvidia-smi) instead of the theoretical 128 GB, and waives the OS reserve, since the free figure already excludes the OS and running processes. Useful when you already have models or containers loaded and want to know what you can still add.

sparkfit llama3.1-8b --live
sparkfit qwen2.5-32b --live

Commands

Command What it does
(none) Pass just a model and get a one-shot smart report with sensible defaults
plan Full unified-memory budget plus a decode-speed roofline for one configuration
advise Recommends the highest-quality quant that fits with a 10% margin
fit Scans the model DB for what fits; --concurrency-scan finds max parallel streams
serve Plans several models co-served at once: shared memory and shared bandwidth
scan Reads live memory when run on the Spark
models Lists the built-in model database

Every command supports --json for machine-readable output (CI, dashboards, scripting).

Key flags

-q/--quant (default q4_k_m for plan, advise, fit; auto in quick mode), -c/--context, -b/--batch, -n/--concurrency, --kv-dtype {fp16,bf16,fp8,int8,q4}, --os-reserve (GB, default 8), --framework (GB, default 2), --total-mem, --bandwidth, --efficiency (default 0.70), --live, --json. For models not in the DB: --params --active --layers --hidden --kv-heads --head-dim.

Methodology

All assumptions are explicit constants near the top of src/sparkfit.py and can be overridden from the command line. Everything below is plain arithmetic, not a benchmark, so you can audit it and change any assumption.

Memory budget

streams      = batch * concurrency
weights      = params_B * 1e9 * bits_per_weight / 8
kv_cache     = kv_per_token * context * streams
activations  = streams * context * hidden * 2 * 2      # two fp16 work buffers
used         = weights + kv_cache + activations + os_reserve + framework
fits         = used <= total_memory

bits_per_weight comes from the QUANT_BITS table, which uses the real on-disk width of each format (for example q4_k_m is 4.85 bits, not 4.0; fp16 is 16). Sizes use decimal GB (1e9), matching how the 128 GB is advertised.

KV-cache per token

Standard attention (MHA, or grouped-query GQA):

kv_per_token = 2 * layers * kv_heads * head_dim * dtype_bytes

The 2 is for K and V; GQA is captured by kv_heads being smaller than the number of attention heads. dtype_bytes is 2 for fp16, 1 for fp8/int8, 0.5 for 4-bit.

Multi-head Latent Attention (MLA, DeepSeek V2/V3/R1) stores one compressed latent per layer instead of per-head K and V:

kv_per_token = layers * (kv_lora_rank + qk_rope_head_dim) * dtype_bytes

This is much smaller. For DeepSeek-R1 at 8k context it is about 0.6 GB, versus tens of GB for a naive per-head estimate.

Hybrid models (Qwen3.5, Jamba, ...) mix linear/SSM layers with periodic full attention; only the full_attention layers keep a growing cache, so just those are counted (read from layer_types or full_attention_interval).

Sliding-window attention (Mistral, Gemma, ...) caps the resident cache at the window size, so very long contexts do not keep growing it.

Decode speed (roofline)

Token generation is memory-bound: each step reads the active weights once plus the resident KV-cache of every concurrent stream. With a fraction eff of peak bandwidth actually reached:

streams           = batch * concurrency
bytes_per_step    = active_weights + kv_per_token * context * streams
tok_s_aggregate   = streams * bandwidth * eff / bytes_per_step
tok_s_per_stream  = tok_s_aggregate / streams

Defaults: bandwidth 273 GB/s, eff 0.70. Weights are read once per step but the KV-cache is read for all streams, so adding concurrency raises aggregate throughput while each stream gets slower. For MoE only the active parameters are read, which is why a large MoE can be far faster than its total size suggests.

Prefill (processing the prompt) is instead compute-bound for long prompts and bandwidth-bound (stream the weights once) for short ones, so the time to first token is modeled as max(compute_time, memory_time):

compute_time = (2 * active_params * prompt + 2 * layers * prompt^2 * hidden)
               / (peak_tflops * precision_mult * mfu)
memory_time  = weights_bytes / (bandwidth * eff)
ttft         = max(compute_time, memory_time)

precision_mult is 2 for fp8 and 4 for fp4 (native low-precision matmul), 1 otherwise; mfu (default 0.30) and peak_tflops are calibratable with --mfu and --compute-tflops. Set --prompt-tokens to size the prompt.

Reserves and parameter estimation

On unified memory the CPU and GPU share one pool, so by default 8 GB (OS/DGX plus the Grace side) and 2 GB (CUDA context plus serving framework) are held back. Both are adjustable; --live sets the OS reserve to 0 because the live-free figure already excludes it.

For Hugging Face auto-fetch, parameters are estimated from config.json by summing embeddings, attention projections, and the MLP or expert layers (including MLA projections and DeepSeek fine-grained MoE). Validated against real configs: Qwen2.5-7B gives 7.62B (actual 7.61B), Mixtral-8x7B gives 46.7B total and 12.9B active, and DeepSeek-V3 gives about 671B total and 37B active.

When you point at a local model directory, sparkfit reads the real on-disk weight size (from the safetensors index or the files) and uses it instead of the estimate, which is exact even for hybrid, multimodal, or mixed-precision models. You can also set it manually with --weights-gb.

Calibrating to your device

The defaults (70% bandwidth efficiency, 8 GB OS reserve, 2 GB framework) are conservative starting points. If you measure real numbers on your Spark, plug them in per run with --efficiency, --bandwidth, --os-reserve, --framework, or set them once via environment variables so they stick:

export SPARKFIT_EFFICIENCY=0.6     # measured decode tok/s divided by the roofline
export SPARKFIT_OS_RESERVE=10
sparkfit llama3.1-8b

To calibrate efficiency, run a known model on your serving stack, note the real decode tok/s, and divide by what sparkfit predicts at --efficiency 1.0.

A robust, architecture-independent shortcut is efficiency ~= measured_tok_s * weights_GB / 273, where weights_GB is the on-disk size of the weight files. Measured example: a DGX Spark serving an FP8 model under vLLM (no other load) reached 6.8 tok/s against 35.9 GB of weights, i.e. about 0.85 efficiency, well above the conservative 0.70 default.

Validation on a real DGX Spark

Measured on a DGX Spark serving FP8 models with vLLM:

  • Bandwidth efficiency: a single stream reached 6.8 tok/s on 35.9 GB of weights, about 0.85 to 0.89 of the 273 GB/s roofline (the 0.70 default is conservative).
  • Concurrency: aggregate throughput scaled almost linearly to 8 parallel streams (6.8, 13.5, 26.4, 51.9 tok/s; per-stream 6.8 down to 6.5), matching sparkfit's batched-decode model within about 5%.
  • Contention: co-serving a second model (a 4B embedding model under load) on the same unified memory cut the generator from 6.8 to 3.8 tok/s. sparkfit assumes a model gets the full 273 GB/s, so when you co-serve, model your share by lowering --bandwidth or --efficiency.

These are capacity-planning estimates, not measurements. For exact numbers, cross-check with sparkfit scan on the real machine.

DGX Spark (GB10) reference

128 GB unified LPDDR5x, 273 GB/s shared CPU+GPU bandwidth, Blackwell GPU with FP4 (about 1 PFLOP / 1000 TOPS), 20-core Grace Arm CPU. The 273 GB/s bandwidth is the dominant constraint for token generation.

Roadmap

  • Local cache of fetched Hugging Face configs.
  • Expanded built-in model database.
  • Optional Go single-binary port.

Contributing

See CONTRIBUTING.md and the Code of Conduct. Issues and pull requests are welcome.

License

MIT.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sparkfit-0.4.0.tar.gz (30.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sparkfit-0.4.0-py3-none-any.whl (25.1 kB view details)

Uploaded Python 3

File details

Details for the file sparkfit-0.4.0.tar.gz.

File metadata

  • Download URL: sparkfit-0.4.0.tar.gz
  • Upload date:
  • Size: 30.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sparkfit-0.4.0.tar.gz
Algorithm Hash digest
SHA256 b8edeccf3c3bca326729d8fc1633c40892771a933b895057e3609f5decc71c75
MD5 dabf0e5b08006b105b58921c4ec2ec92
BLAKE2b-256 0d79114405e460ad06632825273578ed9d52e1e0ecac376d391abf7cfa215c70

See more details on using hashes here.

Provenance

The following attestation bundles were made for sparkfit-0.4.0.tar.gz:

Publisher: release.yml on engineering87/sparkfit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file sparkfit-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: sparkfit-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 25.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sparkfit-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b874020a858f7a84645063335a5081fcdcd6e3c3384bd1f733e11d5ec90a4d66
MD5 1ec7cd2e272fd193ae530698f56f38af
BLAKE2b-256 90770c89d610f3a4e0e87452d1b11ab08290629a7dfcca15d702af39ffcf2532

See more details on using hashes here.

Provenance

The following attestation bundles were made for sparkfit-0.4.0-py3-none-any.whl:

Publisher: release.yml on engineering87/sparkfit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page