National Deep Inference Fabric
Website | nnsight | Discord | Paper
NDIF is the server backend for nnsight. It runs user-submitted intervention code — hooks, captures, edits, generation — against large models on a shared GPU cluster. Researchers point nnsight at an NDIF endpoint and run experiments on models too big to fit on their own hardware.
With no authentication configured — the default — every request is trusted
and its code runs inside the model process. Only when auth is on and a key is
not granted trusted does the code run in a separate process, one fresh
process per request; that isolation is process-based and still being hardened —
see docs/concepts/sandbox-execution.md.
This repo is the server. For the client, see
nnsight; it is an ordinary dependency
here, and just up bind-mounts a local checkout over it for client-side
development.
Quick start
Three ways to stand up your own NDIF. All three need an NVIDIA GPU and a CUDA driver; the container routes also need the NVIDIA container toolkit.
Trust default: with no NDIF_POSTGRES_URL the API is unauthenticated, and an
unauthenticated NDIF runs every request trusted — the submitted block executes
in-process next to the model weights and models load with trust_remote_code. That
is the intended default for running one for yourself. Before anyone else can reach
it, work through docs/runbooks/enable-auth.md.
1. docker run — the published image, whole stack in one container
docker run --gpus all --shm-size 4g -p 8001:8001 -p 9000:9000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ndif/ndif:0.1.0
NDIF_SERVICE defaults to all, so the container starts redis, minio, ray and
the API together. 8001 is the API; 9000 is the object store, which the client
downloads results from, so publish it too. Tags: 0.1.0-cu126, 0.1.0-cu130,
0.1.0 (= cu126) and latest; pick the CUDA line your driver supports. The
entrypoint is the ndif CLI, so any other command works the same way:
docker run --rm ndif/ndif:0.1.0 version # ndif, nnsight, torch+CUDA, transformers, ray
docker run --rm --gpus all ndif/ndif:0.1.0 doctor
2. Docker Compose — the development stack, built from this checkout
Each service in its own container, next to Postgres and the full telemetry set
(Loki, InfluxDB, Prometheus, Grafana). just
wraps the compose commands.
just up # build (first time) + start the whole stack, detached
just logs api # follow a service's logs
just ta # down -> rebuild -> up, after a source change
just down # tear it down
3. From source — the ndif CLI, no Docker
No checkout needed — the package is on PyPI:
pip install torch --index-url https://download.pytorch.org/whl/cu126 # first, or pip picks PyPI's default (CUDA 13) wheel
pip install "ndif[api,ray]" # add metrics,postgres,dashboard as you need them
conda install --override-channels -c conda-forge redis-server minio-server
ndif doctor # versions, binaries, GPU, connectivity
ndif start # redis, minio, ray, api — detached
ndif stop # ...and back down, Ray daemons included
From a checkout, pip install -r requirements.txt first gives you the pinned
dependency set the image is built from, then pip install ".[api,ray]".
redis-server and minio come from conda-forge because MinIO no longer
publishes standalone binaries (--override-channels sidesteps the anaconda
terms-of-service prompt a stock miniconda raises); the other option is copying
the binary out of the quay.io/minio/minio image — see
docs/operating/quickstart.md.
Then run a remote trace
import nnsight
nnsight.CONFIG.API.HOST = "http://localhost:8001"
from nnsight.modeling.transformers import TransformersModel
model = TransformersModel("openai-community/gpt2", task="text-generation")
with model.trace("The Eiffel Tower is in the city of", remote=True):
hidden = model.transformer.h[-1].output.save()
No API key is needed; the first request for a model deploys it. Deeper detail for all three routes is in docs/operating/quickstart.md.
What's running
just ps lists the compose stack. The NDIF pieces:
| Service | Where | What |
|---|---|---|
api |
localhost:8001 |
Accepts nnsight requests, queues them, streams results back. |
ray |
GPU node | Loads models and runs the traced blocks (plain detached Ray actors — NDIF does not use Ray Serve). |
dashboard |
localhost:8081 |
Deploy/evict/status, schedules, request monitor. Compose only; not part of NDIF_SERVICE=all. |
Redis, MinIO (object store), Postgres (API-key auth), and Loki/InfluxDB/Grafana
(telemetry) round out the compose file; see docker/docker-compose.yml.
Development
just ta # down -> rebuild -> up (full refresh after a code change)
just ta ray # ...targeting a single service
just build && just up
The ndif CLI is the image entrypoint (ENTRYPOINT ["ndif"], CMD ["start", "--foreground"]); NDIF_SERVICE selects which service(s) a container runs. It
accepts a space/comma list, and all expands in place — "all dashboard" is the
core stack plus the admin UI. Configuration is read from the environment (see the
environment: blocks in the compose file).
Agent-facing documentation lives in CLAUDE.md and docs/.
Configuration
Everything is configured through NDIF_* environment variables. There is no
central config file — each service/provider reads its own vars (with a working
single-host default) at startup, so a bare just up runs end-to-end with none of
these set. Override them in the compose environment: blocks, a .env file, or
the shell. Empty defaults for the optional providers (Postgres/Loki/Influx) mean
that provider is off until you set its URL.
Core / service
| Variable | Default | Description |
|---|---|---|
NDIF_SERVICE |
all (image ENV) |
Which service(s) this container runs: redis, minio, ray, api, dashboard, all, or a space/comma list. all means redis, minio, ray, api. |
NDIF_ENVIRONMENT |
dev |
Deployment tag attached to logs/metrics. |
NDIF_LOG_LEVEL |
INFO |
Root log level. |
NDIF_HOME |
~/.ndif |
CLI state directory. |
API
| Variable | Default | Description |
|---|---|---|
NDIF_API_URL |
http://localhost:8001 |
Base URL of the API (compose uses http://api:8001). |
NDIF_API_PORT |
8001 |
Port the API binds. |
NDIF_API_WORKERS |
1 |
Gunicorn worker count. |
NDIF_API_TIMEOUT |
120 |
Gunicorn worker timeout (seconds). |
NDIF_API_KEY |
(unset) | API key the dashboard's monitor cron sends with its probe traces (jobs/monitor.py). Not read by the CLI or the API. |
Request queue
| Variable | Default | Description |
|---|---|---|
NDIF_QUEUE_KEY |
queue |
Redis key backing the request queue. |
NDIF_QUEUE_FETCH_TIMEOUT_S |
10 |
Blocking-pop timeout when draining the queue. |
NDIF_QUEUE_FETCH_BATCH_MAX |
32 |
Max requests pulled per fetch. |
Autoscaling
| Variable | Default | Description |
|---|---|---|
NDIF_AUTOSCALING_INTERVAL_S |
5 |
How often the scaler evaluates the queue. |
NDIF_AUTOSCALING_BACKOFF_S |
120 |
Pause after a scale-up so the new replica can warm. |
NDIF_AUTOSCALING_WAIT_THRESHOLD_S |
30 |
Queue wait time that triggers a scale-up. |
NDIF_AUTOSCALING_MAX_REPLICAS |
3 |
Replica ceiling per model. |
Ray / cluster
| Variable | Default | Description |
|---|---|---|
NDIF_RAY_ADDRESS |
ray://localhost:10001 |
Ray client address the API/dashboard connect to. |
NDIF_RAY_HEAD_ADDRESS |
(empty) | Head-node address workers join (empty = start a head). |
NDIF_RAY_HEAD_PORT |
6385 |
Ray GCS head port (offset from Redis's 6379). |
NDIF_RAY_DASHBOARD_PORT |
8265 |
Ray dashboard port. |
NDIF_RAY_DASHBOARD_GRPC_PORT |
52366 |
Ray dashboard gRPC port. |
NDIF_RAY_METRICS_PORT |
8080 |
Ray's --metrics-export-port (the Prometheus scrape target). Not a Ray Serve port. |
NDIF_RAY_OBJECT_MANAGER_PORT |
8076 |
Ray object-manager port. |
NDIF_RAY_RESOURCE_NAME |
(empty) | Custom Ray resource label for this node. |
NDIF_RAY_TEMP_DIR |
/tmp/ray |
Ray temp/session directory. |
NDIF_RAY_HEAD_WAIT_INTERVAL_S |
2 |
Worker poll interval while waiting for the head. |
NDIF_RAY_HEAD_WAIT_RETRIES |
60 |
Worker retries before giving up on the head. |
Controller / deployments
| Variable | Default | Description |
|---|---|---|
NDIF_DEPLOYMENTS |
(empty) | ` |
NDIF_CONTROLLER_SYNC_INTERVAL_S |
30 |
How often the controller re-syncs its node set. Deployment changes are event-driven, not polled. |
NDIF_MINIMUM_DEPLOYMENT_TIME_SECONDS |
3600 |
Minimum lifetime before a model can be evicted. |
NDIF_MODEL_CACHE_PERCENTAGE |
0.9 |
Fraction of the node's host RAM the WARM (off-GPU) model cache may use. Not a GPU knob. |
NDIF_DEFAULT_MODEL_ACTOR_CLASS |
ndif.services.ray.deployments.modeling.base.ModelActor |
Actor class used to serve a model. Compose sets the sandboxed ...ray.sandbox.model.SandboxModelActor. |
NDIF_TP_MODEL_ACTOR_CLASS |
(unset) | Tensor-parallel actor class. Unset means tensor parallelism is off entirely. |
NDIF_DEFAULT_DTYPE |
bfloat16 |
Dtype models load in. |
NDIF_DEFAULT_EXECUTION_TIMEOUT_SECONDS |
(unset) | Per-request execution cap. Unset means no cap — set it before others can submit. |
NDIF_DEFAULT_PADDING_FACTOR |
0.15 |
Batch-padding memory factor. |
NDIF_DEFAULT_PADDING_BIAS |
524288000 |
Batch-padding memory bias in bytes (500 MiB). |
NDIF_MIN_NNSIGHT_VERSION |
(unset) | Minimum client nnsight version accepted. |
NDIF_MIN_PYTHON_VERSION |
(unset) | Minimum client Python version accepted. |
Redis / caches
| Variable | Default | Description |
|---|---|---|
NDIF_REDIS_URL |
redis://localhost:6379 |
Redis connection URL. |
NDIF_ENV_TTL_S |
300 |
TTL of the cached model-environment metadata. |
NDIF_ENV_TIMEOUT_S |
60 |
Timeout awaiting a fresh env entry. |
NDIF_STATUS_TTL_S |
60 |
TTL of the cached deployment status. |
NDIF_STATUS_TIMEOUT_S |
60 |
Timeout awaiting a fresh status entry. |
NDIF_STATUS_CACHE_FREQ_S |
10 |
Refresh frequency of the API's Redis-backed /status cache. |
Object store (S3 / MinIO)
| Variable | Default | Description |
|---|---|---|
NDIF_OBJECT_STORE_URL |
http://localhost:9000 |
S3-compatible endpoint result blobs stage to. |
NDIF_OBJECT_STORE_PUBLIC_URL |
(empty) | Public URL used when presigning (defaults to the endpoint). |
NDIF_OBJECT_STORE_ACCESS_KEY |
minioadmin |
Access key. |
NDIF_OBJECT_STORE_SECRET_KEY |
minioadmin |
Secret key. |
NDIF_OBJECT_STORE_BUCKET |
ndif-results |
Bucket for result blobs. |
NDIF_OBJECT_STORE_REGION |
us-east-1 |
Region sent to the S3 client. |
NDIF_OBJECT_STORE_VERIFY |
true |
Verify TLS to the endpoint. |
NDIF_OBJECT_STORE_CONSOLE_PORT |
9001 |
MinIO web console port (compose only). |
Auth — Postgres (empty URL ⇒ API runs unauthenticated)
| Variable | Default | Description |
|---|---|---|
NDIF_POSTGRES_URL |
(empty) | Connection URL for the user/API-key DB; empty disables auth. |
NDIF_POSTGRES_POOL_MIN |
1 |
Connection-pool minimum size. |
NDIF_POSTGRES_POOL_MAX |
10 |
Connection-pool maximum size. |
NDIF_POSTGRES_COMMAND_TIMEOUT_S |
10.0 |
Per-command timeout (seconds). |
Telemetry — InfluxDB (metrics)
| Variable | Default | Description |
|---|---|---|
NDIF_INFLUX_URL |
(unset — metrics off) | InfluxDB endpoint; set it to turn metrics on. |
NDIF_INFLUX_TOKEN |
(empty) | Write token. |
NDIF_INFLUX_ORG |
ndif |
Influx organization. |
NDIF_INFLUX_BUCKET |
metrics |
Target bucket. |
NDIF_INFLUX_ENABLED |
true |
Master switch for metric writes. |
NDIF_INFLUX_BATCH_SIZE |
500 |
Points buffered before a flush. |
NDIF_INFLUX_FLUSH_INTERVAL_MS |
1000 |
Max time between flushes (ms). |
NDIF_INFLUX_TIMEOUT_MS |
10000 |
Write request timeout (ms). |
Telemetry — Loki (logs) (empty URL ⇒ console-only logging)
| Variable | Default | Description |
|---|---|---|
NDIF_LOKI_URL |
(empty) | Loki push endpoint; empty disables log shipping. |
NDIF_LOKI_LEVEL |
INFO |
Minimum level shipped to Loki. |
NDIF_LOKI_QUEUE_MAX |
10000 |
Max buffered log records before dropping. |
Dashboard
| Variable | Default | Description |
|---|---|---|
NDIF_DASHBOARD_PORT |
8081 |
Port the dashboard binds. |
NDIF_DASHBOARD_USERNAME |
admin |
Admin username. |
NDIF_DASHBOARD_PASSWORD_HASH |
(empty) | Bcrypt hash of the admin password. |
NDIF_DASHBOARD_SESSION_SECRET |
change-me-please-this-is-not-secure |
Cookie-signing secret — set this in prod. |
NDIF_DASHBOARD_SESSION_TTL_DAYS |
7 |
Session cookie lifetime (days). |
NDIF_DASHBOARD_DEV_MODE |
false |
Bypasses the dashboard login entirely. Compose sets it true. |
NDIF_DASHBOARD_API_URL |
http://localhost:8001 |
NDIF API URL (falls back to NDIF_API_URL). |
NDIF_DASHBOARD_DATA_DIR |
~/ndif_dashboard |
Dashboard state directory. |
NDIF_DASHBOARD_FRONTEND_DIST |
<package>/frontend/dist |
Built Vue UI directory to serve. |
NDIF_DASHBOARD_MONITOR_URL |
http://localhost:8001 |
Target the monitor cron probes. |
NDIF_DASHBOARD_MONITOR_CRON |
*/10 * * * * |
Monitor cron schedule. |
NDIF_DASHBOARD_RECONCILE_CRON |
*/2 * * * * |
Reconcile cron schedule. |
Contributing
PRs welcome. Please read the Code of Conduct.
License
MIT © Northeastern University.
Citation
@article{fiottokaufman2024nnsightndifdemocratizingaccess,
title={NNsight and NDIF: Democratizing Access to Foundation Model Internals},
author={Jaden Fiotto-Kaufman and Alexander R Loftus and Eric Todd and Jannik Brinkmann and Caden Juang and Koyena Pal and Can Rager and Aaron Mueller and Samuel Marks and Arnab Sen Sharma and Francesca Lucchetti and Michael Ripa and Adam Belfki and Nikhil Prakash and Sumeet Multani and Carla Brodley and Arjun Guha and Jonathan Bell and Byron Wallace and David Bau},
year={2024},
eprint={2407.14561},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2407.14561},
}
Metadata
Release files for ndif 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ndif-0.1.1.tar.gz | 981.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ndif-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.6 MB
Release files / ndif-0.1.1.tar.gz
| Download URL | ndif-0.1.1.tar.gz |
|---|---|
| Size | 981.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7c7de4acb58d46a1e8785429094c5a0e5e79960ab5e3b62d6d2644176e610e14
|
|
BLAKE2b-256 checksum How to use checksums |
ac5b8389cd56418d232619fce973443b4801bd84a585826e0029042cfca46b12
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / ndif-0.1.1-py3-none-any.whl
| Download URL | ndif-0.1.1-py3-none-any.whl |
|---|---|
| Size | 570.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7113948c382f51db4981d81c4149e17efa5523ed102b361edc2fe4f2ab6d742d
|
|
BLAKE2b-256 checksum How to use checksums |
1c5434bcd738852a872038c96e498e1af50f6c6f626710ecfac8382bbbda37d0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log