ActLens
See inside a language model while it reads your prompt.
ActLens is an interactive viewer for the activations of any supported Hugging Face model. Type a prompt, pick an activation and a layer, and browse the result as a token × channel heatmap: residual stream, attention patterns, Q/K/V, MLP internals and more. It runs locally, in one command.
Features
- Every activation, every layer. Residual stream, attention (Q, K, V, RoPE, patterns, context, output) and MLP (gate, up, SwiGLU, down), captured lazily and cached.
- Four views. Layer shows a token × channel heatmap for one layer; Across layers shows a per-token statistic (L2 norm, |max|, mean, std, kurtosis, or a single channel) for every layer at once; Architecture draws the decoder block and opens any node's activation; Logit lens shows what the model would predict after every layer.
- Attention explorer. A grid of all heads, a full query × key map per head, and a layer × head map of entropy, sink mass and attention distance.
- Find outlier channels. Rank channels by |max|, std or |mean|; jump straight to a head.
- Distributions. Histograms and summary statistics over a region, a channel, a token, or the whole layer.
- Fast to navigate. Pan, zoom, brush and inspect; large activations are pooled server-side so zooming out stays cheap.
- Publication-ready export. PNG and PDF with title, prompt, axes and colorbar at up to 4×.
- Extensible. Add support for a new architecture with a small adapter.
See Examples for a massive activation on the first token, attention sinks and previous-token heads, each a few clicks away.
Quick start
pip install actlens # Python >= 3.10; use a fresh virtual environment
actlens # loads Qwen/Qwen3-0.6B and serves the UI at http://127.0.0.1:8000
The first run downloads the model weights from the Hugging Face Hub. ActLens is developed and tested against torch 2, transformers 5 and nnsight 0.7.
actlens -m Qwen/Qwen2.5-0.5B # another Hugging Face model id or a local path
actlens --device cuda --dtype bfloat16 # device: auto | cpu | cuda | mps; dtype: float32 | float16 | bfloat16
actlens --port 9000 --open # custom port, open the browser when ready
actlens --cache-mb 4096 # activation cache budget (or $ACTLENS_CACHE_MB)
You can load more models from the UI at any time.
Security note. The server binds to
127.0.0.1by default. The API can load any model and has no authentication, so only use--host 0.0.0.0on a network you trust.
Google Colab
No local GPU? Run the model on a free Colab GPU and view the visualization in your own browser.
- Open the notebook and pick a GPU runtime (Runtime → Change runtime type → T4 GPU).
- Run the first cell. It installs ActLens, starts the server and opens a Cloudflare tunnel.
- Click the Open ActLens link it prints. The UI opens in your browser; keep the Colab tab open while you use it.
Or in any notebook:
!pip install -q actlens
from actlens.colab import launch
launch(model="Qwen/Qwen3-0.6B", dtype="float16")
The link contains a random access token, and the server rejects requests without it. Anyone who has the full link can use
your session, so do not share it. Use actlens.colab.stop() to shut everything down. If the page does not load, check
actlens.log.
You can also protect a server of your own with actlens --token (generates a token and prints the URL) or --token VALUE.
Remote server (SSH)
Run the model on a GPU server and view it in your local browser. The server only listens on 127.0.0.1, so forward the port over SSH:
# on the server (inside tmux or screen so it survives a disconnect)
pip install actlens
actlens --device cuda --dtype bfloat16
# on your machine
ssh -L 8000:127.0.0.1:8000 user@server # then open http://127.0.0.1:8000
VS Code Remote-SSH forwards the port for you. Connect to the server, run actlens in the integrated terminal, and
VS Code detects the listening port and lists it in the Ports panel. Click the link there (or the globe icon) to open
the UI in your local browser. No ssh -L needed.
The weights are downloaded on the server, so it needs access to the Hugging Face Hub (set HF_ENDPOINT for a mirror, and
HF_TOKEN for gated models). On a shared machine, add --token so other users cannot use your session.
Using the viewer
Pick an activation and a layer in the selection bar ([ and ] step through layers), then choose a view.
| Interaction | Action |
|---|---|
| Drag | Pan |
| Pinch, or ⌘/Ctrl + scroll | Zoom |
| Shift + drag | Select a region |
| Click | Place the cursor |
| Double-click | Reset the view |
[ / ] |
Previous / next layer |
| Esc | Clear the selection |
For head-structured activations (q k v q_norm k_norm q_rope k_rope attn_ctx) the channel axis is head × head_dim,
with ticks like h3·17 and faint separators between heads; the Head control jumps the window to one head.
Available activations
The picker lists only what the loaded model provides.
| Group | Activations |
|---|---|
| Residual | resid_pre, resid_mid, resid_post |
| Attention | attn_norm, q, k, v, q_norm, k_norm, q_rope, k_rope, attn_pattern, attn_ctx, o |
| MLP | mlp_norm, gate, up, silu, swiglu, mlp_act, down |
Logit lens
Logit lens unembeds the residual stream of one token after every layer, using the model's own final norm and output head (including Gemma's final logit soft-capping), and lists the top-k tokens with their probabilities.
- Pick the token position from the strip above the table, the stream (
resid_pre,resid_midorresid_post) and k. - Track follows one token through the layers (by default the next token of your prompt): its rank and probability get a column of their own, and the side panel plots the rank across layers and the layer where it first becomes top-1. The side panel also plots the entropy of the next-token distribution.
- The last layer's
resid_postreproduces the model's real output. Early layers are often noise, because the residual stream is not yet in the space the output head reads; the lens gets sharper with depth, and more so on some models (GPT-2) than on others. - Only one position is unembedded per request, so it stays cheap even with a large vocabulary. The API is
GET /api/run/{run_id}/logit_lens?act=resid_post&pos=-1&k=10&target=<token id>.
Distribution panel
- Values: histogram and summary (mean, std, percentiles, kurtosis, skew) of the raw values for the visible window, a brushed selection, the cursor's channel or token, or the whole layer.
- Per channel / Per token: one statistic per channel or token, shown as a histogram with a clickable list of the top outliers.
Export
The PNG and PDF buttons render the current view at 1–4×. The histogram and the attention head grid have their own PNG export. PDFs embed the figure as a high-resolution image so CJK tokens render correctly.
Supported models
| Adapter | model_type |
Models |
|---|---|---|
llama |
llama, qwen2, qwen3, mistral, gemma, olmo, stablelm, granite, and models with the same module layout |
Llama, Qwen2/2.5/3, Mistral, Gemma, OLMo, StableLM, Granite, SmolLM |
gpt_neox |
gpt_neox |
Pythia, Dolly, RedPajama-INCITE |
phi3 |
phi3 |
Phi-3, Phi-3.5, Phi-4-mini |
olmo2 |
olmo2 |
OLMo-2 |
gemma2 |
gemma2, gemma3_text |
Gemma-2, Gemma-3 (text-only checkpoints such as 270m / 1b) |
gpt2 |
gpt2 |
GPT-2, DistilGPT-2 |
Qwen3-0.6B and GPT-2 are tested on real checkpoints; every adapter is also tested on a tiny random model. An unsupported model is rejected at load time with a message naming the missing modules.
Want another architecture? See Adding an architecture.
Troubleshooting
- Attention patterns need eager attention. ActLens loads models with
attn_implementation="eager"because SDPA does not return attention probabilities. - Slow first open of an activation. The first time you open an activation, one forward pass captures it for every layer (about 0.1–0.4 s for Qwen3-0.6B); after that it comes from the cache.
- Gated or private models. Accept the license on huggingface.co and run
huggingface-cli login(or setHF_TOKEN) before loading them. The UI shows download progress and, if a load fails, why and what to try. - Running out of memory. Use a smaller model,
--dtype float16orbfloat16, or lower--cache-mb.
Contributing
Development setup, tests, architecture notes and the release process are in CONTRIBUTING.md.
License
Release files for actlens 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| actlens-0.2.0.tar.gz | 510.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| actlens-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 917.6 kB
Release files / actlens-0.2.0.tar.gz
| Download URL | actlens-0.2.0.tar.gz |
|---|---|
| Size | 510.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4209fcbed98b844542eef59106eae154c34636033cfffa00e4501b87ced9a651
|
|
BLAKE2b-256 checksum How to use checksums |
da64c8622689de3bff4018c572fac905535d0ba91582eb0430442e7055e71400
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / actlens-0.2.0-py3-none-any.whl
| Download URL | actlens-0.2.0-py3-none-any.whl |
|---|---|
| Size | 407.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
16897612583a004a6915b8aef1634244448710d574e250763178e297b4a34276
|
|
BLAKE2b-256 checksum How to use checksums |
ef9701b94f5526230b030ebfd0374c4ed9a8b08b6af26c953ffb6135054b2bfb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log