Skip to main content

ActLens

See inside a language model while it reads your prompt.

ActLens is an interactive viewer for the activations of any supported Hugging Face model. Type a prompt, pick an activation and a layer, and browse the result as a token × channel heatmap: residual stream, attention patterns, Q/K/V, MLP internals and more. It runs locally, in one command.

PyPI Python License: MIT

ActLens showing the residual stream of Qwen3-0.6B

Features

  • Every activation, every layer. Residual stream, attention (Q, K, V, RoPE, patterns, context, output) and MLP (gate, up, SwiGLU, down), captured lazily and cached.
  • Four views. Layer shows a token × channel heatmap for one layer; Across layers shows a per-token statistic (L2 norm, |max|, mean, std, kurtosis, or a single channel) for every layer at once; Architecture draws the decoder block and opens any node's activation; Logit lens shows what the model would predict after every layer.
  • Attention explorer. A grid of all heads, a full query × key map per head, and a layer × head map of entropy, sink mass and attention distance.
  • Find outlier channels. Rank channels by |max|, std or |mean|; jump straight to a head.
  • Distributions. Histograms and summary statistics over a region, a channel, a token, or the whole layer.
  • Fast to navigate. Pan, zoom, brush and inspect; large activations are pooled server-side so zooming out stays cheap.
  • Publication-ready export. PNG and PDF with title, prompt, axes and colorbar at up to 4×.
  • Extensible. Add support for a new architecture with a small adapter.

See Examples for a massive activation on the first token, attention sinks and previous-token heads, each a few clicks away.

Quick start

pip install actlens      # Python >= 3.10; use a fresh virtual environment
actlens                  # loads Qwen/Qwen3-0.6B and serves the UI at http://127.0.0.1:8000

The first run downloads the model weights from the Hugging Face Hub. ActLens is developed and tested against torch 2, transformers 5 and nnsight 0.7.

actlens -m Qwen/Qwen2.5-0.5B             # another Hugging Face model id or a local path
actlens --device cuda --dtype bfloat16   # device: auto | cpu | cuda | mps; dtype: float32 | float16 | bfloat16
actlens --port 9000 --open               # custom port, open the browser when ready
actlens --cache-mb 4096                  # activation cache budget (or $ACTLENS_CACHE_MB)

You can load more models from the UI at any time.

Security note. The server binds to 127.0.0.1 by default. The API can load any model and has no authentication, so only use --host 0.0.0.0 on a network you trust.

Google Colab

No local GPU? Run the model on a free Colab GPU and view the visualization in your own browser.

Open in Colab

  1. Open the notebook and pick a GPU runtime (Runtime → Change runtime type → T4 GPU).
  2. Run the first cell. It installs ActLens, starts the server and opens a Cloudflare tunnel.
  3. Click the Open ActLens link it prints. The UI opens in your browser; keep the Colab tab open while you use it.

Or in any notebook:

!pip install -q actlens
from actlens.colab import launch
launch(model="Qwen/Qwen3-0.6B", dtype="float16")

The link contains a random access token, and the server rejects requests without it. Anyone who has the full link can use your session, so do not share it. Use actlens.colab.stop() to shut everything down. If the page does not load, check actlens.log.

You can also protect a server of your own with actlens --token (generates a token and prints the URL) or --token VALUE.

Remote server (SSH)

Run the model on a GPU server and view it in your local browser. The server only listens on 127.0.0.1, so forward the port over SSH:

# on the server (inside tmux or screen so it survives a disconnect)
pip install actlens
actlens --device cuda --dtype bfloat16

# on your machine
ssh -L 8000:127.0.0.1:8000 user@server      # then open http://127.0.0.1:8000

VS Code Remote-SSH forwards the port for you. Connect to the server, run actlens in the integrated terminal, and VS Code detects the listening port and lists it in the Ports panel. Click the link there (or the globe icon) to open the UI in your local browser. No ssh -L needed.

The weights are downloaded on the server, so it needs access to the Hugging Face Hub (set HF_ENDPOINT for a mirror, and HF_TOKEN for gated models). On a shared machine, add --token so other users cannot use your session.

Using the viewer

Pick an activation and a layer in the selection bar ([ and ] step through layers), then choose a view.

Interaction Action
Drag Pan
Pinch, or ⌘/Ctrl + scroll Zoom
Shift + drag Select a region
Click Place the cursor
Double-click Reset the view
[ / ] Previous / next layer
Esc Clear the selection

For head-structured activations (q k v q_norm k_norm q_rope k_rope attn_ctx) the channel axis is head × head_dim, with ticks like h3·17 and faint separators between heads; the Head control jumps the window to one head.

Available activations

The picker lists only what the loaded model provides.

Group Activations
Residual resid_pre, resid_mid, resid_post
Attention attn_norm, q, k, v, q_norm, k_norm, q_rope, k_rope, attn_pattern, attn_ctx, o
MLP mlp_norm, gate, up, silu, swiglu, mlp_act, down

Logit lens

Logit lens unembeds the residual stream of one token after every layer, using the model's own final norm and output head (including Gemma's final logit soft-capping), and lists the top-k tokens with their probabilities.

  • Pick the token position from the strip above the table, the stream (resid_pre, resid_mid or resid_post) and k.
  • Track follows one token through the layers (by default the next token of your prompt): its rank and probability get a column of their own, and the side panel plots the rank across layers and the layer where it first becomes top-1. The side panel also plots the entropy of the next-token distribution.
  • The last layer's resid_post reproduces the model's real output. Early layers are often noise, because the residual stream is not yet in the space the output head reads; the lens gets sharper with depth, and more so on some models (GPT-2) than on others.
  • Only one position is unembedded per request, so it stays cheap even with a large vocabulary. The API is GET /api/run/{run_id}/logit_lens?act=resid_post&pos=-1&k=10&target=<token id>.

Distribution panel

  • Values: histogram and summary (mean, std, percentiles, kurtosis, skew) of the raw values for the visible window, a brushed selection, the cursor's channel or token, or the whole layer.
  • Per channel / Per token: one statistic per channel or token, shown as a histogram with a clickable list of the top outliers.

Export

The PNG and PDF buttons render the current view at 1–4×. The histogram and the attention head grid have their own PNG export. PDFs embed the figure as a high-resolution image so CJK tokens render correctly.

Supported models

Adapter model_type Models
llama llama, qwen2, qwen3, mistral, gemma, olmo, stablelm, granite, and models with the same module layout Llama, Qwen2/2.5/3, Mistral, Gemma, OLMo, StableLM, Granite, SmolLM
gpt_neox gpt_neox Pythia, Dolly, RedPajama-INCITE
phi3 phi3 Phi-3, Phi-3.5, Phi-4-mini
olmo2 olmo2 OLMo-2
gemma2 gemma2, gemma3_text Gemma-2, Gemma-3 (text-only checkpoints such as 270m / 1b)
gpt2 gpt2 GPT-2, DistilGPT-2

Qwen3-0.6B and GPT-2 are tested on real checkpoints; every adapter is also tested on a tiny random model. An unsupported model is rejected at load time with a message naming the missing modules.

Want another architecture? See Adding an architecture.

Troubleshooting

  • Attention patterns need eager attention. ActLens loads models with attn_implementation="eager" because SDPA does not return attention probabilities.
  • Slow first open of an activation. The first time you open an activation, one forward pass captures it for every layer (about 0.1–0.4 s for Qwen3-0.6B); after that it comes from the cache.
  • Gated or private models. Accept the license on huggingface.co and run huggingface-cli login (or set HF_TOKEN) before loading them. The UI shows download progress and, if a load fails, why and what to try.
  • Running out of memory. Use a smaller model, --dtype float16 or bfloat16, or lower --cache-mb.

Contributing

Development setup, tests, architecture notes and the release process are in CONTRIBUTING.md.

License

MIT

Release files for actlens 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for actlens 0.2.0
File Size Uploaded
actlens-0.2.0.tar.gz 510.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for actlens 0.2.0
File Interpreter ABI Platform
actlens-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 917.6 kB

Release files / actlens-0.2.0.tar.gz

Download URL actlens-0.2.0.tar.gz
Size 510.0 kB
Tags Source
SHA-256 checksum
How to use checksums
4209fcbed98b844542eef59106eae154c34636033cfffa00e4501b87ced9a651
BLAKE2b-256 checksum
How to use checksums
da64c8622689de3bff4018c572fac905535d0ba91582eb0430442e7055e71400
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / actlens-0.2.0-py3-none-any.whl

Download URL actlens-0.2.0-py3-none-any.whl
Size 407.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
16897612583a004a6915b8aef1634244448710d574e250763178e297b4a34276
BLAKE2b-256 checksum
How to use checksums
ef9701b94f5526230b030ebfd0374c4ed9a8b08b6af26c953ffb6135054b2bfb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page