Skip to main content

nanoscope

See what your language model learns. You write an nn.Module; nanoscope handles the data, the training loop, evaluation, checkpoints and comparison against baselines.

It has two levels that share one core. Learners call run(). Researchers write a Study with seeds, budgets, parameter matching and preregistration. The level changes what you see, never which code runs.

Learn

pip install git+https://github.com/almajd3713/nanoscope
import torch.nn as nn
from nanoscope import compare, run

class Bigram(nn.Module):
    def __init__(self, vocab_size: int, d_model: int = 32):
        super().__init__()
        self.token_embedding = nn.Embedding(vocab_size, d_model)
        self.head = nn.Linear(d_model, vocab_size, bias=False)

    def forward(self, idx):
        return self.head(self.token_embedding(idx))

result = run(Bigram, preset="tinystories-5min")   # about a minute on a laptop CPU
result.plot()
compare(result, "gpt2")                          # against a shipped 3-seed baseline

Work through the notebooks in order:

  1. 01-first-model: write a bigram model and train it.
  2. 02-gpt2: a real transformer.
  3. 03-modern-block: RoPE, RMSNorm, SwiGLU, GQA, QK-norm, z-loss.
  4. 04-ablations: which part matters, with seeds and confidence intervals.

Token data downloads from the Hub (RedhouaneLazib/nanoscope-tokens) when a preset has it, and is tokenized locally otherwise. Everything is cached under ~/.nanoscope/data.

Research

See the research guide. The research program itself is in docs/project-nanoscope.md.

Build models from blocks

GPT2 and Modern are short compositions of the blocks in nanoscope.blocks, and so can your own models. See docs/blocks.md.

Learn

Guided paths build the models step by step, with checks that say why. See docs/learn.md: nanoscope learn list, learn start, learn check.

Command line

nanoscope presets                                   # available presets
nanoscope run nanoscope/models/gpt2.py:GPT2 --seeds 3
nanoscope compare modern gpt2 --preset tinystories-5min
nanoscope study studies/m1_ablation.py --devices cuda:0
nanoscope report studies/m1_ablation.py             # writes experiments/<name>/
nanoscope bench modern --compile reduce-overhead    # speed, and whether CPU or GPU is the limit
nanoscope status runs                               # what is running and how far along
nanoscope describe nanoscope/models/modern.py:Modern  # shapes, params, FLOPs, memory per module
nanoscope graph my_model.py                         # a model file's architecture, without running it
nanoscope blocks                                    # the blocks models are composed from
nanoscope prepare-data tinystories-5min             # download or tokenize now
nanoscope publish-data tinystories-5min user/repo   # upload tokens to a Hub dataset

Develop

uv sync --all-extras
make test       # offline tests
make lint
make typecheck
make check      # lint, typecheck, then test: what CI runs

Data credits

Tokens are derived from TinyStories (CDLA-Sharing-1.0) and FineWeb-Edu (ODC-By 1.0).

Metadata

Release files for nanoscope-lab 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nanoscope-lab 0.3.0
File Size Uploaded
nanoscope_lab-0.3.0.tar.gz 429.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nanoscope-lab 0.3.0
File Interpreter ABI Platform
nanoscope_lab-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 732.5 kB

Release files / nanoscope_lab-0.3.0.tar.gz

Download URL nanoscope_lab-0.3.0.tar.gz
Size 429.4 kB
Tags Source
SHA-256 checksum
How to use checksums
3d18e17ca9f0d362c23d0471d0ef9e61bfb7c6a348155b3af1ec98c9f77f8c4a
BLAKE2b-256 checksum
How to use checksums
e5a0e2b11b9fce9c5e1021e663f97d704d100efefdae62642d2e400652e30a7c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / nanoscope_lab-0.3.0-py3-none-any.whl

Download URL nanoscope_lab-0.3.0-py3-none-any.whl
Size 303.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
51de7fdd255897e2a35d4b0eaf67c2b38bbdda96e52630c535405209076a852f
BLAKE2b-256 checksum
How to use checksums
4ce957ac0bc0d86e204b3c00b0661996296047df8b0011d374d85a38820229bb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page