Skip to main content

ultragraph

CI python license

A pure-Python (+ numpy) byte-graph that is a 1-bit (ternary) LLM.

genesis 251e6ea · themed after pocoo.vaked.dev

ultragraph architecture — micro (node/edge, 1 byte each) → meso (tree) → macro (ultra-graph)

Three levels:

level unit storage
micro node / edge 1 byte eachint8 activation / ternary weight {-1,0,+1}
meso tree a whole graph == one net/module (a Linear/MLP block)
macro ultra-edge (===) typed wiring between trees → the ultra-graph = the model

Weights are ternary (BitNet b1.58 style); activations are int8. Full-precision "master" weights live in an ad-hoc side store during training; the byte buffers are the deployed state. Training uses a straight-through estimator (STE).

Illustrations

Real outputs from a trained ternary mini-GPT — regenerate with uv run python assets/make_figures.py:

ultra-graph causal attention ternary weight bytes
architecture attention weights

Left: the model as an ultra-graph — trees wired by ultra-edges (===), with residual skips. Middle: real causal self-attention weights (lower-triangular → no peeking at the future). Right: a trained query projection's weight bytes, each ∈ {−1, 0, +1}.

Install

pip install ultragraph-1bit    # then: import ultragraph
# or from source (Python >=3.11):
uv sync

Dunder API

>> is overloaded by operand type:

import numpy as np
from ultragraph import Tree, UltraGraph, Tensor, mlp, SGD

# micro-edges inside a sparse tree
g = Tree(4, "g")
g[0] >> g[1]        # node >> node  -> micro-edge
g[2] = 7            # set a node byte
print(len(g), 2 in g, list(g))

# ultra-edges between trees
ug = UltraGraph()
a = ug.add(Tree.dense(8, 16, "a"))
b = ug.add(Tree.dense(16, 4, "b", act="none"))
a >> b              # tree >> tree -> ultra-edge (plain)
a.wire(b, "residual")

Train a tiny ternary net

ug = mlp([4, 16, 2])                 # dense ternary linear trees wired plain
opt = SGD(ug, lr=0.3, momentum=0.9)
x = Tensor(np.random.randn(32, 4).astype("float32"))
for _ in range(300):
    loss = ug.forward(x).cross_entropy(y)
    opt.zero_grad(); loss.backward(); opt.step()   # step() re-quantizes weights

See examples/char_lm.py (MLP LM), examples/transformer_lm.py (single-head attention), examples/mini_gpt.py (batched multi-head attention + RMSNorm + Adam), examples/gpt_lm.py (the whole stack: ByteTokenizerGPT → train → stream), and examples/mesh_lm.py (a Mesh of GPT experts, gradient accumulation, joint decode) for end-to-end char/byte-level ternary language models.

from ultragraph import Embedding, MultiHeadAttention, RMSNorm, linear_tree, Adam
# pre-norm transformer block over a [B, T, d_model] sequence:
#   x = x + mha(norm1(x));  x = x + ff2(ff1(norm2(x)))

A whole ternary GPT

from ultragraph import GPT

m = GPT(vocab=256, d_model=128, n_layers=4, n_heads=4, max_len=256)  # RoPE + KV-cache
logits = m(ids)                       # ids [B, T] -> logits [B, T, vocab]
out = m.generate([72, 105], n_new=64, temperature=0.8, top_k=40, top_p=0.9,
                 repetition_penalty=1.3, stop=10, seed=0)   # stop on newline byte

for tok in m.generate([72, 105], n_new=64, temperature=0.8, stream=True):
    print(tok, end=" ", flush=True)   # token-by-token
m.save("gpt.npz")                     # fp32 masters; reload onto the same architecture

# a true 1-bit-on-disk checkpoint: bit-packed ternary bytes, no fp32 masters.
m.save_deployed("gpt.q.npz")          # ~10x smaller, inference-only
deployed = GPT.load_deployed("gpt.q.npz")   # byte-exact logits, runs from the trits

The deployed checkpoint stores weights at their true ~1.6 bits/weight density (5 ternary values per byte) plus the tiny fp32 pieces (embedding, norm gains, biases). On an 858k-param model that's 3.4 MB → 334 KB, and deployed(ids) gives logits identical to the trained model — Tree.forward runs straight from the stored bytes.

A mesh of minds

from ultragraph import GPT, Mesh

experts = [GPT(vocab=256, d_model=64, n_layers=2, n_heads=4) for _ in range(4)]
mesh = Mesh(experts, vocab=256, top_k=2)     # a learned router mixes full models
logits = mesh(ids)                           # Σ_e gate(ids)_e · expert_e(ids)
text = mesh.generate([72, 105], n_new=64, temperature=0.8)   # joint KV-cached decode

Mesh lifts nn.MoE's routing to whole networks: a small ternary router reads the sequence and mixes the experts' logits per sequence (soft, or top-k). Router and every expert train together — a graph of minds, still all ternary bytes underneath.

generate decodes with a per-layer KV-cache; since activations are quantized per token, a cached step is byte-for-byte the full-forward result at that position. Positions come from RoPE (rotary embeddings) — relative, and offset-aware so they line up across cached steps.

Guide — common recipes

Goal Command
Install pip install ultragraph-1bit (extras: viz, mcp, wiki)
Train a byte-level GPT uv run python examples/gpt_lm.py
Mixture of full models uv run python examples/mesh_lm.py
1-bit Latin LLM (Anonymus) fetch_gesta.pyanonymus_lm.py
1-bit Hungarian LLM (resumable) fetch_hungarian.pyhungarian_lm.py
Enrich the corpus (no LLM) uv run --extra wiki python examples/enrich_corpus.py
History graph — curated uv run python examples/hungarian_history.py
History graph — live from Wikipedia uv run --extra wiki python examples/hungarian_history_live.py
Serve over MCP (SSE) uv run --extra mcp python mcp_server/server.py
Dev container open in VS Code → Reopen in Container

Ternary language models on real corpora

Two byte-level ternary GPTs trained end-to-end on public-domain text; both deploy to tiny bit-packed checkpoints that run from the trits alone:

  • Latinexamples/anonymus_lm.py on the Gesta Hungarorum of Anonymus (c. 1200), the Hungarian founding chronicle (~94 KB) → 196 KB checkpoint. Sample: GPT.load_deployed("examples/data/anonymus.gpt.npz")"Almus dux … dux cum patis se … terras suis …".
  • Hungarianexamples/hungarian_lm.py on ~450 KB of public-domain Hungarian literature (Arany János + prose, pulled from Project Gutenberg by fetch_hungarian.py). Training is resumable — it saves the fp32 masters + step state and continues across runs (TOTAL/STEPS env vars, periodic checkpoints), so it converges past any wall-clock cap.
from ultragraph import GPT, ByteTokenizer
tok = ByteTokenizer()
m = GPT.load_deployed("examples/data/hungarian.gpt.npz")
print(tok.decode(m.generate(tok.encode("A magyar "), n_new=90, temperature=0.8, top_p=0.9)))

Corpus toolingexamples/enrich_corpus.py is a non-LLM gatherer: it pulls grounded facts from Hungarian Wikipedia (via ultragraph.wiki), turns them into definition + Kérdés:/Válasz: lines, dedups against the corpus, and appends (idempotent, re-runnable).

Knowledge graphs — curated & live

The byte-graph data model doubles as a knowledge graph: entities are Tree nodes, relations are micro-edges, and higher-level structure is ultra-edges (===).

  • Curatedexamples/hungarian_history.py builds Hungarian history as one ultra-graph (13 eras, 59 nodes) and renders a themed timeline SVG.
  • Liveexamples/hungarian_history_live.py builds it dynamically from hu.wikipedia via ultragraph.wiki.build_wiki_graph, scraping pages + links (with an on-disk cache) into a sparse Tree.

MCP server

mcp_server/server.py exposes the library over MCP (SSE transport) with tools anonymus_generate, ultragraph_info, and tokenize_preview:

uv run --extra mcp python mcp_server/server.py   # -> http://127.0.0.1:8000/sse

BMad agents

Three BMad agents live under skills/ to help develop and use the library:

Agent Folder What it does
ByteSmith 🔬 skills/ultragraph-dev/ Byte-graph coding agent — writes autograd ops, wires trees, trains models, explains internals
CorpusCrafter 📚 skills/corpus-trainer/ Automates corpus gathering, Wikipedia enrichment, and byte-level GPT training for any language
GraphViz 📊 skills/viz-doc/ Renders SVG/PNG model visualizations and generates model card reports from checkpoints

Each is a stateless skill (SKILL.md + capability references + customize.toml) activated by describing the task.

Tasks

just test        # pytest
just test-fast   # dependency-free runner (stdlib + numpy)
just demo        # char-LM end-to-end
just viz         # render example SVGs

Layout

ultragraph/quant.py     ternary + int8 quantization, STE
ultragraph/autograd.py  numpy autograd tape; ternary_linear (STE); exp/tanh/sigmoid/gelu/silu
ultragraph/core.py      Node/Edge/Tree/UltraEdge/UltraGraph + dunder API
ultragraph/nn.py        linear_tree, mlp, Attention, MultiHeadAttention, RoPE, RMSNorm, LayerNorm, LearnedPositionalEmbedding, MoE, Dropout, Sequential
ultragraph/model.py     TransformerBlock + GPT (RoPE + KV-cache + .generate + save_deployed) + Mesh (mixture of full models)
ultragraph/optim.py     SGD + Adam (grad clip, weight decay, gradient accumulation) + CosineSchedule
ultragraph/pack.py      dense ternary bit-packing (5 values/byte, ~1.58-bit)
ultragraph/tokenize.py  byte-level tokenizer (ByteTokenizer, vocab 256)
ultragraph/vaked.py      optional vaked lowering (lower_graph, compile_vaked via vendored vakedc)
ultragraph/viz/         svg.py (pure-SVG) + mpl.py (optional matplotlib) — micro / macro / byte-heatmap
ultragraph/io.py        byte-exact save / load (optional packed weights); save_params/load_params
ultragraph/wiki.py      optional MediaWiki client + build_wiki_graph (live hu.wikipedia -> ultra-graph)
mcp_server/server.py    optional MCP server (SSE) exposing the library as tools
examples/               char/gpt/mesh/anonymus/hungarian LMs, history graphs, corpus fetch + enrich
.devcontainer/          VS Code dev container (Python 3.12 + uv + ruff)

Design spec: docs/superpowers/specs/2026-07-10-ultragraph-design.md. Graph-theory reading list (Erdős classics): docs/references.md.

Install from source

git clone https://github.com/peterlodri-sec/ultra-graph
cd ultra-graph
uv sync
just test

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ultragraph_1bit-0.15.0.tar.gz (5.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ultragraph_1bit-0.15.0-py3-none-any.whl (59.6 kB view details)

Uploaded Python 3

File details

Details for the file ultragraph_1bit-0.15.0.tar.gz.

File metadata

  • Download URL: ultragraph_1bit-0.15.0.tar.gz
  • Upload date:
  • Size: 5.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ultragraph_1bit-0.15.0.tar.gz
Algorithm Hash digest
SHA256 3e018026ce02c54e1d6aefb70e949b47eb81f82a78df85b3b4dadbdd58b44789
MD5 d3b0904a0cc4d6fae5449af1a22bc193
BLAKE2b-256 e0060d5e6e432e40d2393da1ef2c7a02ef92dd043c23a842389c718b37a286fa

See more details on using hashes here.

Provenance

The following attestation bundles were made for ultragraph_1bit-0.15.0.tar.gz:

Publisher: publish.yml on peterlodri-sec/ultra-graph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ultragraph_1bit-0.15.0-py3-none-any.whl.

File metadata

File hashes

Hashes for ultragraph_1bit-0.15.0-py3-none-any.whl
Algorithm Hash digest
SHA256 1b84777ba401ccc8ee29435f807e395afbf9dd13dc6efcea638ba4b37502edc7
MD5 89d4e4840698ed0fd714586e17c1e7e4
BLAKE2b-256 43859b1747d080658774ff79f2c908aa9d45dd83df53d43cf2828226a1d0632e

See more details on using hashes here.

Provenance

The following attestation bundles were made for ultragraph_1bit-0.15.0-py3-none-any.whl:

Publisher: publish.yml on peterlodri-sec/ultra-graph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.17.0

2 files

0.16.0

2 files

This release

0.15.0 This release

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page