A minimal, educational deep-learning framework built from scratch on top of NumPy
Project description
tiny-torch
A minimal, educational deep-learning framework built from scratch on top of NumPy.
tiny-torch reimplements the essential pieces of a PyTorch-style workflow — a
tensor with reverse-mode automatic differentiation, a small set of layers,
activation and loss functions, and a data-loading pipeline — in a few hundred
lines of readable Python.
The goal is not performance but clarity: every gradient is computed by hand in an explicit backward class, so you can read exactly how backpropagation flows through the computation graph.
Requirements
- Python >= 3.14
- NumPy >= 2.5.0
- Graphviz >= 0.21
Optional (dev): ipykernel, matplotlib (used by the examples).
Installation
The project uses uv for environment and
dependency management (uv.lock is committed).
uv sync
This creates a .venv and installs the runtime and dev dependencies. Prebuilt
artifacts for tiny_torch-0.1.0 are also available under dist/.
Architecture
The whole library lives under core/ and is organized in a small number of
single-responsibility modules:
core/
├── tensor.py # Tensor: the numpy-backed frontend + operator overloading
├── functions.py # Pure numpy math: activations, softmax, loss functions
├── activations.py # Activation layers (ReLU, Sigmoid, Tanh, GELU, Softmax)
├── losses.py # Loss objects (MSE, CrossEntropy, BinaryCrossEntropy)
├── layers.py # Layer, Linear, Dropout, Sequential
├── optimizer.py # Optimizer, SGD, SGDM, Adam, AdamW
├── graph.py # ComputationalGraph: graphviz visualisation of a Sequential model
├── utils.py # unbroadcast() helper for gradient reduction
├── autograd/ # Reverse-mode automatic differentiation
│ ├── base.py # Function: base class for every backward node
│ ├── arithmetic.py # Add/Sub/Mul/Div/Matmul/Sum/Reshape/Transpose backward
│ ├── activations.py # ReLU/Sigmoid/Tanh/GELU/Softmax backward
│ └── losses.py # MSE/CrossEntropy/BCE backward
├── dataset/ # Data loading pipeline
│ ├── dataset.py # Dataset, TensorDataset, ImageDataset, DataLoader
│ ├── transformation.py # RandomHorizontalFlip, RandomCrop, Compose
│ └── utils.py # image loading helpers
└── training/ # Training loop orchestration
├── trainer.py # Trainer: train_epoch/eval, checkpointing, grad clipping
└── schedulers.py # Schedule, CosineSchedule
The design follows a clear frontend / backend split:
Tensoris the frontend. It wraps a NumPy array, overloads the Python operators (+,-,*,/,@, …) and records the operation that produced it in a_grad_fnattribute.- The
autogradpackage is the backend. Every operation has a matching*Backwardclass (aFunction) that knows how to turn an upstream gradient into the gradients of its inputs.
Layers, activations and losses are thin objects that call into functions.py for
the forward pass and attach the corresponding Function for the backward pass.
Automatic differentiation (core/autograd/)
tiny-torch implements reverse-mode autodiff by building a dynamic graph as
operations execute (define-by-run), then walking it backwards to accumulate
gradients.
The Function node
Every backward node subclasses Function (core/autograd/base.py):
class Function:
def __init__(self, *tensors):
self.saved_tensors = tensors # inputs needed by backward
self.next_functions = [t._grad_fn for t in tensors] # links to parent nodes
def apply(self, grad_output):
"""Turn the upstream gradient into gradients for each input."""
raise NotImplementedError()
saved_tensorsholds the operands captured during the forward pass.next_functionsrecords each operand's own_grad_fn, which is what turns the set of nodes into a traversable graph.apply(grad_output)implements the chain rule for that specific operation and returns one gradient per input.
How the graph is built
When you write c = a + b, Tensor.__add__ computes the numeric result with
NumPy and attaches the backward node:
out = Tensor(self.data + other.data)
out._grad_fn = AddBackward(self, other)
So each output tensor remembers how it was produced. Chaining operations
produces a graph of Function nodes rooted at the final output.
The backward pass
Tensor.backward() (core/tensor.py) drives backpropagation recursively:
- If no gradient is supplied, it seeds
1.0for a scalar output (and raises for non-scalar outputs, matching PyTorch's behaviour). - It accumulates the incoming gradient into
self.grad(gradients add up, which is what makes shared subgraphs correct). - It calls
self._grad_fn.apply(gradient)to get the input gradients, then recurses into each input tensor thatrequires_grad.
Broadcasting is handled by unbroadcast() (core/utils.py), which sums a
gradient back down to the shape of the original operand so that broadcasted
operations (e.g. adding a bias vector to a batch) produce correctly-shaped
gradients.
Managing the graph
Tensor.zero_grad()resets a tensor's accumulated gradient.Tensor.destroy_graph()walks the graph and drops every_grad_fn, freeing the saved tensors so the graph can be garbage-collected between iterations.
Supported backward operations
| Category | Backward classes |
|---|---|
| Arithmetic | AddBackward, SubBackward, MulBackward, DivBackward |
| Linear alg. | MatmulBackward, TransposeBackward |
| Reductions | SumBackward |
| Shape | ReshapeBackward |
| Activations | ReLUBackward, SigmoidBackward, TanhBackward, GELUBackward, SoftmaxBackward |
| Losses | MSELossBackward, CrossEntropyLossBackward, BCELossBackward |
The Tensor class (core/tensor.py)
Tensor is a lightweight wrapper around a np.ndarray (always stored as
float32). It exposes:
- Metadata:
data,shape,size,dtype,requires_grad,grad,_grad_fn. - Operator overloading:
__add__/__radd__,__sub__/__rsub__,__mul__/__rmul__,__truediv__,__matmul__,__pow__,__neg__,__gt__. The autograd-aware operations (+,-,*,/,@) attach a_grad_fn; scalar/ndarrayfast paths return plain results. - Tensor ops:
matmul,reshape(supports-1inference),transpose,sum,mean,max,min. - Autograd control:
backward(),zero_grad(),destroy_graph(). - Interop:
numpy()returns the underlying array.
A convenience path in __init__ lets you build a batched tensor from a list of
tensors — Tensor([t1, t2, ...]) stacks their data automatically.
Note that core.tensor imports the backward classes at the bottom of the file,
after Tensor is defined, to break the circular import between the tensor
frontend and the autograd backend (the backward classes need Tensor at runtime).
Layers (core/layers.py)
All layers derive from the abstract Layer base class, which defines forward(),
makes instances callable, and exposes a parameters property.
| Layer | Description |
|---|---|
Linear |
Fully-connected layer y = xW + b with Xavier weight initialization and optional bias. |
Dropout |
Inverted dropout with keep-probability scaling; a no-op when training=False. |
Sequential |
Chains layers and forwards through them in order; aggregates their parameters. |
Sequential.save_graph(path, arch=True, forward=False, backward=False) renders a
.png of the model via core/graph.py (needs graphviz): a cluster per layer
for the architecture, and — if requested — the forward/backward computational
graphs built from a synthetic input, tensors colour-coded by role
(input/weights/bias/hidden).
Activation functions (core/activations.py)
Each activation is available both as a pure NumPy function (core/functions.py)
and as an autograd-aware Layer:
| Activation | Notes |
|---|---|
ReLU |
max(0, x) |
Sigmoid |
Numerically stable (branch on the sign of the input) |
Tanh |
np.tanh |
GELU |
Sigmoid approximation x · σ(1.702·x) |
Softmax |
Max-shifted for stability; configurable dim |
functions.py also provides a stable log_softmax, used internally by the
cross-entropy loss.
Loss functions (core/losses.py)
| Loss | Input | Notes |
|---|---|---|
MSELoss |
predictions, targets | Mean squared error. |
CrossEntropyLoss |
logits, integer targets | Combines a stable log_softmax with negative log-likelihood; the backward is the classic softmax(logits) − onehot(targets). |
BinaryCrossEntropyLoss |
probabilities, targets | Clips predictions to [1e-7, 1 − 1e-7] to avoid log(0). |
Each loss is callable (loss(pred, target)) and returns a scalar Tensor you can
call .backward() on.
Optimizers (core/optimizer.py)
| Optimizer | Notes |
|---|---|
SGD |
Plain gradient descent with optional L2 weight decay. |
SGDM |
SGD with momentum. |
Adam |
Adaptive moments with bias correction. |
AdamW |
Adam with decoupled weight decay. |
Every optimizer takes model.parameters and a learning rate; step() updates
param.data in place, zero_grad() clears param.grad, and get_state()
returns the optimizer's hyperparameters/buffers for checkpointing.
Training loop (core/training/)
Trainer(trainer.py) wraps a model, loss, optimizer and optional scheduler.train_epoch(dataloader, accumulation_steps=1)runs one epoch (with gradient accumulation and optionalclip_grad_normclipping) andeval(dataloader)runs a no-grad pass, both logging intotrainer.history(train_loss,eval_loss,lr).save()/load()(de)serialize training state to a checkpoint file viapickle.Schedule(schedulers.py) is the abstract base for learning-rate schedules;CosineSchedule(max_lr, min_lr, total_epochs)anneals the learning rate frommax_lrtomin_lrfollowing a cosine curve, applied byTrainerat the end of everytrain_epoch()call.
Data loading (core/dataset/)
The module mirrors the PyTorch Dataset / DataLoader pattern.
Dataset— abstract base defining__len__and__getitem__.TensorDataset— wraps in-memory tensors and validates that they share the same length along dimension 0.ImageDataset— lazily loads images from disk on access (viaload_jpeg), pairing each with its label.DataLoader— iterates aDatasetin mini-batches, with optional shuffling, and collates each batch by stacking samples along a new leading (batch) axis.
Data augmentation transforms live in transformation.py:
RandomHorizontalFlip(p)— flips along the width axis with probabilityp.RandomCrop(height, width, padding)— zero-pads then crops a random window.Compose([...])— chains transforms into a single callable.
Examples
See examples/linear_regression/ for a family
of end-to-end regression scripts and notebooks built on Trainer +
DataLoader:
| Example | Description |
|---|---|
linear/ |
Recovers the slope/intercept of 2·x + 5 with a single Linear(1, 1) layer; compares the learned fit against the closed-form lstsq solution. |
quadratic/ |
Recovers the coefficients of x² + 2·x + 2 via feature expansion ([x, x²]) fed into a Linear(2, 1) layer. |
cubic/ |
Same idea one degree further: recovers 1.2·x³ − 2.3·x² + 2·x + 2 with a Linear(3, 1) layer over [x, x², x³]. |
ill-cond/ |
Notebook comparing closed-form OLS vs. SGD on ill-conditioned (near-collinear) features, showing OLS's coefficients blow up while SGD's stay stable. |
variance/ |
Notebook fitting 20 repeated Linear(1, 1) models at each of nine label-noise levels, showing via boxplots how the loss and learned slope/intercept drift and spread as noise grows. |
The linear_regression/README.md
ties the linear/quadratic/cubic scripts together and explains why the learning
rate has to shrink as the polynomial degree grows.
uv run python examples/linear_regression/linear/linear.py
Roadmap
Planned work is tracked in TODO.md, and includes:
- Loader: parallel loading via multithreading; prefetching of the next batch.
- Autograd: move backward computation entirely onto NumPy arrays (keeping
Tensoras a pure frontend); cache forward intermediates for reuse in backward; add a debug step that reports which node a backward failure occurred on.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file thorcino-0.1.1.tar.gz.
File metadata
- Download URL: thorcino-0.1.1.tar.gz
- Upload date:
- Size: 1.8 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.9 {"installer":{"name":"uv","version":"0.11.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.10","id":"oracular","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c1a8efff2e595a49e4f85a787bca69e889fa59fcb581b3722dff63a3e9a142c4
|
|
| MD5 |
5a1d0870e9f3335131602e6cba3286a7
|
|
| BLAKE2b-256 |
9eee507c0182b484a7ebf5150e5857ff0a95f8b6e68ff1f480f878120b40bc53
|
File details
Details for the file thorcino-0.1.1-py3-none-any.whl.
File metadata
- Download URL: thorcino-0.1.1-py3-none-any.whl
- Upload date:
- Size: 27.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.9 {"installer":{"name":"uv","version":"0.11.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.10","id":"oracular","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f2f2a231204739fad121972fc7381a4b470a6ca7c75ef562a9d01346096fda30
|
|
| MD5 |
a7352beff6c88c9f5adc19de3eddc700
|
|
| BLAKE2b-256 |
5cf5379abb5eb9489ce422e63ba8336251422e4c01188d11e140cbfd2a2e4acf
|