This repository provides tools to generate a linear algebra dataset and code to train an open-source pre-trained model. Our goal is to explore the model's potential for emergent reasoning, inspired by the Deepseek-R1 paper.
Project description
Linalg-Zero
Check out the poster here and the paper here.
Table of Contents
Overview
This repository offers tools for generating a linear algebra problem dataset and training an open-source base model (i.e. Qwen2.5-3B), aiming to explore planning and tool use using SFT and RL, distinct from Deepseek-R1's primary emphasis on reasoning.
The project is simple by design and mostly consists of:
linalg_zero/: contains the scripts to train models as well as generate synthetic data:generate.py: generates the linear algebra dataset and splits.distillation.py: runs the distillation pipeline to create multi-turn tool-use data.sft_train.py: performs a simple SFT of a model on a dataset.grpo_train.py: trains a model with GRPO on a given dataset.
Makefile: contains easy-to-run commands for the dataset and training workflows using previous scripts.
Main Phases
We use the DeepSeek-R1 tech report as a loose guide, but the project phases are:
- Step 1: generate a linear algebra dataset with controlled difficulty and tool-call metadata.
- Step 2: distill multi-turn tool-use data from a teacher model.
- Step 3: SFT the base model on the dataset to teach the tool-calling format.
- Step 4: GRPO fine-tune on the tool-use tasks, using a curriculum.
Installation
We use uv as the dependency management tool.
First, to install uv, follow the UV Installation Guide.
To run the experiments install the dependencies using:
- For generation/distillation:
make install-data-gen - For SFT:
make install-sft - For RL:
make install-grpo
Next, log into your Hugging Face and Weights and Biases accounts as follows:
huggingface-cli login
wandb login
Quickstart
After installing dependencies above, run the commands below. For modifications, see the config files.
# Phase 1: Generate dataset
uv run python linalg_zero/generate.py --dataset_name atomwalk12/linalgzero --push_dataset
# Phase 2: Distillation (setup once)
cp linalg_zero/config/distillation/env.example.sh env.sh
# Edit env.sh to set HF_TOKEN and ARGILLA_API_KEY.
source env.sh
# Terminal A
uv run python linalg_zero/distillation/launch_server.py --config linalg_zero/config/distillation/vllm_qwen3_32b.yaml
# Terminal B (new terminal; source env.sh again)
source env.sh
uv run python linalg_zero/distillation.py --config linalg_zero/config/distillation/vllm_qwen3_32b.yaml
# Phase 3: SFT
uv run python linalg_zero/sft_train.py --config linalg_zero/config/sft/qwen2.5-3B/lora.yaml
# Phase 4: GRPO
uv run python linalg_zero/grpo_train.py --config-name runpod.yaml
Training requires the dataset to follow the strict OpenAI tool-calling format (see this link). We provide scripts to prepare and validate the data accordingly:
linalg_zero/sft/scripts/prepare_dataset.py: prepares the SFT dataset.grpo/scripts/prepare_dataset.py: prepares and validates the GRPO dataset.
Results
We provide a recipe to encourage planning and tool-use capabilities in the Qwen2.5-3B model, starting from a pre-trained (not instruction-tuned) base model.
This yields models like Linalg-Zero-SFT and Linalg-Zero-GRPO, with the following downstream performance on the test set:
| Metric | LinAlgZero-SFT | LinAlgZero-GRPO |
|---|---|---|
| Optimal Trajectory | 89.87% | 90.26% |
| Correctness | 91.86% | 92.63% |
| Format Validity | 96.15% | 96.66% |
| Tool Success | 100.00% | 100.00% |
Artifacts
| Artifact | Link |
|---|---|
| SFT checkpoint | atomwalk12/LinalgZero-SFT |
| GRPO checkpoint | atomwalk12/LinAlgZero-GRPO |
| Base dataset | atomwalk12/linalgzero |
| Distilled dataset (clean) | atomwalk12/linalgzero-distilled-clean |
| SFT dataset | atomwalk12/linalgzero-sft |
| GRPO dataset | atomwalk12/linalgzero-grpo |
Reproducibility
- Distillation: H100 80GB on Runpod with Qwen/Qwen3-32B-FP8; 14 hours at $2.39/hr (~$25).
- SFT: Local 24GB RTX 4090 with Qwen/Qwen2.5-3B.
- GRPO: RTX 6000 Ada on Runpod, improving on the SFT baseline; 57 hours at $0.77/hr (~$50).
- Total: ~$75 using a mix of cloud GPUs and local training.
Citation
If you find this project is useful in your own work, please consider citing as follows:
@misc{openr1,
title = {Linalg-Zero: Distilling Neurosymbolic Reasoning for Linear Algebra in Small Language Models},
url = {https://github.com/atomwalk12/linalg-zero},
author = {{Razvan F. Vasile}},
month = {March},
year = {2026}
}
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file linalg_zero-1.0.0.tar.gz.
File metadata
- Download URL: linalg_zero-1.0.0.tar.gz
- Upload date:
- Size: 6.5 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.6.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
feccdd33e1d97044a9273404882d95f1c2588d2b2eadbb09c4b329eeb9884d74
|
|
| MD5 |
c95e233d5b1f9d8fa32440c2b2daefc2
|
|
| BLAKE2b-256 |
1619702658f1f4640ec8043b6e8f4743062ea24fac9533d44efe10920c8a086c
|
File details
Details for the file linalg_zero-1.0.0-py3-none-any.whl.
File metadata
- Download URL: linalg_zero-1.0.0-py3-none-any.whl
- Upload date:
- Size: 666.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.6.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
54f7a57ec560dff505ec1f55cc4b0220d664b8c216222a56f09b5756087bf752
|
|
| MD5 |
f3700562b8a43c17c77c7939f0b7b328
|
|
| BLAKE2b-256 |
3ea0b0a2fa908d1c736807c55c069dcc8212e0d8c39505d961ef15e0c6a2b5e2
|