Skip to main content

This repository provides tools to generate a linear algebra dataset and code to train an open-source pre-trained model. Our goal is to explore the model's potential for emergent reasoning, inspired by the Deepseek-R1 paper.

Project description

Release Build status Commit activity License

Linalg-Zero

Check out the poster here and the paper here.

poster

Table of Contents

  1. Overview
  2. Main Phases
  3. Installation
  4. Quickstart
  5. Results
  6. Reproducibility

Overview

This repository offers tools for generating a linear algebra problem dataset and training an open-source base model (i.e. Qwen2.5-3B), aiming to explore planning and tool use using SFT and RL, distinct from Deepseek-R1's primary emphasis on reasoning.

The project is simple by design and mostly consists of:

  • linalg_zero/: contains the scripts to train models as well as generate synthetic data:
    • generate.py: generates the linear algebra dataset and splits.
    • distillation.py: runs the distillation pipeline to create multi-turn tool-use data.
    • sft_train.py: performs a simple SFT of a model on a dataset.
    • grpo_train.py: trains a model with GRPO on a given dataset.
  • Makefile: contains easy-to-run commands for the dataset and training workflows using previous scripts.

Main Phases

We use the DeepSeek-R1 tech report as a loose guide, but the project phases are:

  • Step 1: generate a linear algebra dataset with controlled difficulty and tool-call metadata.
  • Step 2: distill multi-turn tool-use data from a teacher model.
  • Step 3: SFT the base model on the dataset to teach the tool-calling format.
  • Step 4: GRPO fine-tune on the tool-use tasks, using a curriculum.

Installation

We use uv as the dependency management tool. First, to install uv, follow the UV Installation Guide.

To run the experiments install the dependencies using:

  • For generation/distillation: make install-data-gen
  • For SFT: make install-sft
  • For RL: make install-grpo

Next, log into your Hugging Face and Weights and Biases accounts as follows:

huggingface-cli login
wandb login

Quickstart

After installing dependencies above, run the commands below. For modifications, see the config files.

# Phase 1: Generate dataset
uv run python linalg_zero/generate.py --dataset_name atomwalk12/linalgzero --push_dataset

# Phase 2: Distillation (setup once)
cp linalg_zero/config/distillation/env.example.sh env.sh
# Edit env.sh to set HF_TOKEN and ARGILLA_API_KEY.
source env.sh

# Terminal A
uv run python linalg_zero/distillation/launch_server.py --config linalg_zero/config/distillation/vllm_qwen3_32b.yaml

# Terminal B (new terminal; source env.sh again)
source env.sh
uv run python linalg_zero/distillation.py --config linalg_zero/config/distillation/vllm_qwen3_32b.yaml

# Phase 3: SFT
uv run python linalg_zero/sft_train.py --config linalg_zero/config/sft/qwen2.5-3B/lora.yaml

# Phase 4: GRPO
uv run python linalg_zero/grpo_train.py --config-name runpod.yaml

Training requires the dataset to follow the strict OpenAI tool-calling format (see this link). We provide scripts to prepare and validate the data accordingly:

  • linalg_zero/
    • sft/scripts/prepare_dataset.py: prepares the SFT dataset.
    • grpo/scripts/prepare_dataset.py: prepares and validates the GRPO dataset.

Results

We provide a recipe to encourage planning and tool-use capabilities in the Qwen2.5-3B model, starting from a pre-trained (not instruction-tuned) base model.

This yields models like Linalg-Zero-SFT and Linalg-Zero-GRPO, with the following downstream performance on the test set:

Metric LinAlgZero-SFT LinAlgZero-GRPO
Optimal Trajectory 89.87% 90.26%
Correctness 91.86% 92.63%
Format Validity 96.15% 96.66%
Tool Success 100.00% 100.00%

Artifacts

Artifact Link
SFT checkpoint atomwalk12/LinalgZero-SFT
GRPO checkpoint atomwalk12/LinAlgZero-GRPO
Base dataset atomwalk12/linalgzero
Distilled dataset (clean) atomwalk12/linalgzero-distilled-clean
SFT dataset atomwalk12/linalgzero-sft
GRPO dataset atomwalk12/linalgzero-grpo

Reproducibility

  • Distillation: H100 80GB on Runpod with Qwen/Qwen3-32B-FP8; 14 hours at $2.39/hr (~$25).
  • SFT: Local 24GB RTX 4090 with Qwen/Qwen2.5-3B.
  • GRPO: RTX 6000 Ada on Runpod, improving on the SFT baseline; 57 hours at $0.77/hr (~$50).
  • Total: ~$75 using a mix of cloud GPUs and local training.

Citation

If you find this project is useful in your own work, please consider citing as follows:

@misc{openr1,
    title = {Linalg-Zero: Distilling Neurosymbolic Reasoning for Linear Algebra in Small Language Models},
    url = {https://github.com/atomwalk12/linalg-zero},
    author = {{Razvan F. Vasile}},
    month = {March},
    year = {2026}
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

linalg_zero-1.0.0.tar.gz (6.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

linalg_zero-1.0.0-py3-none-any.whl (666.1 kB view details)

Uploaded Python 3

File details

Details for the file linalg_zero-1.0.0.tar.gz.

File metadata

  • Download URL: linalg_zero-1.0.0.tar.gz
  • Upload date:
  • Size: 6.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.6.14

File hashes

Hashes for linalg_zero-1.0.0.tar.gz
Algorithm Hash digest
SHA256 feccdd33e1d97044a9273404882d95f1c2588d2b2eadbb09c4b329eeb9884d74
MD5 c95e233d5b1f9d8fa32440c2b2daefc2
BLAKE2b-256 1619702658f1f4640ec8043b6e8f4743062ea24fac9533d44efe10920c8a086c

See more details on using hashes here.

File details

Details for the file linalg_zero-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for linalg_zero-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 54f7a57ec560dff505ec1f55cc4b0220d664b8c216222a56f09b5756087bf752
MD5 f3700562b8a43c17c77c7939f0b7b328
BLAKE2b-256 3ea0b0a2fa908d1c736807c55c069dcc8212e0d8c39505d961ef15e0c6a2b5e2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page