multilingual-gsm-symbolic
A Python package for generating synthetic multilingual math problems from symbolic templates. Allows you to create more than a thousand examples from just one problem and allows you to test if the LLMs actually understand the problem or whether it was just lucky pattern-matching.
⏳ Installation
pip install multilingual-gsm-symbolic
👩💻 Get started
from multilingual_gsm_symbolic import load_data, available_languages
# see possible languages
print(available_languages())
# {'eng': {'number of samples': 100}, 'dan': {'number of samples': 100}, ...}
# Load English templates
templates = load_data("eng")
# Generate concrete questions from a template
questions = templates[0].generate_questions(n=10)
for q in questions:
print(q.question)
print(q.answer)
print()
Running experiments
You might often be interested in some sort of variation upon the dataset. E.g. does the performance degredation happens only due to the changes names:
# We can also control the synthetic generation:
# fix numeric variables and only vary names/strings
defaults = templates[0].get_default_assignments()
number_vars = {var: val for var, val in defaults.items() if not isinstance(val, str)}
questions = templates[0].generate_questions(n=5, fixed=number_vars, verbose=False)
You can also inspect the available combinations directly:
# get up to 100 unique numeric assignments for a template
combinations = templates[0].get_combinations(limit=100, only_numeric=True)
print(len(combinations))
print(combinations[:3])
You could imagine similar ablations, but adding spelling errors, introducing irrelevant task information like "Hey just a small math question: {question}" or similar.
📋 Template format
Templates are TOML files with the following fields:
| Field | Description |
|---|---|
question |
Concrete question (the original example) |
answer |
Concrete answer with calculation steps |
question_annotated |
Template with variable placeholders and #init / #conditions / #answer sections |
answer_annotated |
Answer template with inline expressions |
source-language |
Source language code; omitted for English originals |
initial_translation_model |
Initial translation model; omitted for English originals |
human-validated |
Human review description or "in progress"; omitted if unreviewed |
error-analysis |
Error analysis performed; omitted if absent |
Computational validation is enforced by CI and is not stored in templates. Absent metadata fields are omitted from TOML and load as None in Python.
Annotated question syntax
{variable, default_value} — placeholder in the question text
#init:
- $var = range(low, high) — variable sampled from a range
- $var = sample([a, b, c]) — variable sampled from a list
#conditions:
- is_int(x / y) — constraint that must hold for a combination to be valid
#answer: x * y + z — Python expression evaluated to produce the numeric answer
Example: fog bank problem
question = "A fog bank rolls in over a city at 3 miles/hour. The city is 42 miles wide. How many hours will it take for the fog bank to cover the city?"
answer = "At 3 miles/hour, it will take 42/3=14 hours for the fog to cover the city."
id_orig = 0
id_shuffled = 0
creation = "example"
language = "eng"
question_annotated = """
A fog bank rolls in over a city at {speed,3} miles/hour. The city is {width,42} miles wide. How many hours will it take for the fog bank to cover the city?
#init:
- $speed = range(1, 20)
- $width = range(2, 100)
#conditions:
- is_int(width / speed)
#answer: width // speed
"""
answer_annotated = "At {speed} miles/hour, it will take {width}/{speed}={width//speed} hours for the fog to cover the city."
Example: shopping problem
question = "A store sells apples for $2 each and oranges for $3 each. If you buy 4 apples and 5 oranges, how much do you spend?"
answer = "You spend 4*2 + 5*3 = 8 + 15 = $23."
id_orig = 0
id_shuffled = 0
creation = "example"
language = "eng"
question_annotated = """
A store sells apples for ${apple_price,2} each and oranges for ${orange_price,3} each. If you buy {n_apples,4} apples and {n_oranges,5} oranges, how much do you spend?
#init:
- $apple_price = range(1, 10)
- $orange_price = range(1, 10)
- $n_apples = range(1, 20)
- $n_oranges = range(1, 20)
#conditions:
- True
#answer: apple_price * n_apples + orange_price * n_oranges
"""
answer_annotated = "You spend {n_apples}*{apple_price} + {n_oranges}*{orange_price} = {n_apples*apple_price} + {n_oranges*orange_price} = ${apple_price*n_apples + orange_price*n_oranges}."
Writing a custom template
Writing a custom template
Here is a complete example — a "speed × time = distance" problem with randomised values and a divisibility constraint:
question = "A car travels at 60 mph for 3 hours. How far does it travel?"
answer = """
Distance = speed × time = 60 × 3 = 180 miles.
#### 180
"""
id_orig = 0
id_shuffled = 0
creation = "example"
language = "eng"
question_annotated = """
A car travels at {speed,60} mph for {hours,3} hours. How far does it travel?
#init:
- $speed = range(20, 100, 10)
- $hours = range(1, 9)
#conditions:
- is_int(speed * hours / 10)
#answer: speed * hours
"""
answer_annotated = """
Distance = speed × time = {speed} × {hours} = {speed * hours} miles.
#### {speed * hours}
"""
Save it as a .toml file and load it directly:
from multilingual_gsm_symbolic.templates import AnnotatedQuestion
template = AnnotatedQuestion.from_toml("my_template.toml")
questions = template.generate_questions(n=5)
for q in questions:
print(q.question)
print(q.answer)
Init functions available in #init lines:
| Function | Returns |
|---|---|
range(start, end[, step]) |
integers in [start, end) |
arange(start, end[, step]) |
evenly-spaced floats |
sample(items[, n]) |
one item (or n items) from a list |
sample_sequential(items, n) |
n consecutive items from a list |
range_str(start, end, step, word_list) |
(word, int) pairs, e.g. ("three", 3) |
Condition functions available in #conditions lines:
| Function | Returns |
|---|---|
is_int(x) |
True if x is a whole number |
divides(a, b) |
True if a % b == 0 |
Fraction(x) |
fraction string, e.g. "3/4" |
🗃️ Data
The English templates are derived from Apple's GSM-Symbolic paper, from which the remainder is derived. For example, the Danish templates were initially translated using GPT-5.4, localized and reviewed by native speakers, and validated computationally and using Claude Opus. The original concrete problems are from GSM8k.
You can see the available languages as follows:
from multilingual_gsm_symbolic import available_languages, load_data
# see possible languages
print(available_languages())
# {'eng': {'number of samples': 100}, 'dan': {'number of samples': 100}, ...}
# Inspect template metadata:
template = load_data("dan")[0]
template.source_language # 'eng'
template.initial_translation_model # 'gpt-5.4'
template.human_validated # 'by three native speakers'
template.error_analysis
The following table shows the list of validated languages:
| Language | Computationally validated | Human validated | Error analysis |
|---|---|---|---|
ara |
✓ | by a native speaker | ✓ |
dan |
✓ | by three native speakers | ✓ |
deu |
✓ | by two native speakers | ✓ |
eng |
✓ | by a native speaker | ✓ |
est |
✓ | by a native speaker | ✓ |
fra |
✓ | by a native speaker | ✓ |
hin |
✓ | by a native speaker | ✓ |
isl |
✓ | by native speakers | ✓ |
ita |
✓ | by two native speakers | ✓ |
jpn |
✓ | by a native speaker | ✓ |
mar |
✓ | by a native speaker | ✓ |
nld |
✓ | by a native speaker | ✓ |
pan |
✓ | by two native speakers | ✓ |
rus |
✓ | by a native speaker | ✓ |
swe |
✓ | by 5 native speakers | ✓ |
ukr |
✓ | by a native speaker | ✓ |
urd |
✓ | by a native speaker | ✓ |
zho |
✓ | by a native speaker | ✓ |
Want to add a new language?
Want to add a new language or validate an existing one? Great to hear. The src/multilingual_gsm_symbolic/data/templates/ folder contains the templates for each language, and src/scripts/translate_templates.py can be used to translate the templates from one language to another. We have already pre-generated a few languages; see the templates folder for which ones. Each language folder also contains an instruction.toml with the prompt used to evaluate models (available through load_instruction), which should be validated along with the templates. Once you have validated the examples you can submit a PR with the changes.
📖 API reference
function load_data
load_data(language="eng", directory=None) → list[AnnotatedQuestion]
Load symbolic templates.
| Argument | Type | Description |
|---|---|---|
language |
str |
Language code, e.g. "eng" (default) or "dan" |
directory |
Path | None |
Override the bundled data; load templates from this path instead |
| RETURNS | list[AnnotatedQuestion] |
The loaded templates |
function load_replacements
load_replacements(language="eng") → dict
Load language-specific named values (e.g. lists of names, places) used inside templates.
| Argument | Type | Description |
|---|---|---|
language |
str |
Language code, e.g. "eng" (default) |
| RETURNS | dict |
Mapping of replacement name → value list |
function load_instruction
load_instruction(language="eng") → str
Load the language-specific instruction used to prompt a model with a question (solve step by step and put the final answer in \boxed{}). It is meant to be followed by a blank line and the question. Raises FileNotFoundError if the language has no instruction yet.
| Argument | Type | Description |
|---|---|---|
language |
str |
Language code, e.g. "eng" (default) |
| RETURNS | str |
The instruction text |
function load_gsm
load_gsm(language="eng", directory=None) → list[GSMProblem]
Load the bundled concrete problems for a given language.
| Argument | Type | Description |
|---|---|---|
language |
str |
Language code, e.g. "eng" (default) |
directory |
Path | None |
Override the bundled data directory |
| RETURNS | list[GSMProblem] |
The loaded concrete problems |
class AnnotatedQuestion
Core class representing a symbolic template. Constructed from a TOML template file via AnnotatedQuestion.from_toml(path).
method AnnotatedQuestion.generate_questions
Generate concrete Question instances from the template.
| Argument | Type | Description |
|---|---|---|
n |
int |
Number of questions to generate |
replacements |
dict | None |
Replacement values; loaded automatically if omitted |
seed |
int | None |
Random seed for reproducibility |
fixed |
dict | None |
Variables to hold constant; only the remaining variables are sampled |
| RETURNS | list[Question] |
The generated questions |
method AnnotatedQuestion.get_default_assignments
Extract the default variable values from the question template placeholders.
| Argument | Type | Description |
|---|---|---|
| RETURNS | dict |
Mapping of variable name → default value |
method AnnotatedQuestion.get_combinations
Enumerate unique valid assignments for a template.
| Argument | Type | Description |
|---|---|---|
replacements |
dict | None |
Replacement values; loaded automatically if omitted |
only_numeric |
bool |
If True, keep only numeric variables in each returned assignment |
fixed |
dict | None |
Variables to hold constant while enumerating combinations |
limit |
int | None |
Stop after this many unique combinations |
| RETURNS | list[dict] |
Unique valid assignments for the template |
method AnnotatedQuestion.format_question
Render the question text for a given variable assignment.
| Argument | Type | Description |
|---|---|---|
assignments |
dict |
Variable name → value mapping |
language |
str |
Language code for rendered text |
| RETURNS | str |
The rendered question string |
method AnnotatedQuestion.format_answer
Render the answer text for a given variable assignment.
| Argument | Type | Description |
|---|---|---|
assignments |
dict |
Variable name → value mapping |
language |
str |
Language code for rendered text |
| RETURNS | str |
The rendered answer string |
class Question
Dataclass holding a single generated problem.
| Attribute | Type | Description |
|---|---|---|
question |
str |
The rendered question text |
answer |
str |
The rendered answer text |
id_orig |
int |
Index of the original template |
id_shuffled |
int |
Index within the shuffled sample |
Acknowledgement
If you use this dataset please cite the paper:
@misc{enevoldsen2026multilingualgsmsymbolicdeterminescapability,
title={Multilingual GSM-Symbolic: What determines capability transfer across languages?},
author={Kenneth Enevoldsen and Riley Herchert and Sofie Mosegaard and Dan Saattrup Smart and Simon Enni and Isaac Chung and Sofie Bruun and Ayush Sunil Munot and Max Müller-Eberstein and Adnan El-Assadi and Elisa Bassignana and Gianluca Barmina and Hafsteinn Einarsson and Iben Nyholm Debess and Linda Freienthal and Lukas Galke Poech and Mike Zhang and Nicolas Legrand and Vladimir Salnikov and Yevhen Kostiuk and Zafar Hussain and Sagandeep Kaur and Agnes Toftgård and Marie Mattson and Kristoffer Nielbo},
year={2026},
eprint={2610.03367},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.03367},
}
The symbolic template engine and the danish subset were originally developed as part of the m-gsm-symbolic project at the Centre for Humanities Computing by:
The initial template format was derived from Apple's GSM-Symbolic paper and the original concrete problems are from GSM8k.
Metadata
Release files for multilingual-gsm-symbolic 0.5.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| multilingual_gsm_symbolic-0.5.5.tar.gz | 6.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| multilingual_gsm_symbolic-0.5.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.5 MB
Release files / multilingual_gsm_symbolic-0.5.5.tar.gz
| Download URL | multilingual_gsm_symbolic-0.5.5.tar.gz |
|---|---|
| Size | 6.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
65f2dcfa26e85ba02d209f4f809f3a7dbd2f42d14f520f0cd2f2dc7be6ac7738
|
|
BLAKE2b-256 checksum How to use checksums |
9ffe238d69ec4e8e0d0fe44cba31d2a34e2fe53f0407ef2db308d53762f053a8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / multilingual_gsm_symbolic-0.5.5-py3-none-any.whl
| Download URL | multilingual_gsm_symbolic-0.5.5-py3-none-any.whl |
|---|---|
| Size | 12.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
89d3f52317cf06f995e42415e7c6b83ed7d78d1d480d8a2165d6ffa3db4f9227
|
|
BLAKE2b-256 checksum How to use checksums |
ba1d23db088795c71a6f15e81fe59a33eaadd7cc702df0bdac9c98c26d4a4e40
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log