Skip to main content

TQB++ Synthetic Data Generation Toolkit

This directory contains the Python script and utilities for generating synthetic Question-Answering (QA) micro-benchmarks for the Tiny QA Benchmark++ (TQB++) suite.

Overview

The generator toolkit is a core component of TQB++, enabling the creation of bespoke tiny QA datasets. It is implemented as a Python script (approximately 200 lines of core logic) that leverages the LiteLLM library for provider-agnostic calls to various Large Language Models (LLMs).

Reference: See Section 3 ("Synthetic Data Generation Toolkit") of the TQB++ paper for a detailed description.

Features

  • Provider Agnostic: Uses LiteLLM to connect to a wide range of LLM providers (e.g., OpenAI, Anthropic, Cohere, Google, and any OpenAI-compatible API).
  • Customizable Output: Users can specify parameters to tailor the generated datasets:
    • --num: Number of questions to generate.
    • --languages: Comma-separated list of language codes (e.g., en,fr,ja).
    • --categories: Topics/domains for the questions (e.g., history,math,science).
    • --difficulty: Desired difficulty level (e.g., easy,medium,hard).
    • --provider: The LLM endpoint/model to use for generation.
  • Structured Output: The generator is prompted to produce JSON formatted output adhering to the TQB schema (text, label, context, tags: {category, difficulty}).
  • Few-Shot Prompting: Includes few-shot exemplars in the prompt to guide the LLM on the desired format and content style.
  • Schema Validation: Performs basic validation of the generated JSON structure.
  • Provenance Tracking: Stamps each generated item with a SHA-256 hash for reproducibility and provenance.

How to Run

The generator script (e.g., generator.py or by invoking the package tinyqabenchmarkpp.generate) can be run from the command line.

Important Note on Temperature: When using OpenAI reasoning models for generation (e.g., openai/o4-mini), you need to use temperature=1.0 as the script has no logic to handle this. Default is 0 to encourage reproduceability as detailed in the TQB++ paper (Appendix A.1).

Here are conceptual examples:

1. Using a specific OpenAI model:

python -m tinyqabenchmarkpp.generate \
    --num 10 \
    --languages "en" \
    --categories "science" \
    --difficulty "medium" \
    --provider "openai/gpt-3.5-turbo-0125" \
    --output_dir "./data/packs/science_en_10.jsonl" 
    # --temperature 1.0 # Explicitly set if needed, often a default

2. Using an OpenRouter model:

LiteLLM allows you to use models hosted on OpenRouter. You'll need to set your OPENROUTER_API_KEY environment variable.

# Ensure OPENROUTER_API_KEY is set in your environment
export OPENROUTER_API_KEY="your_openrouter_key_here"

python -m tinyqabenchmarkpp.generate \
    --num 15 \
    --languages "de" \
    --categories "history" \
    --difficulty "easy" \
    --provider "openrouter/google/gemma-7b-it" \
    --output_dir "./data/packs/history_de_15.jsonl"

3. Using a local Ollama model:

To use a model served locally via Ollama, ensure your Ollama server is running and the desired model is pulled (e.g., ollama pull llama3).

python -m tinyqabenchmarkpp.generate \
    --num 5 \
    --languages "es" \
    --categories "literature" \
    --difficulty "hard" \
    --provider "ollama/llama3" \
    --output_dir "./data/packs/literature_es_5_hard.jsonl" \
    --base_url "http://localhost:11434" # Specify your Ollama API base URL

(Note: Actual script name, package invocation, and parameters might vary slightly. Refer to the script's help message (python -m tinyqabenchmarkpp.generate --help) for precise usage and all available options, including how to pass API keys if not using environment variables.)

Generation Process

  1. System Prompt: Instructs the LLM to output structured JSON according to the TQB schema.
  2. Few-Shot Exemplars: Provides 2 examples to the LLM.
  3. User Prompt: Specifies the number, language, category, and difficulty of questions required.
  4. LLM Call: Sends the request to the chosen LLM via LiteLLM.
  5. Parsing & Validation: Parses the LLM response, validates the JSON structure, and includes a retry mechanism (up to 3 attempts) for malformed outputs.
  6. Hashing & Storage: Stores a SHA-256 hash of each item and saves the output.

License

This toolkit is part of the TQB++ project and is licensed under the Apache License 2.0. See the main LICENSE file in the root of the repository for details.

Metadata

Release files for tinyqabenchmarkpp 1.2.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tinyqabenchmarkpp 1.2.3
File Size Uploaded
tinyqabenchmarkpp-1.2.3.tar.gz 13.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tinyqabenchmarkpp 1.2.3
File Interpreter ABI Platform
tinyqabenchmarkpp-1.2.3-py3-none-any.whl Python 3 none any Details

Total release size: 26.9 kB

Release files / tinyqabenchmarkpp-1.2.3.tar.gz

Download URL tinyqabenchmarkpp-1.2.3.tar.gz
Size 13.9 kB
Tags Source
SHA-256 checksum
How to use checksums
33e21a9d6492c29472d04fdabd63c53e0a29e4f88c56a4e1969d9097e618dfa8
BLAKE2b-256 checksum
How to use checksums
3bfc5b8e9f5d88d0e9f37eb0f934ed578e860e66090a438896ea3433e5ae7b09
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 19, 2025.

Transparency log

Release files / tinyqabenchmarkpp-1.2.3-py3-none-any.whl

Download URL tinyqabenchmarkpp-1.2.3-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
de93323945dcb823838035789f870861a575c5a076690d69fb0499a42c22d97e
BLAKE2b-256 checksum
How to use checksums
1b165e13b531097fb14108b6ce6c11956138b923447496979b42232352da35b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 19, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

1.2.3 This release

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page