Skip to main content
Datafast Logo

Generate text datasets for LLMs in minutes, not weeks.

Intended use cases

  • Get initial evaluation text data instead of starting your LLM project blind.
  • Increase diversity and coverage of an existing dataset by generating more data.
  • Experiment and test quickly LLM-based application PoCs.
  • Make your own datasets to fine-tune and evaluate language models for your application.

🌟 Star this repo if you find this useful!

Supported Dataset Types

  • ✅ Text Classification Dataset
  • ✅ Raw Text Generation Dataset
  • ✅ Instruction Dataset (Ultrachat-like)
  • ✅ Multiple Choice Question (MCQ) Dataset
  • ✅ Preference Dataset
  • ⏳ more to come...

Supported LLM Providers

Currently we support the following LLM providers:

  • ✔︎ OpenAI
  • ✔︎ Anthropic
  • ✔︎ Google Gemini
  • ✔︎ Ollama (local LLM server)
  • ✔︎ Mistral AI
  • ⏳ more to come...

Try it in Colab:

Open In Colab

Installation

pip install datafast

Quick Start

1. Environment Setup

Make sure you have created a .env file with your API keys. HF token is needed if you want to push the dataset to your HF hub. Other keys depends on which LLM providers you use.

GEMINI_API_KEY=XXXX
OPENAI_API_KEY=sk-XXXX
ANTHROPIC_API_KEY=sk-ant-XXXXX
MISTRAL_API_KEY=XXXX
HF_TOKEN=hf_XXXXX

2. Import Dependencies

from datafast.datasets import ClassificationDataset
from datafast.schema.config import ClassificationDatasetConfig, PromptExpansionConfig
from datafast.llms import OpenAIProvider, AnthropicProvider, GeminiProvider
from dotenv import load_dotenv

# Load environment variables
load_dotenv() # <--- your API keys

3. Configure Dataset

# Configure the dataset for text classification
config = ClassificationDatasetConfig(
    classes=[
        {"name": "positive", "description": "Text expressing positive emotions or approval"},
        {"name": "negative", "description": "Text expressing negative emotions or criticism"}
    ],
    num_samples_per_prompt=5,
    output_file="outdoor_activities_sentiments.jsonl",
    languages={
        "en": "English", 
        "fr": "French"
    },
    prompts=[
        (
            "Generate {num_samples} reviews in {language_name} which are diverse "
            "and representative of a '{label_name}' sentiment class. "
            "{label_description}. The reviews should be {{style}} and in the "
            "context of {{context}}."
        )
    ],
    expansion=PromptExpansionConfig(
        placeholders={
            "context": ["hike review", "speedboat tour review", "outdoor climbing experience"],
            "style": ["brief", "detailed"]
        },
        combinatorial=True
    )
)

4. Setup LLM Providers

# Create LLM providers
providers = [
    OpenAIProvider(model_id="gpt-5-mini-2025-08-07"),
    AnthropicProvider(model_id="claude-haiku-4-5-20251001"),
    GeminiProvider(model_id="gemini-2.0-flash")
]

5. Generate and Push Dataset

# Generate dataset and local save
dataset = ClassificationDataset(config)
dataset.generate(providers)

# Optional: Push to Hugging Face Hub
dataset.push_to_hub(
    repo_id="YOUR_USERNAME/YOUR_DATASET_NAME",
    train_size=0.6
)

Next Steps

Check out our guides for different dataset types:

Key Features

  • Easy-to-use and simple interface 🚀
  • Multi-lingual datasets generation 🌍
  • Multiple LLMs used to boost dataset diversity 🤖
  • Flexible prompt: use our default prompts or provide your own custom prompts 📝
  • Prompt expansion: Combinatorial variation of prompts to maximize diversity 🔄
  • Hugging Face Integration: Push generated datasets to the Hub 🤗

Contributing

Contributions are welcome! If you are new to the project, pick an issue labelled "good first issue".

How to proceed?

  1. Pick an issue
  2. Comment on the issue to let others know you are working on it
  3. Fork the repository
  4. Clone your fork locally
  5. Create a new branch and give it a name like feature/my-awsome-feature
  6. Make your changes
  7. If you feel like it, write a few tests for your changes
  8. To run the current tests, you can run pytest in the root directory. Don't pay attention to UserWarning: Pydantic serializer warnings. Note that for the LLMs test to run successfully you'll need to have:
  • openai API key
  • anthropic API key
  • gemini API key
  • mistral API key
  • an ollama server running (use ollama serve from command line)
  1. Commit your change, push to your fork and create a pull request from your fork branch to datafast main branch.
  2. Explain your pull request in a clear and concise way, I'll review it as soon as possible.

Roadmap:

  • RAG datasets
  • Personas
  • Seeds
  • More types of instructions datasets (not just ultrachat)
  • More LLM providers
  • Deduplication, filtering
  • Dataset cards generation

Creator

Made with ❤️ by Patrick Fleith.


This is volunteer work, star this repo to show your support! 🙏

Project Details

  • Status: Work in Progress (APIs may change)
  • License: Apache 2.0

Metadata

Release files for datafast 0.0.36

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datafast 0.0.36
File Size Uploaded
datafast-0.0.36.tar.gz 72.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datafast 0.0.36
File Interpreter ABI Platform
datafast-0.0.36-py3-none-any.whl Python 3 none any Details

Total release size: 144.0 kB

Release files / datafast-0.0.36.tar.gz

Download URL datafast-0.0.36.tar.gz
Size 72.7 kB
Tags Source
SHA-256 checksum
How to use checksums
46e6567754d7c0de50501d1d7636ac57a21f6da8964e9d5c660200a36a20672c
BLAKE2b-256 checksum
How to use checksums
b25086331d3a789f2741d447ce8b761d66af24ffa82ffcd1853737ff0a8d555f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / datafast-0.0.36-py3-none-any.whl

Download URL datafast-0.0.36-py3-none-any.whl
Size 71.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ece6ed61d602026c8d29849ca9b8cb2804e309cc043d7dda0686efff264c2c9d
BLAKE2b-256 checksum
How to use checksums
3abd130e2e3ed1b5239405a0a7b11e2396fb89293dbfb90385cb9c08097651fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.0.36 This release

2 release files

0.0.34

2 release files

0.0.33

2 release files

0.0.32

2 release files

0.0.29

2 release files

0.0.28

2 release files

0.0.27

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.24

2 release files

0.0.23

2 release files

0.0.22

2 release files

0.0.21

2 release files

0.0.19

2 release files

0.0.18

2 release files

0.0.17

2 release files

0.0.16

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page