Data Simulator
data-simulator is a lightweight Python library for generating synthetic datasets from your own corpus — perfect for testing, evaluating, or fine-tuning LLM Applications.
Motivation
Real documents contain a mix of useful and irrelevant content. When generating synthetic data, this leads to:
- Queries that real users would never ask
- Test sets that don't reflect actual usage
- Wasted effort optimizing for the wrong things
Data Simulator filters out low-quality content first, then generates realistic queries and answers that match how your system will actually be used.
Getting Started
Install from PyPI:
pip install llm-data-simulator
Or install it locally:
git clone https://github.com/langwatch/data-simulator.git
cd data-simulator
pip install -e .
Run the built-in test script:
python test.py
Example test.py
from data_simulator import DataSimulator
from dotenv import load_dotenv
import os
from data_simulator.utils import display_results
load_dotenv()
generator = DataSimulator(api_key=os.getenv("OPENAI_API_KEY"))
results = generator.generate_from_docs(
file_paths=["test_data/nike_10k.pdf"],
context="You're a financial support assistant for Nike, helping a financial analyst decide whether to invest in the stock.",
example_queries="how much revenue did nike make last year\nwhat risks does nike face\nwhat are nike's top 3 priorities"
)
display_results(results)
Output Format
{
"id": "chunk_42",
"document": "Nike reported annual revenue of $44.5 billion for fiscal year 2022, an increase of 5% compared to the previous year.",
"query": "What was Nike's revenue growth in 2022?",
"answer": "Nike's revenue grew by 5% in fiscal year 2022, reaching $44.5 billion."
}
Project Structure
The project follows a modular, object-oriented design:
simulator.py: Contains the mainDataSimulatorclass that orchestrates the data generation processllm.py: Houses theLLMProcessorclass that handles all LLM-related operationsdocument_processor.py: Provides theDocumentProcessorclass for loading and chunking documentsprompts.py: Stores all prompt templates used for LLM interactionsutils.py: Contains utility functions likedisplay_resultsfor formatting output
License
MIT License
Metadata
Release files for llm-data-simulator 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_data_simulator-0.1.0.tar.gz | 9.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_data_simulator-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.4 kB
Release files / llm_data_simulator-0.1.0.tar.gz
| Download URL | llm_data_simulator-0.1.0.tar.gz |
|---|---|
| Size | 9.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ebb5f0ec410195202521f535f1d3a59f7c0735131977fd0aa27f5684816dd336
|
|
BLAKE2b-256 checksum How to use checksums |
28bf2340780ce43bb1584f1fd7fa3a8726c69424d254df7eaa03a0e15f83a712
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.3
|
Release files / llm_data_simulator-0.1.0-py3-none-any.whl
| Download URL | llm_data_simulator-0.1.0-py3-none-any.whl |
|---|---|
| Size | 9.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6f45aa81a2c7cbe0a7bbe3235079845fe6fecfaa411ee365cd6e300a4b45dc22
|
|
BLAKE2b-256 checksum How to use checksums |
ce2355355c865ffc1a1df3c416d7ed9bcbb3a0fd1fde122d9ef91e1fa14c7995
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.3
|