General framework for synthetic data generation

These details have not been verified by PyPI

Project description

🎨 NeMo Data Designer

Tokens

Generate high-quality synthetic datasets from scratch or using your own seed data.

Welcome!

Data Designer helps you create synthetic datasets that go beyond simple LLM prompting. Whether you need diverse statistical distributions, meaningful correlations between fields, or validated high-quality outputs, Data Designer provides a flexible framework for building production-grade synthetic data.

What can you do with Data Designer?

Generate diverse data using statistical samplers, LLMs, or existing seed datasets
Control relationships between fields with dependency-aware generation
Validate quality with built-in Python, SQL, and custom local and remote validators
Score outputs using LLM-as-a-judge for quality assessment
Iterate quickly with preview mode before full-scale generation

📣 Heads-up: async engine

Data Designer now runs pipelines on a cell-level async engine that overlaps independent columns and adapts concurrency per (provider, model). On most pipelines this is faster with no config changes; on slow self-hosted endpoints, set inference_parameters.timeout to your real per-request latency. See Architecture & Performance → Async Engine for the behaviors worth knowing about.

If you hit anything unexpected, please open an issue.

Quick Start

1. Install

pip install data-designer

Or install from source:

git clone https://github.com/NVIDIA-NeMo/DataDesigner.git
cd DataDesigner
make install

2. Set your API key

Start with one of our default model providers:

Grab your API key(s) using the above links and set one or more of the following environment variables:

export NVIDIA_API_KEY="your-api-key-here"

export OPENAI_API_KEY="your-openai-api-key-here"

export OPENROUTER_API_KEY="your-openrouter-api-key-here"

3. Start generating data!

import data_designer.config as dd
from data_designer.interface import DataDesigner

# Initialize with default settings
data_designer = DataDesigner()
config_builder = dd.DataDesignerConfigBuilder()

# Add a product category
config_builder.add_column(
    dd.SamplerColumnConfig(
        name="product_category",
        sampler_type=dd.SamplerType.CATEGORY,
        params=dd.CategorySamplerParams(
            values=["Electronics", "Clothing", "Home & Kitchen", "Books"],
        ),
    )
)

# Generate personalized customer reviews
config_builder.add_column(
    dd.LLMTextColumnConfig(
        name="review",
        model_alias="nvidia-text",
        prompt="Write a brief product review for a {{ product_category }} item you recently purchased.",
    )
)

# Preview your dataset
preview = data_designer.preview(config_builder=config_builder)
preview.display_sample_record()

What's next?

📚 Learn more

Getting Started – Install, configure, and generate your first dataset
Tutorial Notebooks – Step-by-step interactive tutorials
Column Types – Explore samplers, LLM columns, validators, and more
Validators – Learn how to validate generated data with Python, SQL, and remote validators
Model Configuration – Configure custom models and providers
Person Sampling – Learn how to sample realistic person data with demographic attributes

📝 Documentation

Data Designer documentation now lives on Fern at docs.nvidia.com/nemo/datadesigner.

Contributors should edit docs prose under fern/. Tutorial notebook source remains in docs/notebook_source/*.py; generated notebooks and Fern artifacts are not the source of truth. The legacy MkDocs archive remains available on GitHub Pages for releases 0.5.7 and older.

🔧 Configure models via CLI

data-designer config providers # Configure model providers
data-designer config models    # Set up your model configurations
data-designer config list      # View current settings

🤖 Agent Skill

Data Designer has a skill for coding agents. Just describe the dataset you want, and your agent handles schema design, validation, and generation. While the skill should work with other coding agents that support skills, our development and testing has focused on Claude Code at this stage.

Install via skills.sh (be sure to select Claude Code as an additional agent):

npx skills add NVIDIA-NeMo/DataDesigner

After installation, type /data-designer or describe the dataset you want and the skill will kick in.

🤝 Get involved

This repository supports agent-assisted development — see CONTRIBUTING.md for the recommended workflow.

Contributing Guide – How to contribute, including agent-assisted workflows
GitHub Issues – Report bugs or make a feature request

Telemetry

Data Designer collects telemetry to help us improve the library for developers. This data is not used to track any individual user behavior. It is used to see an aggregation of which models are the most popular for SDG. We will share this usage data with the community.

Disable with NEMO_TELEMETRY_ENABLED=false. More details →

Top models (YTD)

Aggregate model usage across synthetic data generation jobs, year-to-date 1/1/2026–6/1/2026:

Top models used for synthetic data generation

Last updated on June 1, 2026

License

Apache License 2.0 – see LICENSE for details.

Citation

If you use NeMo Data Designer in your research, please cite it using the following BibTeX entry:

@misc{nemo-data-designer,
  author = {The NeMo Data Designer Team, NVIDIA},
  title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data},
  howpublished = {\url{https://github.com/NVIDIA-NeMo/DataDesigner}},
  year = {2025},
  note = {GitHub Repository},
}

Telemetry & privacy

NeMo Data Designer includes an optional function to share anonymous telemetry data with NVIDIA for product improvement. Data collected is limited to names of models used and token counts (input and output). No user or device information is collected. This data is used to prioritize product improvements and will be shared in aggregate with the community. It is not used to track any individual user behavior.

You may opt out of telemetry collection at any time. Opting out applies only to data collection by the NeMo Data Designer library itself.

Use of third-party endpoints, including NVIDIA Build: NeMo Data Designer can be configured to use various inference endpoints, including build.nvidia.com (NVIDIA Build). If you choose to use NVIDIA Build or any other third-party endpoint, that endpoint's own terms of service and privacy practices apply independently of this library. Any opt-out you exercise within NeMo Data Designer does not extend to data collection by your chosen endpoint. NVIDIA Build is intended for evaluation and testing purposes only and may not be used in production environments. Do not submit any confidential information or personal data when using NVIDIA Build.

Project details

These details have not been verified by PyPI

Release history Release notifications | RSS feed

This version

0.7.0

Jun 26, 2026

0.7.0rc1 pre-release

Jun 26, 2026

0.6.1

Jun 1, 2026

0.6.1rc2 pre-release

Jun 1, 2026

0.6.1rc1 pre-release

Jun 1, 2026

0.6.0

May 13, 2026

0.6.0rc9 pre-release

May 13, 2026

0.6.0rc8 pre-release

May 13, 2026

0.6.0rc7 pre-release

May 12, 2026

0.6.0rc6 pre-release

May 12, 2026

0.6.0rc5 pre-release

May 12, 2026

0.6.0rc4 pre-release

May 12, 2026

0.6.0rc3 pre-release

May 12, 2026

0.6.0rc2 pre-release

May 12, 2026

0.6.0rc1 pre-release

May 12, 2026

0.5.9

Apr 28, 2026

0.5.8

Apr 27, 2026

0.5.8rc2 pre-release

Apr 27, 2026

0.5.8rc1 pre-release

Apr 24, 2026

0.5.7

Apr 17, 2026

0.5.7rc1 pre-release

Apr 17, 2026

0.5.6

Apr 9, 2026

0.5.6rc1 pre-release

Apr 9, 2026

0.5.5

Apr 2, 2026

0.5.5rc1 pre-release

Apr 2, 2026

0.5.4

Mar 25, 2026

0.5.4rc4 pre-release

Mar 25, 2026

0.5.4rc3 pre-release

Mar 19, 2026

0.5.4rc2 pre-release

Mar 17, 2026

0.5.4rc1 pre-release

Mar 15, 2026

0.5.3

Mar 12, 2026

0.5.3rc4 pre-release

Mar 12, 2026

0.5.3rc3 pre-release

Mar 12, 2026

0.5.3rc2 pre-release

Mar 12, 2026

0.5.3rc1 pre-release

Mar 12, 2026

0.5.2

Mar 5, 2026

0.5.1

Feb 20, 2026

0.5.1rc1 pre-release

Feb 20, 2026

0.5.0

Feb 11, 2026

0.5.0rc4 pre-release

Feb 11, 2026

0.5.0rc3 pre-release

Feb 10, 2026

0.5.0rc2 pre-release

Feb 5, 2026

0.5.0rc1 pre-release

Feb 3, 2026

0.4.0

Jan 31, 2026

0.4.0rc3 pre-release

Jan 31, 2026

0.4.0rc2 pre-release

Jan 29, 2026

0.4.0rc1 pre-release

Jan 28, 2026

0.3.8

Jan 27, 2026

0.3.8rc2 pre-release

Jan 26, 2026

0.3.8rc1 pre-release

Jan 21, 2026

0.3.7

Jan 17, 2026

0.3.6

Jan 17, 2026

0.3.5

Jan 16, 2026

0.3.4

Jan 14, 2026

0.3.3

Jan 12, 2026

0.3.2

Jan 9, 2026

0.3.1

Jan 8, 2026

0.3.0

Jan 8, 2026

0.2.3 yanked

Jan 7, 2026

Reason this release was yanked:

Potential exposure to litellm v1.82.8

0.2.2 yanked

Dec 30, 2025

Reason this release was yanked:

Potential exposure to litellm v1.82.8

0.2.1

Dec 19, 2025

0.2.0

Dec 17, 2025

0.1.5

Dec 11, 2025

0.1.4

Dec 8, 2025

0.1.3

Dec 3, 2025

0.1.2

Nov 24, 2025

0.1.1

Nov 21, 2025

0.1.0

Nov 20, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

data_designer-0.7.0.tar.gz (208.4 kB view details)

Uploaded Jun 26, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

data_designer-0.7.0-py3-none-any.whl (150.3 kB view details)

Uploaded Jun 26, 2026 Python 3

File details

Details for the file data_designer-0.7.0.tar.gz.

File metadata

Download URL: data_designer-0.7.0.tar.gz
Upload date: Jun 26, 2026
Size: 208.4 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for data_designer-0.7.0.tar.gz
Algorithm	Hash digest
SHA256	`c322682d7d14674e31d936be8db74e71e4eb31c9ab09e3ceef3c0fc84a5ddbeb`
MD5	`186fd82d4383eebd832e24e4a61caddb`
BLAKE2b-256	`3cdbd8ac36a0d46790cc793eec7a61b8a349e6a18ae28c5fbff3c43271097c69`

See more details on using hashes here.

File details

Details for the file data_designer-0.7.0-py3-none-any.whl.

File metadata

Download URL: data_designer-0.7.0-py3-none-any.whl
Upload date: Jun 26, 2026
Size: 150.3 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for data_designer-0.7.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`a97c65c45773ffd00d8a38ed0a3ab82c42fb627198b196e3f62ac3f30f9aee23`
MD5	`ba60a4e32bafddfbea96ef622c77f979`
BLAKE2b-256	`39be1d6b1736919346cbda39d347c866947b7afdd6f14c38b73d675b6180b271`

See more details on using hashes here.

data-designer 0.7.0

Navigation

Verified details

Maintainers

Unverified details

Meta

Classifiers

Project description

🎨 NeMo Data Designer

Welcome!

What can you do with Data Designer?

📣 Heads-up: async engine

Quick Start

1. Install

2. Set your API key

3. Start generating data!

What's next?

📚 Learn more

📝 Documentation

🔧 Configure models via CLI

🤖 Agent Skill

🤝 Get involved

Telemetry

Top models (YTD)

License

Citation

Telemetry & privacy

Project details

Verified details

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes