Skip to main content
Misata

Misata

The relational synthetic data engine that satisfies exact business outcomes.

Generate complete multi-table databases with foreign keys that resolve, math that reconciles, and domain-authentic prose with zero Lorem Ipsum. From a single sentence, YAML schema, or live database.

PyPI version Python versions CI License Open in Colab Paper smithery badge Misata Studio

Prefer a visual interface? Try Misata Studio to design schemas on an interactive canvas and generate datasets directly in your browser.


The Paradigm Shift

Most synthetic data tools take existing data and imitate it. But in modern engineering, you rarely have clean data to start with—or you need data designed around a specific target outcome:

  • "Monthly revenue rises from $50k to $200k with a Q3 slump."
  • "Fraud rate starts at 1% in Q1 and climbs to 6% by Q4."
  • "Every customer's total_spent strictly equals the sum of their order line items."

Misata works in reverse: you declare the outcome, and Misata solves for the individual rows that hit it to $0.00 error, guaranteed by a closed-form Gamma conditional-sum mechanism (arXiv:2606.08736). No machine learning models, no real data required, and zero hallucinated foreign keys.


⚡ 30-Second Quickstart

Install via pip:

pip install misata

CLI

Generate a complete relational dataset from a single description:

misata generate \
  --story "Brazilian fintech with R$ payments, CPF verification, and 3% fraud" \
  --rows 1000 \
  --output-dir ./demo_data

Outputs clean CSVs and an oracle_report.json verifying zero foreign key orphans, constraint satisfaction, and statistical fidelity.

Python

import misata

# Generate multi-table DataFrames from plain English
data = misata.generate("SaaS startup with 500 users, monthly subscriptions, and 12% churn")

users_df = data["users"]
subscriptions_df = data["subscriptions"]

Enrich Existing Data in 1 Line

Replace boring or blank columns in an existing DataFrame with domain-authentic prose:

import misata
import pandas as pd

df = pd.read_csv("support_cases.csv")
# Automatically detects semantic columns (subjects, resolution notes, memos, error traces)
df_enriched = misata.enrich_text(df, seed=42)

🚀 The 6 Unique Capabilities of Misata

What separates Misata from legacy libraries like Faker and imitation models like SDV:

1. Exact Outcome Conformance ($0.00 Error)

Declare the aggregate outcome (revenue curve, churn rate, seasonal surge, default curve), and Misata generates micro-rows whose monthly or annual sums match your target to $0.00 error. While off-the-shelf synthesizers miss aggregate targets by 74–86%, Misata hits them provably (read the research paper).

2. Instant Sandbox Oracle for AI Coding Agents

Misata includes a native Model Context Protocol (MCP) server. Cursor, Claude Code, and Windsurf can call create_sandbox to spin up an isolated SQLite database seeded with realistic relational data in 2 seconds. AI agents can test SQL queries and test code against real tables instead of guessing schema.

pip install "misata[mcp]" && misata mcp install --client all

3. Topological DAG Relational Integrity (0 Orphan Foreign Keys)

Parent-child relationships across 10+ tables are resolved via topological rank ordering. Self-referential hierarchies, composite unique constraints, and temporal causality (signup_date <= order_date <= payment_date <= refund_date) are strictly enforced.

4. Cross-Vertical Text Realism (Zero Lorem Ipsum)

No Latin placeholder gibberish. Combinatorial microtext generators provide authentic text across 15+ real-world industries:

  • E-Commerce: Category-conditioned product descriptions (electronics, apparel, home, beauty), authentic return reasons, delivery notes.
  • B2B SaaS: Support ticket subjects, multi-line issue descriptions, agent resolution notes, competitor churn reasons.
  • FinTech: Bank statement descriptors, ACH/wire remittance memos, AML audit overrides.
  • Healthcare: Clinical SOAP progress notes, chief complaints, discharge instructions, medical dosage schedules (TID, QID, PRN).
  • Customer Reviews: Sentiment calibrated provably to 1-to-5 star ratings.

5. Vectorized Speed (100x–500x Faster Than Faker)

While Faker iterates row-by-row in pure Python (5,000–15,000 rows/s), Misata uses vectorized NumPy array operations. It generates 500,000 to 16,000,000 rows/second on a single CPU core.

6. Introspect & Seed Live Databases (misata seed)

Point Misata directly at a PostgreSQL, MySQL, or SQLite database. Misata introspects foreign keys, check constraints, enums, and column types, generates a topological insertion plan, and seeds production-like data directly back into your database without writing schema code.


🎯 Dozens of Real-World Use Cases

Misata is built for engineers, testers, data teams, and founders across dozens of everyday workloads:

🛠️ Data Engineering & ETL Pipelines

  • Known-Answer Pipeline Testing: Declare exact KPI targets (e.g. $1.2M Q4 revenue), generate synthetic raw tables, and verify your dbt, Spark, or SQL transforms return the exact expected figure.
  • High-Throughput Stress Testing: Churn 10,000,000+ rows in seconds to test partition boundaries, shuffle performance, and warehouse scaling.
  • Schema Migration Dry Runs: Rehearse destructive column backfills and foreign key additions against realistic data before deploying to production.
  • Deterministic CI/CD Fixtures: Reproducible seeds ensure test assertions never flake across continuous integration runs.

🤖 AI Coding Agents & LLM Development

  • Agent SQL Sandboxes: Provide Cursor, Claude Code, and Windsurf with isolated, pre-seeded databases to validate SQL queries without production access.
  • Text-to-SQL Benchmark Generation: Build complex relational schema benches with diverse joins to evaluate fine-tuned coding models.
  • Evaluation Database Packs (Evalpacks): Create verified eval datasets with independent DuckDB answer keys where the ground truth cannot be wrong.

🗄️ Application Development & Database Seeding

  • Local & Staging Environment Seeding: Seed staging Postgres, MySQL, or SQLite databases with realistic customer histories and 0 broken foreign keys (misata seed).
  • ORM Model Fixtures: Generate rich fixtures matching your Prisma schema or SQLAlchemy declarative models.
  • Multi-Tenant Isolation Verification: Test row-level security (RLS) and tenant isolation rules without cross-tenant key leakage.
  • Incremental Data Growth (generate_diff): Add 5,000 new rows to an existing dataset while auto-offsetting IDs and preserving referential integrity.

📊 BI, Product Demos & Sales Engineering

  • Board-Ready Dashboard Demos: Populate Tableau, PowerBI, and Metabase dashboards with convincing seasonality (Black Friday spikes, summer dips) rather than flat random noise.
  • Sales Engineering Prototypes: Demo customer-facing analytics with authentic company names, human names, and transaction histories with zero PII exposure.
  • Feature Previews: Preview upcoming charts, cohorts, and metrics before production customer data accumulates.

💳 FinTech, Banking & Payments

  • Double-Entry Ledger Balancing: Generate accounting transactions where total debits strictly equal credits across every ledger account.
  • Credit Risk & Loan Tapes: Calibrate delinquency curves, credit score distributions, and default rates for credit portfolio testing.
  • AML & Fraud Detection Testing: Inforce exact fraud incidence rates (e.g. 2.4%) with authentic transaction memos and AML audit trails.
  • Payment Remittance: Simulate SWIFT, ACH, and card transactions with valid routing numbers, CVVs, and statement descriptors.

🏥 Healthcare & Clinical Informatics

  • HIPAA Safe-Harbor Synthetic Cohorts: Generate realistic patient populations, vital signs, and encounter histories with zero PHI liability.
  • Clinical NLP Model Evaluation: Evaluate healthcare LLMs against authentic SOAP notes, chief complaints, and discharge summaries.
  • Ward & Scheduling Simulation: Simulate hospital appointment grids with realistic 15-minute intervals, business hours, and weekend dips.

📦 E-Commerce & Supply Chain Logistics

  • Multi-Category Catalog Modeling: Produce realistic item specs and descriptions conditioned on category (electronics, apparel, home, industrial).
  • Return & Refund Workflows: Simulate return logistics with authentic return reasons, restocking milestones, and customer refund dates.
  • Route & Fleet Optimization: Compute realistic routes with Haversine distance calculations and valid secondary addresses (Apt, Suite, Bldg).

🛡️ Cybersecurity & IT Infrastructure

  • Network Intrusion Datasets: Generate netflow logs, port scans, and DDoS traffic patterns for security tool benchmarking.
  • System Exception & Error Analysis: Populate observability dashboards with realistic deadlocks, HTTP 504 timeouts, and connection pool exhaustion logs.
  • Compliance Audit Logging: Simulate SOC2/HIPAA access logs with documented managerial access override justifications.

🔬 Machine Learning & Statistical Research

  • Synthetic Twins from CSV (misata.mimic): Clone distributions and correlations from sensitive CSVs without copying a single original row.
  • Hierarchical Cluster Modeling (ICC): Generate multi-site data with specified Intraclass Correlation Coefficients for mixed-effects regression.
  • Time-Series Autocorrelation (AR1): Generate longitudinal entity trajectories that maintain realistic temporal memory.

⚡ Why You Should Never Use Faker Again

Faker was built over a decade ago for single-attribute mock values. For modern applications, it introduces critical failure modes:

Problem in 2026 Faker Reality Misata 0.9.6.60 Advantage
Relational Topology ✗ 0 concept of databases or FKs; manual glue code required ✓ Strict topological DAG; 0 orphan FKs guaranteed
Cross-Column Coherence ✗ Incoherent (e.g. "Male" name, mismatched email, invalid city) ✓ Coherent identities, addresses, and causality
Text Realism ✗ 2,000-year-old Latin "Lorem Ipsum" or robotic templates ✓ 15+ domain microtext pools (SOAP notes, tickets, memos)
Mathematical Consistency ✗ Violates basic accounting (price * qty != total) ✓ Exact mathematical formulas and balanced ledgers
Performance ✗ ~10k rows/s (single-threaded Python loops) ✓ 500k to 16M rows/s (Vectorized NumPy engine)
Aggregate Targets ✗ Impossible (uniform random noise) ✓ Exact closed-form outcome conformance ($0.00 error)
Database Seeding ✗ Manual SQL scripts or ORM boilerplate ✓ One-command introspection and seeding (misata seed)

Read the complete Faker vs SDV vs Misata Guide for full benchmarks and code comparisons.


🛠️ Eight Ways to Generate Data

Misata fits whatever workflow you already use:

Input Mode Best For Learn More
1. Plain English Story Rapid prototyping, zero configuration Story Guide
2. YAML Schema-as-Code Committing versioned data definitions to git YAML Guide
3. Live Database Seeding Introspecting and populating Postgres, MySQL, SQLite Database Seeding Guide
4. Python Dict Schema Programmatic in-memory generation in Python scripts Dict Schema Guide
5. dbt Project Schemas Generating fixtures directly from schema.yml dbt Seeding Guide
6. Prisma Schema Next.js and Node.js developers seeding full-stack apps Prisma Guide
7. Multi-Provider LLMs Groq, OpenAI, Claude, Gemini, or Ollama-driven schemas LLM Guide
8. Incremental Growth Appending rows with offset IDs and preserved FKs Incremental Guide

🌐 20+ Built-in Industry Domains

Generate domain-complete schemas with tuned statistical distributions out of the box:

SaaS · E-Commerce · FinTech · Healthcare · Logistics · Credit Risk · HR & People · Streaming Media · Insurance · CRM & Sales · Food Delivery · Travel & Hospitality · Gaming · Crypto & DeFi · Predictive Maintenance · Network Intrusion · Islamic Finance · EdTech · Real Estate · Contact Centers · Manufacturing SPC

See the Complete Domain Catalog.


⚡ Performance

Measured on standard Apple M-series hardware (single CPU core, no GPU):

Workload Row Count Generation Time Throughput
Single table (lognormal distribution) 1,000,000 0.06 s ~16M rows/s
Star schema (5 tables, 4 FK dependencies) 1,055,030 1.54 s ~687k rows/s
Multi-table enterprise database 100,000 0.42 s ~240k rows/s

📚 Documentation Index

For in-depth guides, API references, and architecture deep dives:


📄 Research & Citation

The closed-form exact-outcome conformance engine is formalised in arXiv preprint 2606.08736:

@article{rasin2026declarative,
  title   = {Declarative Outcome-Conformant Synthesis: Exact, Closed-Form
             Specification Satisfaction and a Conformance Benchmark},
  author  = {Rasin, Muhammed},
  year    = {2026},
  url     = {https://arxiv.org/abs/2606.08736v1}
}

🤝 Contributing & Development Status

Misata is currently under massive, rapid development to push synthetic realism to its absolute limit: expanding real-world domain knowledge, deepening seed pool vocabularies, elevating textual column realism, and advancing statistical fidelity across every industry vertical.

Contributions from domain experts, data engineers, and researchers are warmly welcomed! Whether you want to:

  • Enrich Text & Vocabulary Pools: Add authentic seeds and grammar rules for specialized domains in misata/vocab_seeds.py and misata/microtext.py.
  • Contribute a Domain Capsule: Expand built-in industry templates (healthcare, legal, banking, engineering, supply chain).
  • Advance Statistical Fidelity: Improve multi-variate copulas, time-series dynamics, or outcome-curve solvers.
  • Report Edge Cases & Realism Flaws: Open an issue or discussion whenever generated values don't look 100% human-authentic.
git clone https://github.com/rasinmuhammed/misata
cd Misata
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest -q

Misata is open-source under the MIT License.

Metadata

Release files for misata 0.9.6.60

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for misata 0.9.6.60
File Size Uploaded
misata-0.9.6.60.tar.gz 1.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for misata 0.9.6.60
File Interpreter ABI Platform
misata-0.9.6.60-py3-none-any.whl Python 3 none any Details

Total release size: 2.1 MB

Release files / misata-0.9.6.60.tar.gz

Download URL misata-0.9.6.60.tar.gz
Size 1.1 MB
Tags Source
SHA-256 checksum
How to use checksums
f50f9650d29a8e37135cee8994be954dc5d727fdfb19e1be051f7bb660db03b8
BLAKE2b-256 checksum
How to use checksums
9a99a57fbeb826d47a15a522784494bb698b7e8a74392a1eeb44c367e540ae80
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / misata-0.9.6.60-py3-none-any.whl

Download URL misata-0.9.6.60-py3-none-any.whl
Size 928.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8556405ea30516bc48be6ce1e83162143c45e313ee54b54337ab383356896000
BLAKE2b-256 checksum
How to use checksums
f44b88560622d4a733d751fffbadd2bcc9a5a943f3d6213f95cd4eaaf206d495
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.9.6.60 This release

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page