Skip to main content

MergenDB Banner

MergenDB

PyPI version npm version Python Versions License: MIT Tests Author

MergenDB is an ultra-compact, high-performance embedded columnar database engine designed to process massive analytical workloads and multi-million row table scans on resource-constrained hardware. It delivers strict zero external runtime dependencies -- requiring no C compilers, no native C++ binaries, and no bulky runtimes across both Python and Node.js.

Whether querying a 10-million row dataset on a 500 MB RAM VPS, analyzing telemetry streams on an edge Raspberry Pi, running real-time analytical reporting in Node.js/TypeScript, or managing hierarchical databases through the browser in Mergen Studio, MergenDB provides columnar speed with bounded memory guarantees.


What Is New in v0.7.1

  1. 500 Automated Resilience Tests Across Python & Node.js (1,000+ Total Scenarios):

    • Scaled both Python and Node.js test suites to 500 individual resilience scenarios each.
    • Tests extreme edge cases: ragged short/long rows, dirty null tokens (NULL, \N, NaN, nil), null bytes (\x00), unclosed quotes, escaped SQL quotes (O\'Connor), extreme floats, scientific notations, dynamic schema mutations, out-of-order columns, and multi-format export/re-import roundtrips.
    • Guaranteed 100% pass rate with zero crashes, robust error recovery, and bounded RAM usage (< 30 MB).
  2. Radiant Amber-Orange Visual Identity:

    • Brand new futuristic cinematic banner and geometric falcon emblem logo in warm obsidian and glowing amber-orange tones, active through the v0.8.x series.
  3. Synchronous Multi-Platform Distribution:

    • Released mergendb 0.7.1 and mergendb-studio 0.7.1 to PyPI.
    • Published mergendb@0.7.1 to npm.

What Is New in v0.7.0

  1. New Visual Identity & Radiant Amber-Orange Aesthetic:

    • Brand new futuristic cinematic banner and iconic geometric falcon emblem logo in glowing obsidian, fiery orange, and warm amber tones.
    • Designed to serve as the signature visual design across all repositories and platforms through the v0.7.x lifecycle.
  2. +100 Extreme Resilience & Fault Tolerance Tests (Python & Node.js):

    • Added 100 comprehensive edge-case test scenarios per platform testing corrupt headers, truncated blocks, mixed numeric/string comparisons, zero division prevention, irregular line breaks, extreme Unicode charsets, and full mutation-query-export roundtrips.
    • Total automated test suite now exceeds 510 tests (195 Python + 318 Node.js) with 100% pass rate.
  3. Harden Query Engine & Cross-Platform Path Normalization:

    • Safe type coercion in ExpressionEvaluator arithmetic and comparison operators, preventing unexpected TypeError or zero division exceptions on dirty data.
    • Fully normalized Windows path escaping in Table.query and SQL converter pipelines, eliminating backslash escape collisions.

What Is New in v0.6.10

  1. Decoupled Mergen Studio Web UI (Optional Upgrade Package):

    • The web interface has been decoupled from the core engine into an optional upgrade package (mergendb-studio), keeping the core engine minimal, lightweight, and headless.
    • Installable on demand via pip install "mergendb[studio]" or via the CLI runner mergen studio install.
  2. High-Concurrency RWLock & Multi-Threading Architecture:

    • Introduced fine-grained per-table RWLock and TableLockManager (mergendb/storage/lock.py).
    • Multiple concurrent readers execute queries without blocking each other, while writer operations (mutations, inserts, DDL) are executed with atomic staging file replacement and exponential backoff retry.
    • Strict bounded memory (< 30 MB peak RAM) safe on 500 MB VPS headless environments without GPU.
  3. 200-Combination Fault-Tolerant I/O Engine (Python & Node.js):

    • Resilient import engine handles dirty null tokens ("", "NULL", "none", "N/A", "NaN", "\\N", "nil", "-"), ragged rows, null bytes (\x00), multi-encodings (UTF-8, Latin1, CP1254, BOM), and corrupt JSON/JSONL/SQL dump fragments.
    • Single malformed records or syntax errors never cause the entire document to be discarded or ruined.
    • Validated across 200 distinct test combinations in both Python and Node.js SDK test suites.

What Is New in v0.6.8

  1. HTTP Streaming Export Payload Fix:

    • Resolved an issue in MergenDB Server where manual chunk-length framing headers were injected into streaming file downloads (_send_response_streaming_download), causing chunk byte counts to be saved into downloaded CSV/JSON/SQL files.
    • Configured protocol_version = "HTTP/1.1" and switched to raw streaming byte payloads, terminated cleanly upon stream closure (Connection: close). Multi-gigabyte downloads now write 100% clean, valid data files with zero corrupt header bytes.
  2. Mergen Studio Active Table Indicator in Export Tab:

    • Added an active table indicator badge (Selected Table: <name> (<N> rows)) directly inside the Export tab.
    • Prevents accidental exports of unintended tables by giving clear, unambiguous visual confirmation of which table is queued for export before clicking the Export button.
  3. Multi-Runtime Packaging & Version Alignment:

    • Bumped and synchronized all Python wheels/sdist and Node.js SDK npm packages to v0.6.8.

What Is New in v0.6.7

  1. Mergen Studio Table Selection & Navigation Fix:

    • Restored internationalization engine (EN/DE/TR) and resolved setLanguage reference errors in the browser client.
    • Synchronized top navigation dropdown (Table:) with the active database and active table selections.
    • Fully enabled one-click table browsing and structure views across all databases and subtables.
  2. Node.js SDK Multi-Database Context:

    • TableHandle now properly encapsulates its parent database context, routing all schema, query, update, delete, column mutations, export, and import commands to the intended database.
  3. CLI REPL Absolute Path Export & Terminal Polish:

    • REPL EXPORT now explicitly prints full absolute filesystem paths on export start and finish.
    • Cleaned terminal progress bar carriage-return output to eliminate leftover progress telemetry.

What Is New in v0.6.6

  1. Hierarchical Database Containers & Nested Sub-tables:

    • Organize data just like modern RDBMS platforms: Databases -> Tables -> Nested Sub-tables (e.g. enterprise.employees.engineering).
    • Store root records or partition sub-groups into isolated columnar files while preserving relational hierarchy.
    • Comprehensive SQL support: SHOW DATABASES;, CREATE DATABASE <name>;, DROP DATABASE <name>;, USE <name>;, SHOW TABLES [FROM <name>];.
    • Native dot-notation resolution across Python (mergendb.database(), table.create_subtable()), Node.js (client.database(), table.createSubtable()), CLI, REST server, and Studio Web UI.
  2. Zero-Memory Chunked Streaming Import & Export Engine:

    • Dedicated DataExporter streaming engine utilizing HTTP Chunked Transfer Encoding (Transfer-Encoding: chunked).
    • Tables of arbitrary size stream directly to disk or network sockets without ever accumulating full datasets into memory buffers.
    • Dedicated /import_stream endpoint accepts raw chunked streams in 64 KB blocks directly from network sockets to temporary disk files, eliminating browser V8 heap bloat and preventing GPU/compositor memory crashes.
    • Live 0% to 100% upload progress telemetry with transferred byte counters in Mergen Studio.
  3. Mergen Studio Web Dashboard Enhancements:

    • Complete phpMyAdmin-style tree hierarchy sidebar with collapsible database, table, and sub-table nodes.
    • Visual creation modals: + DB, + Table, and + Sub.
    • Streamlined DOM rendering without dataset string serialization, protecting browser memory.
    • Comprehensive internationalization: English (Default), German (Deutsch), and Turkish (Turkce).
    • Strict Zero-Emoji policy enforced across all interfaces, logs, and documentation.
  4. Node.js & TypeScript SDK 100% Parity:

    • Added DatabaseHandle, database(), listDatabases(), createDatabase(), dropDatabase().
    • Added createSubtable(), subtable(), listSubtables().
    • Added zero-memory file stream operations: exportToFile(destPath) and importFile(filePath) powered by standard library streams (pipe).

Zero-Dependency Installation

MergenDB requires zero external packages or compilers (dependencies: {}). It runs purely on the standard library of Python and Node.js.

Python Engine & CLI

# Install via PyPI
pip install --upgrade mergendb

# Run interactive CLI REPL directly:
python -m mergendb

# Start HTTP server & Mergen Studio:
python -m mergendb serve 8765

# Run embedded diagnostics & hardware benchmark:
python -m mergendb test

Node.js & TypeScript SDK

# Install SDK via npm
npm install mergendb

# Run server or open studio directly via npx:
npx mergendb serve 8765
npx mergendb studio 8765

Requires Python 3.8+ and/or Node.js 16+. Works out-of-the-box on Windows, macOS, Linux, and Docker with zero additional setup.


Why MergenDB? (The Problem with Row Stores)

Traditional embedded databases like SQLite store data row-by-row ([id, name, age, address, notes, ...]). When running analytical queries:

SELECT name, balance FROM users WHERE balance > 1000;

Even though you only care about name and balance, row stores must read every single column of every row off disk -- including massive text fields like address and notes. On a 10-million row database, that translates to gigabytes of unnecessary disk I/O and heavy memory exhaustion.

MergenDB utilizes the columnar approach:

  1. Column-Isolated I/O: Every column is stored and compressed independently. Unqueried columns are never read from disk.
  2. ZoneMap Pruning: Every block records min_value and max_value. If a block cannot contain matching rows, it is skipped with zero disk reads.
  3. 1024-bit Block Bloom Filters: Instant single-pass lookup index skips blocks that do not contain a queried ID, text, or UUID.
  4. Demand-Driven Late Materialization (LazyColumnDict): In multi-column filters (WHERE status = 'ACTIVE' AND balance > 50), MergenDB evaluates status first. If no rows in the block match, balance and all other columns are never decompressed.
  5. Strictly Bounded Memory: Data streams in small, tunable blocks (1,024-8,192 rows). Memory usage stays under 15-20 MB RAM, whether the database is 100 MB or 100 GB.

MergenDB Ecosystem Architecture

+---------------------------------------------------------------------------------+
|                                 CLIENT LAYER                                    |
|   Python Library (mergendb)   |   Node.js / TS SDK   |   Mergen Studio (Web)    |
|   db.find() / db.sql()        |   db.sql`...`        |   phpMyAdmin UI Tree     |
+---------------------------------------+-----------------------------------------+
                                        | HTTP / REST (Zero-Dependency)
+---------------------------------------v-----------------------------------------+
|                                SERVER ENGINE                                    |
|   Multi-threaded HTTP Server  |  Chunked Stream Transfer    |  Progress Stream  |
|   GET /export (Chunked)       |  POST /import_stream (Raw)  |  Hierarchical DB  |
+---------------------------------------+-----------------------------------------+
                                        | Analytical AST / Execution Plans
+---------------------------------------v-----------------------------------------+
|                              ANALYTICAL ENGINE                                  |
|   In-Memory Hash JOINs        |  Multi-Column GROUP BY / HAVING                 |
|   ZoneMap & Bloom Pruning     |  Vectorized Column Evaluation                   |
+---------------------------------------+-----------------------------------------+
                                        | Zero-Copy mmap & Block I/O
+---------------------------------------v-----------------------------------------+
|                        STORAGE & ADAPTIVE COMPRESSION                           |
|   Bit-Packed Booleans         |  Delta / Frame-of-Reference (FoR)               |
|   Block Dictionary Encoding   |  Run-Length Encoding (RLE)                      |
|   Secondary Zlib Stream       |  ZoneMap & Bloom Header (.mgdb)                 |
+---------------------------------------------------------------------------------+

Hierarchical Database & Nested Sub-tables System

MergenDB supports multi-tier hierarchical data management matching traditional relational databases while retaining columnar performance:

[DB] enterprise
 |-- [TBL] departments (1,200 rows)
 \-- [TBL] employees (4,500 rows)
      |-- [SUB] engineering (320 rows)
      \-- [SUB] marketing (150 rows)

Python API

import mergendb

# 1. Create or open database container
enterprise = mergendb.create_database("enterprise")

# 2. Create tables inside database
departments = enterprise.create_table("departments", [
    ("id", "INT64"),
    ("name", "STRING"),
    ("location", "STRING")
])
employees = enterprise.create_table("employees", [
    ("id", "INT64"),
    ("name", "STRING"),
    ("role", "STRING")
])

# 3. Insert records directly
employees.insert([
    {"id": 1, "name": "Alice", "role": "Staff Engineer"},
    {"id": 2, "name": "Bob", "role": "Data Scientist"},
])

# 4. Create nested sub-tables inside a table
engineering = employees.create_subtable("engineering", [
    ("employee_id", "INT64"),
    ("project_code", "STRING"),
    ("clearance_level", "INT32")
])
engineering.insert([{"employee_id": 1, "project_code": "ATLAS", "clearance_level": 4}])

# 5. Access via dot-notation
tbl = mergendb.connect("enterprise.employees.engineering")
results = tbl.find(employee_id=1)
print(results)

SQL Commands

SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
SHOW TABLES;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0);
SELECT * FROM employees;

Zero-Memory Streaming Import & Export

When handling multi-gigabyte datasets, traditional engines often buffer entire payloads into RAM, leading to memory exhaustion and browser compositor crashes. MergenDB resolves this with true streaming architecture:

1. Chunked Export (GET /export)

The server reads column blocks and yields encoded byte chunks directly into the HTTP response socket using standard HTTP Chunked Transfer Encoding. Server-side memory usage remains strictly bounded (< 1 MB RAM) regardless of table size.

# Stream table directly to disk
curl -N "http://localhost:8765/export?table=enterprise.employees&format=csv" -o employees.csv
curl -N "http://localhost:8765/export?table=enterprise.employees&format=json" -o employees.json
curl -N "http://localhost:8765/export?table=enterprise.employees&format=sql" -o employees.sql

2. Zero-Memory Import (POST /import_stream)

Mergen Studio streams the raw native File object directly over an HTTP socket using XMLHttpRequest. The server buffers incoming bytes in 64 KB chunks directly to a temporary file on disk, parses it with DataImporter, and indexes columnar blocks without consuming V8 heap memory.

# Direct zero-memory streaming upload
curl -X POST "http://localhost:8765/import_stream?table=enterprise.employees&format=csv" \
  --data-binary @large_dataset.csv

5-Minute Quickstart

1. Python API

import mergendb

# 1. Connect to table (Auto-created if not exists)
db = mergendb.connect("telemetry.mgdb")

# 2. Insert records (Schema is auto-inferred)
db.insert([
    {"id": 1, "sensor": "TEMP-01", "reading": 23.5, "active": True},
    {"id": 2, "sensor": "TEMP-02", "reading": 28.1, "active": True},
    {"id": 3, "sensor": "TEMP-01", "reading": 24.0, "active": False},
])

# 3. Document-Style Queries
active_sensors = db.find(sensor="TEMP-01", active=True)
first = db.find_one(sensor="TEMP-02")
matches = db.search("TEMP")  # Substring search across all text columns

# 4. Analytical SQL
res = db.sql("SELECT sensor, COUNT(*), AVG(reading) FROM telemetry GROUP BY sensor")
res.show()

# 5. Zero-Memory File Streaming
db.export_csv("telemetry_export.csv")
mergendb.from_csv("telemetry_export.csv", "backup.mgdb")

# 6. Live Hardware Diagnostics
mergendb.benchmark()

2. Node.js & TypeScript SDK

const { connect } = require('mergendb');

async function main() {
  const client = connect('http://localhost:8765');

  // Database Container Operations
  await client.createDatabase('analytics');
  const analyticsDb = client.database('analytics');

  // Create Table
  await analyticsDb.createTable('metrics', [
    { name: 'id', type: 'INT64' },
    { name: 'host', type: 'STRING' },
    { name: 'cpu_usage', type: 'FLOAT64' }
  ]);

  const metrics = analyticsDb.table('metrics');
  await metrics.insert([
    { id: 1, host: 'prod-api-1', cpu_usage: 42.5 },
    { id: 2, host: 'prod-api-2', cpu_usage: 78.1 }
  ]);

  // Nested Sub-table
  await metrics.createSubtable('hourly', [
    { name: 'metric_id', type: 'INT64' },
    { name: 'val', type: 'FLOAT64' }
  ]);
  const hourly = metrics.subtable('hourly');
  await hourly.insert([{ metric_id: 1, val: 41.2 }]);

  // Zero-Memory File Streaming
  await metrics.exportToFile('metrics_stream.csv', 'csv');
  await metrics.importFile('metrics_stream.csv', 'csv');

  // Analytical Query
  const summary = await client.sql`SELECT host, AVG(cpu_usage) FROM analytics.metrics GROUP BY host`;
  console.table(summary.rows);
}

main().catch(console.error);

3. Mergen CLI REPL

# Launch interactive REPL
mergen

# Inside REPL:
SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0), (2, 'Bob', 82000.0);
SELECT * FROM employees;
EXPORT employees TO CSV;
EXPORT employees TO JSON;
EXPORT employees TO SQL;

Adaptive Compression Pipeline

When writing column blocks, MergenDB analyzes incoming values and applies the most optimal encoding scheme:

  1. Bit-Packed Booleans: Stores boolean flags at 1 bit per value (8 values per byte).
  2. Delta / FoR (Frame-of-Reference): Stores sequential numbers as offsets from min_value, reducing 8-byte integers to 1- or 2-byte deltas.
  3. Block Dictionary Encoding: Optimal for low-cardinality text (gender, country, status). Stores unique strings once in a block dictionary and encodes rows as 1-byte indices.
  4. Run-Length Encoding (RLE): Collapses repeated identical values into (count, value) pairs.
  5. Secondary Zlib Compression: Applied to compressed byte streams for secondary compaction.

Performance Benchmarks

Tested on an Intel Core i5 / i7 with 100,000 mixed telemetry records (12 columns: integers, floats, timestamps, statuses, long strings):

Storage Format Disk Size Space Saved 2-Column Query Disk Read Peak RAM
JSON Lines (.jsonl) 19.5 MB 0% (Baseline) 19.5 MB Unbounded
SQLite 3 (.db) 8.1 MB 58.4% 8.1 MB (reads full row) ~30 MB
MergenDB (.mgdb) 1.6 MB 91.5% 0.29 MB (pruned) < 15 MB RAM
  • Exact Filter Scan Throughput: ~50,000,000 rows/sec (single core)
  • Substring (LIKE '%term%') Scan: ~10,000,000 rows/sec (single core)
  • SQL Streaming Import Speed: ~70,000-120,000 rows/sec on standard SSD

Version Changelog & Release Progression

Version Milestone Key Deliverables Status
v0.5.8 Analytical SQL & JOINs In-Memory Hash JOIN (INNER/LEFT), Multi-Column GROUP BY, HAVING Released
v0.5.9 Embedded Web UI Initial Mergen Studio web interface Released
v0.6.0 Universal Node.js SDK Zero-dependency Node.js/TypeScript SDK + CLI runner Released
v0.6.1 phpMyAdmin Overhaul 100% CLI feature parity in browser and Node.js SDK, 106 tests Released
v0.6.2 Multi-Runtime & Zero-Dependency python -m mergendb & npx mergendb runners, autoStart: true, GET /query Released
v0.6.3 Official Identity & i18n Official Logo, strict zero-emoji policy, English/German/Turkish localization Released
v0.6.4 Hierarchical Architecture Database containers, nested sub-tables, and streaming export/import Released
v0.6.5 Zero-Memory Streaming Engine Zero-Memory Chunked Streaming Engine, crash-free browser upload Released
v0.6.6 Hierarchical Containers & Streaming Database containers, sub-tables, and chunked streaming import/export Released
v0.6.7 Studio Navigation & Table Selection Restored table selection, top nav sync, i18n fix, Node.js multi-database routing Released
v0.6.8 Streaming Raw Payload & Export UX Clean streaming export raw payload, active table indicator in Studio export Current Release

Running the Test Suite

MergenDB includes an embedded test suite with 107 comprehensive tests (89 Python unit tests covering storage, compression algorithms, query planning, Bloom filters, hierarchical databases, sub-tables, and analytical joins + 18 end-to-end Node.js SDK integration tests):

# Run Python unit tests via unittest
python -m unittest discover -s tests

# Run Node.js Client SDK integration tests
node sdks/nodejs/test.js

License

Distributed under the MIT License. See LICENSE for details.

Developed by Ugur Turker Kebeci.

Metadata

Release files for mergendb 0.7.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mergendb 0.7.1
File Size Uploaded
mergendb-0.7.1.tar.gz 156.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mergendb 0.7.1
File Interpreter ABI Platform
mergendb-0.7.1-py3-none-any.whl Python 3 none any Details

Total release size: 288.6 kB

Release files / mergendb-0.7.1.tar.gz

Download URL mergendb-0.7.1.tar.gz
Size 156.3 kB
Tags Source
SHA-256 checksum
How to use checksums
68fc3b1138fbb65a6cdcd02aa51752014c48044b4c572537c1d1cca5f7ad509c
BLAKE2b-256 checksum
How to use checksums
ae4fb5180c8a52a1fce807f9e4e8507d0ad7bd1cfb3ec5df1c868bb25f4cf659
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.7rc1

Release files / mergendb-0.7.1-py3-none-any.whl

Download URL mergendb-0.7.1-py3-none-any.whl
Size 132.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8d57fffba0c764233c980a8193439d626e216179aff249c9f4d54e35f012c487
BLAKE2b-256 checksum
How to use checksums
e84f00361a089c0af9784a8b836a150a7536b2690f2274a70b5d457a54baebf3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.7rc1

Release history Release notifications | RSS feed

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.4

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

This release

0.7.1 This release

2 release files

0.7.0

2 release files

0.6.10

2 release files

0.6.9

2 release files

0.6.8

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page