MergenDB
MergenDB is an ultra-compact, high-performance embedded columnar database engine designed to process massive analytical workloads and multi-million row table scans on resource-constrained hardware. It delivers strict zero external runtime dependencies -- requiring no C compilers, no native C++ binaries, and no bulky runtimes across both Python and Node.js.
Whether querying a 10-million row dataset on a 500 MB RAM VPS, analyzing telemetry streams on an edge Raspberry Pi, running real-time analytical reporting in Node.js/TypeScript, or managing hierarchical databases through the browser in Mergen Studio, MergenDB provides columnar speed with bounded memory guarantees.
What Is New in v0.7.1
-
500 Automated Resilience Tests Across Python & Node.js (1,000+ Total Scenarios):
- Scaled both Python and Node.js test suites to 500 individual resilience scenarios each.
- Tests extreme edge cases: ragged short/long rows, dirty null tokens (
NULL,\N,NaN,nil), null bytes (\x00), unclosed quotes, escaped SQL quotes (O\'Connor), extreme floats, scientific notations, dynamic schema mutations, out-of-order columns, and multi-format export/re-import roundtrips. - Guaranteed 100% pass rate with zero crashes, robust error recovery, and bounded RAM usage (< 30 MB).
-
Radiant Amber-Orange Visual Identity:
- Brand new futuristic cinematic banner and geometric falcon emblem logo in warm obsidian and glowing amber-orange tones, active through the v0.8.x series.
-
Synchronous Multi-Platform Distribution:
- Released
mergendb0.7.1 andmergendb-studio0.7.1 to PyPI. - Published
mergendb@0.7.1to npm.
- Released
What Is New in v0.7.0
-
New Visual Identity & Radiant Amber-Orange Aesthetic:
- Brand new futuristic cinematic banner and iconic geometric falcon emblem logo in glowing obsidian, fiery orange, and warm amber tones.
- Designed to serve as the signature visual design across all repositories and platforms through the v0.7.x lifecycle.
-
+100 Extreme Resilience & Fault Tolerance Tests (Python & Node.js):
- Added 100 comprehensive edge-case test scenarios per platform testing corrupt headers, truncated blocks, mixed numeric/string comparisons, zero division prevention, irregular line breaks, extreme Unicode charsets, and full mutation-query-export roundtrips.
- Total automated test suite now exceeds 510 tests (195 Python + 318 Node.js) with 100% pass rate.
-
Harden Query Engine & Cross-Platform Path Normalization:
- Safe type coercion in
ExpressionEvaluatorarithmetic and comparison operators, preventing unexpectedTypeErroror zero division exceptions on dirty data. - Fully normalized Windows path escaping in
Table.queryand SQL converter pipelines, eliminating backslash escape collisions.
- Safe type coercion in
What Is New in v0.6.10
-
Decoupled Mergen Studio Web UI (Optional Upgrade Package):
- The web interface has been decoupled from the core engine into an optional upgrade package (
mergendb-studio), keeping the core engine minimal, lightweight, and headless. - Installable on demand via
pip install "mergendb[studio]"or via the CLI runnermergen studio install.
- The web interface has been decoupled from the core engine into an optional upgrade package (
-
High-Concurrency RWLock & Multi-Threading Architecture:
- Introduced fine-grained per-table
RWLockandTableLockManager(mergendb/storage/lock.py). - Multiple concurrent readers execute queries without blocking each other, while writer operations (mutations, inserts, DDL) are executed with atomic staging file replacement and exponential backoff retry.
- Strict bounded memory (< 30 MB peak RAM) safe on 500 MB VPS headless environments without GPU.
- Introduced fine-grained per-table
-
200-Combination Fault-Tolerant I/O Engine (Python & Node.js):
- Resilient import engine handles dirty null tokens (
"","NULL","none","N/A","NaN","\\N","nil","-"), ragged rows, null bytes (\x00), multi-encodings (UTF-8, Latin1, CP1254, BOM), and corrupt JSON/JSONL/SQL dump fragments. - Single malformed records or syntax errors never cause the entire document to be discarded or ruined.
- Validated across 200 distinct test combinations in both Python and Node.js SDK test suites.
- Resilient import engine handles dirty null tokens (
What Is New in v0.6.8
-
HTTP Streaming Export Payload Fix:
- Resolved an issue in MergenDB Server where manual chunk-length framing headers were injected into streaming file downloads (
_send_response_streaming_download), causing chunk byte counts to be saved into downloaded CSV/JSON/SQL files. - Configured
protocol_version = "HTTP/1.1"and switched to raw streaming byte payloads, terminated cleanly upon stream closure (Connection: close). Multi-gigabyte downloads now write 100% clean, valid data files with zero corrupt header bytes.
- Resolved an issue in MergenDB Server where manual chunk-length framing headers were injected into streaming file downloads (
-
Mergen Studio Active Table Indicator in Export Tab:
- Added an active table indicator badge (
Selected Table: <name> (<N> rows)) directly inside the Export tab. - Prevents accidental exports of unintended tables by giving clear, unambiguous visual confirmation of which table is queued for export before clicking the Export button.
- Added an active table indicator badge (
-
Multi-Runtime Packaging & Version Alignment:
- Bumped and synchronized all Python wheels/sdist and Node.js SDK npm packages to
v0.6.8.
- Bumped and synchronized all Python wheels/sdist and Node.js SDK npm packages to
What Is New in v0.6.7
-
Mergen Studio Table Selection & Navigation Fix:
- Restored internationalization engine (EN/DE/TR) and resolved
setLanguagereference errors in the browser client. - Synchronized top navigation dropdown (
Table:) with the active database and active table selections. - Fully enabled one-click table browsing and structure views across all databases and subtables.
- Restored internationalization engine (EN/DE/TR) and resolved
-
Node.js SDK Multi-Database Context:
TableHandlenow properly encapsulates its parent database context, routing all schema, query, update, delete, column mutations, export, and import commands to the intended database.
-
CLI REPL Absolute Path Export & Terminal Polish:
- REPL
EXPORTnow explicitly prints full absolute filesystem paths on export start and finish. - Cleaned terminal progress bar carriage-return output to eliminate leftover progress telemetry.
- REPL
What Is New in v0.6.6
-
Hierarchical Database Containers & Nested Sub-tables:
- Organize data just like modern RDBMS platforms: Databases -> Tables -> Nested Sub-tables (e.g.
enterprise.employees.engineering). - Store root records or partition sub-groups into isolated columnar files while preserving relational hierarchy.
- Comprehensive SQL support:
SHOW DATABASES;,CREATE DATABASE <name>;,DROP DATABASE <name>;,USE <name>;,SHOW TABLES [FROM <name>];. - Native dot-notation resolution across Python (
mergendb.database(),table.create_subtable()), Node.js (client.database(),table.createSubtable()), CLI, REST server, and Studio Web UI.
- Organize data just like modern RDBMS platforms: Databases -> Tables -> Nested Sub-tables (e.g.
-
Zero-Memory Chunked Streaming Import & Export Engine:
- Dedicated
DataExporterstreaming engine utilizing HTTP Chunked Transfer Encoding (Transfer-Encoding: chunked). - Tables of arbitrary size stream directly to disk or network sockets without ever accumulating full datasets into memory buffers.
- Dedicated
/import_streamendpoint accepts raw chunked streams in 64 KB blocks directly from network sockets to temporary disk files, eliminating browser V8 heap bloat and preventing GPU/compositor memory crashes. - Live 0% to 100% upload progress telemetry with transferred byte counters in Mergen Studio.
- Dedicated
-
Mergen Studio Web Dashboard Enhancements:
- Complete phpMyAdmin-style tree hierarchy sidebar with collapsible database, table, and sub-table nodes.
- Visual creation modals: + DB, + Table, and + Sub.
- Streamlined DOM rendering without dataset string serialization, protecting browser memory.
- Comprehensive internationalization: English (Default), German (Deutsch), and Turkish (Turkce).
- Strict Zero-Emoji policy enforced across all interfaces, logs, and documentation.
-
Node.js & TypeScript SDK 100% Parity:
- Added
DatabaseHandle,database(),listDatabases(),createDatabase(),dropDatabase(). - Added
createSubtable(),subtable(),listSubtables(). - Added zero-memory file stream operations:
exportToFile(destPath)andimportFile(filePath)powered by standard library streams (pipe).
- Added
Zero-Dependency Installation
MergenDB requires zero external packages or compilers (dependencies: {}). It runs purely on the standard library of Python and Node.js.
Python Engine & CLI
# Install via PyPI
pip install --upgrade mergendb
# Run interactive CLI REPL directly:
python -m mergendb
# Start HTTP server & Mergen Studio:
python -m mergendb serve 8765
# Run embedded diagnostics & hardware benchmark:
python -m mergendb test
Node.js & TypeScript SDK
# Install SDK via npm
npm install mergendb
# Run server or open studio directly via npx:
npx mergendb serve 8765
npx mergendb studio 8765
Requires Python 3.8+ and/or Node.js 16+. Works out-of-the-box on Windows, macOS, Linux, and Docker with zero additional setup.
Why MergenDB? (The Problem with Row Stores)
Traditional embedded databases like SQLite store data row-by-row ([id, name, age, address, notes, ...]). When running analytical queries:
SELECT name, balance FROM users WHERE balance > 1000;
Even though you only care about name and balance, row stores must read every single column of every row off disk -- including massive text fields like address and notes. On a 10-million row database, that translates to gigabytes of unnecessary disk I/O and heavy memory exhaustion.
MergenDB utilizes the columnar approach:
- Column-Isolated I/O: Every column is stored and compressed independently. Unqueried columns are never read from disk.
- ZoneMap Pruning: Every block records
min_valueandmax_value. If a block cannot contain matching rows, it is skipped with zero disk reads. - 1024-bit Block Bloom Filters: Instant single-pass lookup index skips blocks that do not contain a queried ID, text, or UUID.
- Demand-Driven Late Materialization (
LazyColumnDict): In multi-column filters (WHERE status = 'ACTIVE' AND balance > 50), MergenDB evaluatesstatusfirst. If no rows in the block match,balanceand all other columns are never decompressed. - Strictly Bounded Memory: Data streams in small, tunable blocks (1,024-8,192 rows). Memory usage stays under 15-20 MB RAM, whether the database is 100 MB or 100 GB.
MergenDB Ecosystem Architecture
+---------------------------------------------------------------------------------+
| CLIENT LAYER |
| Python Library (mergendb) | Node.js / TS SDK | Mergen Studio (Web) |
| db.find() / db.sql() | db.sql`...` | phpMyAdmin UI Tree |
+---------------------------------------+-----------------------------------------+
| HTTP / REST (Zero-Dependency)
+---------------------------------------v-----------------------------------------+
| SERVER ENGINE |
| Multi-threaded HTTP Server | Chunked Stream Transfer | Progress Stream |
| GET /export (Chunked) | POST /import_stream (Raw) | Hierarchical DB |
+---------------------------------------+-----------------------------------------+
| Analytical AST / Execution Plans
+---------------------------------------v-----------------------------------------+
| ANALYTICAL ENGINE |
| In-Memory Hash JOINs | Multi-Column GROUP BY / HAVING |
| ZoneMap & Bloom Pruning | Vectorized Column Evaluation |
+---------------------------------------+-----------------------------------------+
| Zero-Copy mmap & Block I/O
+---------------------------------------v-----------------------------------------+
| STORAGE & ADAPTIVE COMPRESSION |
| Bit-Packed Booleans | Delta / Frame-of-Reference (FoR) |
| Block Dictionary Encoding | Run-Length Encoding (RLE) |
| Secondary Zlib Stream | ZoneMap & Bloom Header (.mgdb) |
+---------------------------------------------------------------------------------+
Hierarchical Database & Nested Sub-tables System
MergenDB supports multi-tier hierarchical data management matching traditional relational databases while retaining columnar performance:
[DB] enterprise
|-- [TBL] departments (1,200 rows)
\-- [TBL] employees (4,500 rows)
|-- [SUB] engineering (320 rows)
\-- [SUB] marketing (150 rows)
Python API
import mergendb
# 1. Create or open database container
enterprise = mergendb.create_database("enterprise")
# 2. Create tables inside database
departments = enterprise.create_table("departments", [
("id", "INT64"),
("name", "STRING"),
("location", "STRING")
])
employees = enterprise.create_table("employees", [
("id", "INT64"),
("name", "STRING"),
("role", "STRING")
])
# 3. Insert records directly
employees.insert([
{"id": 1, "name": "Alice", "role": "Staff Engineer"},
{"id": 2, "name": "Bob", "role": "Data Scientist"},
])
# 4. Create nested sub-tables inside a table
engineering = employees.create_subtable("engineering", [
("employee_id", "INT64"),
("project_code", "STRING"),
("clearance_level", "INT32")
])
engineering.insert([{"employee_id": 1, "project_code": "ATLAS", "clearance_level": 4}])
# 5. Access via dot-notation
tbl = mergendb.connect("enterprise.employees.engineering")
results = tbl.find(employee_id=1)
print(results)
SQL Commands
SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
SHOW TABLES;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0);
SELECT * FROM employees;
Zero-Memory Streaming Import & Export
When handling multi-gigabyte datasets, traditional engines often buffer entire payloads into RAM, leading to memory exhaustion and browser compositor crashes. MergenDB resolves this with true streaming architecture:
1. Chunked Export (GET /export)
The server reads column blocks and yields encoded byte chunks directly into the HTTP response socket using standard HTTP Chunked Transfer Encoding. Server-side memory usage remains strictly bounded (< 1 MB RAM) regardless of table size.
# Stream table directly to disk
curl -N "http://localhost:8765/export?table=enterprise.employees&format=csv" -o employees.csv
curl -N "http://localhost:8765/export?table=enterprise.employees&format=json" -o employees.json
curl -N "http://localhost:8765/export?table=enterprise.employees&format=sql" -o employees.sql
2. Zero-Memory Import (POST /import_stream)
Mergen Studio streams the raw native File object directly over an HTTP socket using XMLHttpRequest. The server buffers incoming bytes in 64 KB chunks directly to a temporary file on disk, parses it with DataImporter, and indexes columnar blocks without consuming V8 heap memory.
# Direct zero-memory streaming upload
curl -X POST "http://localhost:8765/import_stream?table=enterprise.employees&format=csv" \
--data-binary @large_dataset.csv
5-Minute Quickstart
1. Python API
import mergendb
# 1. Connect to table (Auto-created if not exists)
db = mergendb.connect("telemetry.mgdb")
# 2. Insert records (Schema is auto-inferred)
db.insert([
{"id": 1, "sensor": "TEMP-01", "reading": 23.5, "active": True},
{"id": 2, "sensor": "TEMP-02", "reading": 28.1, "active": True},
{"id": 3, "sensor": "TEMP-01", "reading": 24.0, "active": False},
])
# 3. Document-Style Queries
active_sensors = db.find(sensor="TEMP-01", active=True)
first = db.find_one(sensor="TEMP-02")
matches = db.search("TEMP") # Substring search across all text columns
# 4. Analytical SQL
res = db.sql("SELECT sensor, COUNT(*), AVG(reading) FROM telemetry GROUP BY sensor")
res.show()
# 5. Zero-Memory File Streaming
db.export_csv("telemetry_export.csv")
mergendb.from_csv("telemetry_export.csv", "backup.mgdb")
# 6. Live Hardware Diagnostics
mergendb.benchmark()
2. Node.js & TypeScript SDK
const { connect } = require('mergendb');
async function main() {
const client = connect('http://localhost:8765');
// Database Container Operations
await client.createDatabase('analytics');
const analyticsDb = client.database('analytics');
// Create Table
await analyticsDb.createTable('metrics', [
{ name: 'id', type: 'INT64' },
{ name: 'host', type: 'STRING' },
{ name: 'cpu_usage', type: 'FLOAT64' }
]);
const metrics = analyticsDb.table('metrics');
await metrics.insert([
{ id: 1, host: 'prod-api-1', cpu_usage: 42.5 },
{ id: 2, host: 'prod-api-2', cpu_usage: 78.1 }
]);
// Nested Sub-table
await metrics.createSubtable('hourly', [
{ name: 'metric_id', type: 'INT64' },
{ name: 'val', type: 'FLOAT64' }
]);
const hourly = metrics.subtable('hourly');
await hourly.insert([{ metric_id: 1, val: 41.2 }]);
// Zero-Memory File Streaming
await metrics.exportToFile('metrics_stream.csv', 'csv');
await metrics.importFile('metrics_stream.csv', 'csv');
// Analytical Query
const summary = await client.sql`SELECT host, AVG(cpu_usage) FROM analytics.metrics GROUP BY host`;
console.table(summary.rows);
}
main().catch(console.error);
3. Mergen CLI REPL
# Launch interactive REPL
mergen
# Inside REPL:
SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0), (2, 'Bob', 82000.0);
SELECT * FROM employees;
EXPORT employees TO CSV;
EXPORT employees TO JSON;
EXPORT employees TO SQL;
Adaptive Compression Pipeline
When writing column blocks, MergenDB analyzes incoming values and applies the most optimal encoding scheme:
- Bit-Packed Booleans: Stores boolean flags at 1 bit per value (8 values per byte).
- Delta / FoR (Frame-of-Reference): Stores sequential numbers as offsets from
min_value, reducing 8-byte integers to 1- or 2-byte deltas. - Block Dictionary Encoding: Optimal for low-cardinality text (gender, country, status). Stores unique strings once in a block dictionary and encodes rows as 1-byte indices.
- Run-Length Encoding (RLE): Collapses repeated identical values into
(count, value)pairs. - Secondary Zlib Compression: Applied to compressed byte streams for secondary compaction.
Performance Benchmarks
Tested on an Intel Core i5 / i7 with 100,000 mixed telemetry records (12 columns: integers, floats, timestamps, statuses, long strings):
| Storage Format | Disk Size | Space Saved | 2-Column Query Disk Read | Peak RAM |
|---|---|---|---|---|
JSON Lines (.jsonl) |
19.5 MB | 0% (Baseline) | 19.5 MB | Unbounded |
SQLite 3 (.db) |
8.1 MB | 58.4% | 8.1 MB (reads full row) | ~30 MB |
MergenDB (.mgdb) |
1.6 MB | 91.5% | 0.29 MB (pruned) | < 15 MB RAM |
- Exact Filter Scan Throughput: ~50,000,000 rows/sec (single core)
- Substring (
LIKE '%term%') Scan: ~10,000,000 rows/sec (single core) - SQL Streaming Import Speed: ~70,000-120,000 rows/sec on standard SSD
Version Changelog & Release Progression
| Version | Milestone | Key Deliverables | Status |
|---|---|---|---|
| v0.5.8 | Analytical SQL & JOINs | In-Memory Hash JOIN (INNER/LEFT), Multi-Column GROUP BY, HAVING |
Released |
| v0.5.9 | Embedded Web UI | Initial Mergen Studio web interface | Released |
| v0.6.0 | Universal Node.js SDK | Zero-dependency Node.js/TypeScript SDK + CLI runner | Released |
| v0.6.1 | phpMyAdmin Overhaul | 100% CLI feature parity in browser and Node.js SDK, 106 tests | Released |
| v0.6.2 | Multi-Runtime & Zero-Dependency | python -m mergendb & npx mergendb runners, autoStart: true, GET /query |
Released |
| v0.6.3 | Official Identity & i18n | Official Logo, strict zero-emoji policy, English/German/Turkish localization | Released |
| v0.6.4 | Hierarchical Architecture | Database containers, nested sub-tables, and streaming export/import | Released |
| v0.6.5 | Zero-Memory Streaming Engine | Zero-Memory Chunked Streaming Engine, crash-free browser upload | Released |
| v0.6.6 | Hierarchical Containers & Streaming | Database containers, sub-tables, and chunked streaming import/export | Released |
| v0.6.7 | Studio Navigation & Table Selection | Restored table selection, top nav sync, i18n fix, Node.js multi-database routing | Released |
| v0.6.8 | Streaming Raw Payload & Export UX | Clean streaming export raw payload, active table indicator in Studio export | Current Release |
Running the Test Suite
MergenDB includes an embedded test suite with 107 comprehensive tests (89 Python unit tests covering storage, compression algorithms, query planning, Bloom filters, hierarchical databases, sub-tables, and analytical joins + 18 end-to-end Node.js SDK integration tests):
# Run Python unit tests via unittest
python -m unittest discover -s tests
# Run Node.js Client SDK integration tests
node sdks/nodejs/test.js
License
Distributed under the MIT License. See LICENSE for details.
Developed by Ugur Turker Kebeci.
Metadata
Release files for mergendb 0.7.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mergendb-0.7.1.tar.gz | 156.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mergendb-0.7.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 288.6 kB
Release files / mergendb-0.7.1.tar.gz
| Download URL | mergendb-0.7.1.tar.gz |
|---|---|
| Size | 156.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
68fc3b1138fbb65a6cdcd02aa51752014c48044b4c572537c1d1cca5f7ad509c
|
|
BLAKE2b-256 checksum How to use checksums |
ae4fb5180c8a52a1fce807f9e4e8507d0ad7bd1cfb3ec5df1c868bb25f4cf659
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.8.7rc1
|
Release files / mergendb-0.7.1-py3-none-any.whl
| Download URL | mergendb-0.7.1-py3-none-any.whl |
|---|---|
| Size | 132.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8d57fffba0c764233c980a8193439d626e216179aff249c9f4d54e35f012c487
|
|
BLAKE2b-256 checksum How to use checksums |
e84f00361a089c0af9784a8b836a150a7536b2690f2274a70b5d457a54baebf3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.8.7rc1
|