Skip to main content

MergenDB Banner

MergenDB

PyPI version npm version Python Versions License: MIT Tests Author

MergenDB is an ultra-compact, high-performance embedded columnar database engine designed to process massive analytical workloads and multi-million row table scans on resource-constrained hardware. It delivers strict zero external runtime dependencies -- requiring no C compilers, no native C++ binaries, and no bulky runtimes across both Python and Node.js.

Whether querying a 10-million row dataset on a 500 MB RAM VPS, analyzing telemetry streams on an edge Raspberry Pi, running real-time analytical reporting in Node.js/TypeScript, or managing hierarchical databases through the browser in Mergen Studio, MergenDB provides columnar speed with bounded memory guarantees.


What Is New in v0.6.7

  1. Mergen Studio Table Selection & Navigation Fix:

    • Restored internationalization engine (EN/DE/TR) and resolved setLanguage reference errors in the browser client.
    • Synchronized top navigation dropdown (Table:) with the active database and active table selections.
    • Fully enabled one-click table browsing and structure views across all databases and subtables.
  2. Node.js SDK Multi-Database Context:

    • TableHandle now properly encapsulates its parent database context, routing all schema, query, update, delete, column mutations, export, and import commands to the intended database.
  3. CLI REPL Absolute Path Export & Terminal Polish:

    • REPL EXPORT now explicitly prints full absolute filesystem paths on export start and finish.
    • Cleaned terminal progress bar carriage-return output to eliminate leftover progress telemetry.

What Is New in v0.6.6

  1. Hierarchical Database Containers & Nested Sub-tables:

    • Organize data just like modern RDBMS platforms: Databases -> Tables -> Nested Sub-tables (e.g. enterprise.employees.engineering).
    • Store root records or partition sub-groups into isolated columnar files while preserving relational hierarchy.
    • Comprehensive SQL support: SHOW DATABASES;, CREATE DATABASE <name>;, DROP DATABASE <name>;, USE <name>;, SHOW TABLES [FROM <name>];.
    • Native dot-notation resolution across Python (mergendb.database(), table.create_subtable()), Node.js (client.database(), table.createSubtable()), CLI, REST server, and Studio Web UI.
  2. Zero-Memory Chunked Streaming Import & Export Engine:

    • Dedicated DataExporter streaming engine utilizing HTTP Chunked Transfer Encoding (Transfer-Encoding: chunked).
    • Tables of arbitrary size stream directly to disk or network sockets without ever accumulating full datasets into memory buffers.
    • Dedicated /import_stream endpoint accepts raw chunked streams in 64 KB blocks directly from network sockets to temporary disk files, eliminating browser V8 heap bloat and preventing GPU/compositor memory crashes.
    • Live 0% to 100% upload progress telemetry with transferred byte counters in Mergen Studio.
  3. Mergen Studio Web Dashboard Enhancements:

    • Complete phpMyAdmin-style tree hierarchy sidebar with collapsible database, table, and sub-table nodes.
    • Visual creation modals: + DB, + Table, and + Sub.
    • Streamlined DOM rendering without dataset string serialization, protecting browser memory.
    • Comprehensive internationalization: English (Default), German (Deutsch), and Turkish (Turkce).
    • Strict Zero-Emoji policy enforced across all interfaces, logs, and documentation.
  4. Node.js & TypeScript SDK 100% Parity:

    • Added DatabaseHandle, database(), listDatabases(), createDatabase(), dropDatabase().
    • Added createSubtable(), subtable(), listSubtables().
    • Added zero-memory file stream operations: exportToFile(destPath) and importFile(filePath) powered by standard library streams (pipe).

Zero-Dependency Installation

MergenDB requires zero external packages or compilers (dependencies: {}). It runs purely on the standard library of Python and Node.js.

Python Engine & CLI

# Install via PyPI
pip install --upgrade mergendb

# Run interactive CLI REPL directly:
python -m mergendb

# Start HTTP server & Mergen Studio:
python -m mergendb serve 8765

# Run embedded diagnostics & hardware benchmark:
python -m mergendb test

Node.js & TypeScript SDK

# Install SDK via npm
npm install mergendb

# Run server or open studio directly via npx:
npx mergendb serve 8765
npx mergendb studio 8765

Requires Python 3.8+ and/or Node.js 16+. Works out-of-the-box on Windows, macOS, Linux, and Docker with zero additional setup.


Why MergenDB? (The Problem with Row Stores)

Traditional embedded databases like SQLite store data row-by-row ([id, name, age, address, notes, ...]). When running analytical queries:

SELECT name, balance FROM users WHERE balance > 1000;

Even though you only care about name and balance, row stores must read every single column of every row off disk -- including massive text fields like address and notes. On a 10-million row database, that translates to gigabytes of unnecessary disk I/O and heavy memory exhaustion.

MergenDB utilizes the columnar approach:

  1. Column-Isolated I/O: Every column is stored and compressed independently. Unqueried columns are never read from disk.
  2. ZoneMap Pruning: Every block records min_value and max_value. If a block cannot contain matching rows, it is skipped with zero disk reads.
  3. 1024-bit Block Bloom Filters: Instant single-pass lookup index skips blocks that do not contain a queried ID, text, or UUID.
  4. Demand-Driven Late Materialization (LazyColumnDict): In multi-column filters (WHERE status = 'ACTIVE' AND balance > 50), MergenDB evaluates status first. If no rows in the block match, balance and all other columns are never decompressed.
  5. Strictly Bounded Memory: Data streams in small, tunable blocks (1,024-8,192 rows). Memory usage stays under 15-20 MB RAM, whether the database is 100 MB or 100 GB.

MergenDB Ecosystem Architecture

+---------------------------------------------------------------------------------+
|                                 CLIENT LAYER                                    |
|   Python Library (mergendb)   |   Node.js / TS SDK   |   Mergen Studio (Web)    |
|   db.find() / db.sql()        |   db.sql`...`        |   phpMyAdmin UI Tree     |
+---------------------------------------+-----------------------------------------+
                                        | HTTP / REST (Zero-Dependency)
+---------------------------------------v-----------------------------------------+
|                                SERVER ENGINE                                    |
|   Multi-threaded HTTP Server  |  Chunked Stream Transfer    |  Progress Stream  |
|   GET /export (Chunked)       |  POST /import_stream (Raw)  |  Hierarchical DB  |
+---------------------------------------+-----------------------------------------+
                                        | Analytical AST / Execution Plans
+---------------------------------------v-----------------------------------------+
|                              ANALYTICAL ENGINE                                  |
|   In-Memory Hash JOINs        |  Multi-Column GROUP BY / HAVING                 |
|   ZoneMap & Bloom Pruning     |  Vectorized Column Evaluation                   |
+---------------------------------------+-----------------------------------------+
                                        | Zero-Copy mmap & Block I/O
+---------------------------------------v-----------------------------------------+
|                        STORAGE & ADAPTIVE COMPRESSION                           |
|   Bit-Packed Booleans         |  Delta / Frame-of-Reference (FoR)               |
|   Block Dictionary Encoding   |  Run-Length Encoding (RLE)                      |
|   Secondary Zlib Stream       |  ZoneMap & Bloom Header (.mgdb)                 |
+---------------------------------------------------------------------------------+

Hierarchical Database & Nested Sub-tables System

MergenDB supports multi-tier hierarchical data management matching traditional relational databases while retaining columnar performance:

[DB] enterprise
 |-- [TBL] departments (1,200 rows)
 \-- [TBL] employees (4,500 rows)
      |-- [SUB] engineering (320 rows)
      \-- [SUB] marketing (150 rows)

Python API

import mergendb

# 1. Create or open database container
enterprise = mergendb.create_database("enterprise")

# 2. Create tables inside database
departments = enterprise.create_table("departments", [
    ("id", "INT64"),
    ("name", "STRING"),
    ("location", "STRING")
])
employees = enterprise.create_table("employees", [
    ("id", "INT64"),
    ("name", "STRING"),
    ("role", "STRING")
])

# 3. Insert records directly
employees.insert([
    {"id": 1, "name": "Alice", "role": "Staff Engineer"},
    {"id": 2, "name": "Bob", "role": "Data Scientist"},
])

# 4. Create nested sub-tables inside a table
engineering = employees.create_subtable("engineering", [
    ("employee_id", "INT64"),
    ("project_code", "STRING"),
    ("clearance_level", "INT32")
])
engineering.insert([{"employee_id": 1, "project_code": "ATLAS", "clearance_level": 4}])

# 5. Access via dot-notation
tbl = mergendb.connect("enterprise.employees.engineering")
results = tbl.find(employee_id=1)
print(results)

SQL Commands

SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
SHOW TABLES;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0);
SELECT * FROM employees;

Zero-Memory Streaming Import & Export

When handling multi-gigabyte datasets, traditional engines often buffer entire payloads into RAM, leading to memory exhaustion and browser compositor crashes. MergenDB resolves this with true streaming architecture:

1. Chunked Export (GET /export)

The server reads column blocks and yields encoded byte chunks directly into the HTTP response socket using standard HTTP Chunked Transfer Encoding. Server-side memory usage remains strictly bounded (< 1 MB RAM) regardless of table size.

# Stream table directly to disk
curl -N "http://localhost:8765/export?table=enterprise.employees&format=csv" -o employees.csv
curl -N "http://localhost:8765/export?table=enterprise.employees&format=json" -o employees.json
curl -N "http://localhost:8765/export?table=enterprise.employees&format=sql" -o employees.sql

2. Zero-Memory Import (POST /import_stream)

Mergen Studio streams the raw native File object directly over an HTTP socket using XMLHttpRequest. The server buffers incoming bytes in 64 KB chunks directly to a temporary file on disk, parses it with DataImporter, and indexes columnar blocks without consuming V8 heap memory.

# Direct zero-memory streaming upload
curl -X POST "http://localhost:8765/import_stream?table=enterprise.employees&format=csv" \
  --data-binary @large_dataset.csv

5-Minute Quickstart

1. Python API

import mergendb

# 1. Connect to table (Auto-created if not exists)
db = mergendb.connect("telemetry.mgdb")

# 2. Insert records (Schema is auto-inferred)
db.insert([
    {"id": 1, "sensor": "TEMP-01", "reading": 23.5, "active": True},
    {"id": 2, "sensor": "TEMP-02", "reading": 28.1, "active": True},
    {"id": 3, "sensor": "TEMP-01", "reading": 24.0, "active": False},
])

# 3. Document-Style Queries
active_sensors = db.find(sensor="TEMP-01", active=True)
first = db.find_one(sensor="TEMP-02")
matches = db.search("TEMP")  # Substring search across all text columns

# 4. Analytical SQL
res = db.sql("SELECT sensor, COUNT(*), AVG(reading) FROM telemetry GROUP BY sensor")
res.show()

# 5. Zero-Memory File Streaming
db.export_csv("telemetry_export.csv")
mergendb.from_csv("telemetry_export.csv", "backup.mgdb")

# 6. Live Hardware Diagnostics
mergendb.benchmark()

2. Node.js & TypeScript SDK

const { connect } = require('mergendb');

async function main() {
  const client = connect('http://localhost:8765');

  // Database Container Operations
  await client.createDatabase('analytics');
  const analyticsDb = client.database('analytics');

  // Create Table
  await analyticsDb.createTable('metrics', [
    { name: 'id', type: 'INT64' },
    { name: 'host', type: 'STRING' },
    { name: 'cpu_usage', type: 'FLOAT64' }
  ]);

  const metrics = analyticsDb.table('metrics');
  await metrics.insert([
    { id: 1, host: 'prod-api-1', cpu_usage: 42.5 },
    { id: 2, host: 'prod-api-2', cpu_usage: 78.1 }
  ]);

  // Nested Sub-table
  await metrics.createSubtable('hourly', [
    { name: 'metric_id', type: 'INT64' },
    { name: 'val', type: 'FLOAT64' }
  ]);
  const hourly = metrics.subtable('hourly');
  await hourly.insert([{ metric_id: 1, val: 41.2 }]);

  // Zero-Memory File Streaming
  await metrics.exportToFile('metrics_stream.csv', 'csv');
  await metrics.importFile('metrics_stream.csv', 'csv');

  // Analytical Query
  const summary = await client.sql`SELECT host, AVG(cpu_usage) FROM analytics.metrics GROUP BY host`;
  console.table(summary.rows);
}

main().catch(console.error);

3. Mergen CLI REPL

# Launch interactive REPL
mergen

# Inside REPL:
SHOW DATABASES;
CREATE DATABASE enterprise;
USE enterprise;
CREATE TABLE employees (id BIGINT, name TEXT, salary DOUBLE);
INSERT INTO employees VALUES (1, 'Alice', 95000.0), (2, 'Bob', 82000.0);
SELECT * FROM employees;
EXPORT employees TO CSV;
EXPORT employees TO JSON;
EXPORT employees TO SQL;

Adaptive Compression Pipeline

When writing column blocks, MergenDB analyzes incoming values and applies the most optimal encoding scheme:

  1. Bit-Packed Booleans: Stores boolean flags at 1 bit per value (8 values per byte).
  2. Delta / FoR (Frame-of-Reference): Stores sequential numbers as offsets from min_value, reducing 8-byte integers to 1- or 2-byte deltas.
  3. Block Dictionary Encoding: Optimal for low-cardinality text (gender, country, status). Stores unique strings once in a block dictionary and encodes rows as 1-byte indices.
  4. Run-Length Encoding (RLE): Collapses repeated identical values into (count, value) pairs.
  5. Secondary Zlib Compression: Applied to compressed byte streams for secondary compaction.

Performance Benchmarks

Tested on an Intel Core i5 / i7 with 100,000 mixed telemetry records (12 columns: integers, floats, timestamps, statuses, long strings):

Storage Format Disk Size Space Saved 2-Column Query Disk Read Peak RAM
JSON Lines (.jsonl) 19.5 MB 0% (Baseline) 19.5 MB Unbounded
SQLite 3 (.db) 8.1 MB 58.4% 8.1 MB (reads full row) ~30 MB
MergenDB (.mgdb) 1.6 MB 91.5% 0.29 MB (pruned) < 15 MB RAM
  • Exact Filter Scan Throughput: ~50,000,000 rows/sec (single core)
  • Substring (LIKE '%term%') Scan: ~10,000,000 rows/sec (single core)
  • SQL Streaming Import Speed: ~70,000-120,000 rows/sec on standard SSD

Version Changelog & Release Progression

Version Milestone Key Deliverables Status
v0.5.8 Analytical SQL & JOINs In-Memory Hash JOIN (INNER/LEFT), Multi-Column GROUP BY, HAVING Released
v0.5.9 Embedded Web UI Initial Mergen Studio web interface Released
v0.6.0 Universal Node.js SDK Zero-dependency Node.js/TypeScript SDK + CLI runner Released
v0.6.1 phpMyAdmin Overhaul 100% CLI feature parity in browser and Node.js SDK, 106 tests Released
v0.6.2 Multi-Runtime & Zero-Dependency python -m mergendb & npx mergendb runners, autoStart: true, GET /query Released
v0.6.3 Official Identity & i18n Official Logo, strict zero-emoji policy, English/German/Turkish localization Released
v0.6.4 Hierarchical Architecture Database containers, nested sub-tables, and streaming export/import Released
v0.6.5 Zero-Memory Streaming Engine Zero-Memory Chunked Streaming Engine, crash-free browser file upload with live % progress bar, 107 tests (100% pass) Current Release

Running the Test Suite

MergenDB includes an embedded test suite with 107 comprehensive tests (89 Python unit tests covering storage, compression algorithms, query planning, Bloom filters, hierarchical databases, sub-tables, and analytical joins + 18 end-to-end Node.js SDK integration tests):

# Run Python unit tests via unittest
python -m unittest discover -s tests

# Run Node.js Client SDK integration tests
node sdks/nodejs/test.js

License

Distributed under the MIT License. See LICENSE for details.

Developed by Ugur Turker Kebeci.

Metadata

Release files for mergendb 0.6.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mergendb 0.6.7
File Size Uploaded
mergendb-0.6.7.tar.gz 135.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mergendb 0.6.7
File Interpreter ABI Platform
mergendb-0.6.7-py3-none-any.whl Python 3 none any Details

Total release size: 259.6 kB

Release files / mergendb-0.6.7.tar.gz

Download URL mergendb-0.6.7.tar.gz
Size 135.5 kB
Tags Source
SHA-256 checksum
How to use checksums
c8ba1eafd7273cf078f40188b4c60a0952019af0ea2e73e1901819daa67fa011
BLAKE2b-256 checksum
How to use checksums
5ab39c079c66332d90eaf0d432979c18dcf54c63bd33fff9e7355f292947e027
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.7rc1

Release files / mergendb-0.6.7-py3-none-any.whl

Download URL mergendb-0.6.7-py3-none-any.whl
Size 124.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
369e935361361a58ae2d49b2575ab0f5c7dd92ae7a35db146c38e9c9f0b58853
BLAKE2b-256 checksum
How to use checksums
dbc8fa7f30d9a8f62a2d754301eb0078af4351105847779ada99ef7606b7d43b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.7rc1

Release history Release notifications | RSS feed

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.4

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.10

2 release files

0.6.9

2 release files

0.6.8

2 release files

This release

0.6.7 This release

2 release files

0.6.6

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page