Skip to main content

MariaDB Binlog Indexer for Faster Querying

Project description

MariaDB Binlog Indexer

This tool helps to index mariadb binlogs and store them in compact format to query faster.

File Format

  • DuckDB - Store all the metadata of binlogs.
  • Parquet - Store the actual queries.

DB Schema

┌─────────────┬─────────────┬─────────┬─────────┬─────────┬─────────┐
│ column_name │ column_type │  null   │   key   │ default │  extra  │
│   varchar   │   varchar   │ varchar │ varchar │ varchar │ varchar │
├─────────────┼─────────────┼─────────┼─────────┼─────────┼─────────┤
│ binlog      │ VARCHAR     │ YES     │ NULL    │ NULL    │ NULL    │
│ db_name     │ VARCHAR     │ YES     │ NULL    │ NULL    │ NULL    │
│ table_name  │ VARCHAR     │ YES     │ NULL    │ NULL    │ NULL    │
│ timestamp   │ INTEGER     │ YES     │ NULL    │ NULL    │ NULL    │
│ type        │ VARCHAR     │ YES     │ NULL    │ NULL    │ NULL    │
│ row_id      │ INTEGER     │ YES     │ NULL    │ NULL    │ NULL    │
│ event_size  │ INTEGER     │ YES     │ NULL    │ NULL    │ NULL    │
└─────────────┴─────────────┴─────────┴─────────┴─────────┴─────────┘

Usage

Before starting up please note down,

  • All the timestamps are in seconds UTC
  • Available query types are SELECT, INSERT, UPDATE, DELETE, OTHER

Also, check the python docstrings for more details.

Initialize Indexer

from mariadb_binlog_indexer import Indexer

indexer = Indexer(
    base_path="<path to store indexer related files>",
    db_name="metadata.db", # The name of the database to store metadata
)

Index New Binlog

indexer.add(
    binlog_path="<path to binlog>",
    batch_size=10000, # The batch size to insert binlog in duckdb
)

Remove Indexes of Binlog

indexer.remove(
    binlog_path="<path to binlog>",
)

Generate Timeline

This function will provide a summary of binlog event in a given time range. It will split the time range into 30 parts and provide event related information for each part.

indexer.get_timeline(
    start_timestamp=1746534427, # The start timestamp in seconds UTC
    end_timestamp=1756534427, # The end timestamp in seconds UTC,
    type="INSERT", # Optional
    database="test_db", # Optional
)

Get Row Ids

This function will help to get the row ids of each binlog based on our request. The purpose of this function is to help finding required row ids beforehand to implement pagination on the other end.

Only reason of doing this is to reduce the search in the parquet files.

indexer.get_row_ids(
    start_timestamp=1746534427, # The start timestamp in seconds UTC
    end_timestamp=1756534427, # The end timestamp in seconds UTC,
    type="INSERT", # Optional
    database="test_db", # Optional
    table="test_table", # Optional
    search_str="test", # Optional
)

Get Queries from Parquet Files

This function will help to get the queries from parquet files.

indexer.get_queries(
    row_ids={
        "binlog_1": [101, 102, 103],
        "binlog_2": [104, 105, 106],
    },
    database="test_db", # Optional
)

You need to provide database name as filtering purpose only.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file mariadb_binlog_indexer-0.0.10-py3-none-manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for mariadb_binlog_indexer-0.0.10-py3-none-manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 c1722008dff515d134197e62e4e85d439c5d23ec1fdc7522ce2fa3079f3659c6
MD5 af2e069b615121d8494d9a13c16a78c8
BLAKE2b-256 7bd61cb82c6da9b65d3103c80f50e507a36103972b9b577c3659ba91bee591a1

See more details on using hashes here.

File details

Details for the file mariadb_binlog_indexer-0.0.10-py3-none-manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for mariadb_binlog_indexer-0.0.10-py3-none-manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 274340f2124fc08c6e96f4f33b399d70f8c34c0c9d825a916b451c4acc895b50
MD5 916d1a7d52930b6c366aeddae3fc568b
BLAKE2b-256 cf36dd79d4b1a719261270f1ad8aeb1bb47bf8b950cbe07b3258df56e1cde327

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page