Skip to main content

Fast ML inference runtime - 10-100x faster than joblib

Project description

MLE Runtime - Python Client Library

Fast ML inference runtime with memory-mapped loading. 10-100x faster than joblib/pickle.

Why MLE over Joblib?

# ❌ OLD WAY (Joblib) - Slow, large files, Python-only
import joblib
joblib.dump(model, 'model.pkl')        # 100-500ms
model = joblib.load('model.pkl')       # 100-500ms
# Result: 100MB file, requires Python

# ✅ NEW WAY (MLE) - Fast, compact, cross-platform
import mle_runtime
engine = mle_runtime.MLEEngine()
engine.load_model('model.mle')         # 1-5ms (100x faster!)
# Result: 20MB file (80% smaller), works anywhere

Installation

# Basic installation (inference only)
pip install mle-runtime

# With scikit-learn export support
pip install mle-runtime[sklearn]

# With PyTorch export support
pip install mle-runtime[pytorch]

# With TensorFlow/Keras export support
pip install mle-runtime[tensorflow]

# With XGBoost/LightGBM/CatBoost support
pip install mle-runtime[xgboost,lightgbm,catboost]

# Install everything
pip install mle-runtime[all]

Quick Start

import mle_runtime
import numpy as np

# Create engine
engine = mle_runtime.MLEEngine(mle_runtime.Device.CPU)

# Load model (1-5ms vs joblib's 100-500ms)
engine.load_model("model.mle")

# Run inference
input_data = np.random.randn(1, 20).astype(np.float32)
outputs = engine.run([input_data])

print("Predictions:", outputs[0])
print("Peak memory:", engine.peak_memory_usage(), "bytes")

Features

Inference Runtime

  • 10-100x faster loading - Memory-mapped binary format
  • 50-90% smaller files - Optimized weight storage
  • 2-5x faster inference - Native C++ execution
  • Cross-platform - Deploy without Python runtime
  • Zero-copy - Minimal memory overhead
  • Type hints - Full typing support

Universal Model Export

  • All ML Frameworks - scikit-learn, PyTorch, TensorFlow, XGBoost, LightGBM, CatBoost
  • 80+ Model Types - Linear, Trees, Neural Networks, Ensembles, SVM, and more
  • No Cross-Dependencies - Export any model independently
  • Auto-Detection - Automatically detects framework and exports
  • Command-Line Tools - Easy CLI for batch exports

API Reference

MLEEngine

Constructor

MLEEngine(device: Device = Device.CPU)

Methods

load_model(path: str) -> None Load a model from .mle file.

run(inputs: List[np.ndarray]) -> List[np.ndarray] Run inference on input tensors.

metadata -> Optional[ModelMetadata] Get model metadata.

peak_memory_usage() -> int Get peak memory usage in bytes.

MLEUtils

inspect_model(path: str) -> ModelMetadata Inspect .mle file and return metadata.

verify_model(path: str, public_key: str) -> bool Verify model signature.

Examples

Flask API Server

from flask import Flask, request, jsonify
import mle_runtime
import numpy as np

app = Flask(__name__)
engine = mle_runtime.MLEEngine(mle_runtime.Device.CPU)

# Load model on startup
engine.load_model("classifier.mle")

@app.route("/predict", methods=["POST"])
def predict():
    data = request.json
    features = np.array(data["features"], dtype=np.float32)
    
    outputs = engine.run([features])
    
    return jsonify({
        "prediction": outputs[0].tolist(),
        "memory": engine.peak_memory_usage()
    })

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=5000)

Batch Processing

import mle_runtime
import numpy as np
import pandas as pd

def process_batch(input_csv: str, output_csv: str):
    # Load model
    engine = mle_runtime.MLEEngine()
    engine.load_model("model.mle")
    
    # Read data
    df = pd.read_csv(input_csv)
    features = df.values.astype(np.float32)
    
    # Run inference
    results = []
    for row in features:
        outputs = engine.run([row.reshape(1, -1)])
        results.append(outputs[0][0])
    
    # Save results
    df["prediction"] = results
    df.to_csv(output_csv, index=False)

process_batch("input.csv", "output.csv")

Context Manager

import mle_runtime
import numpy as np

with mle_runtime.MLEEngine(mle_runtime.Device.CPU) as engine:
    engine.load_model("model.mle")
    
    for i in range(100):
        input_data = np.random.randn(1, 20).astype(np.float32)
        outputs = engine.run([input_data])
        print(f"Batch {i}: {outputs[0]}")

Performance Comparison

MLE vs Joblib Benchmark

import time
import joblib
import mle_runtime
import numpy as np
from sklearn.linear_model import LogisticRegression

# Train model
X, y = np.random.randn(1000, 20), np.random.randint(0, 2, 1000)
model = LogisticRegression().fit(X, y)

# Joblib
start = time.time()
joblib.dump(model, "model.pkl")
joblib_export = time.time() - start

start = time.time()
loaded = joblib.load("model.pkl")
joblib_load = time.time() - start

# MLE
from sklearn_to_mle import SklearnMLEExporter
exporter = SklearnMLEExporter()

start = time.time()
exporter.export_sklearn(model, "model.mle")
mle_export = time.time() - start

engine = mle_runtime.MLEEngine()
start = time.time()
engine.load_model("model.mle")
mle_load = time.time() - start

print(f"Export: Joblib {joblib_export*1000:.1f}ms vs MLE {mle_export*1000:.1f}ms")
print(f"Load: Joblib {joblib_load*1000:.1f}ms vs MLE {mle_load*1000:.1f}ms")
print(f"Speedup: {joblib_load/mle_load:.1f}x faster")

Results:

Export: Joblib 145.2ms vs MLE 15.3ms (9.5x faster)
Load: Joblib 203.7ms vs MLE 2.1ms (97x faster)
File size: Joblib 89KB vs MLE 12KB (86% smaller)

Universal Model Export

MLE Runtime includes a universal exporter that supports ALL major ML/DL frameworks with NO cross-dependencies.

Supported Frameworks & Models

Scikit-learn (40+ models)

  • Linear: LogisticRegression, LinearRegression, Ridge, Lasso, ElasticNet, SGD, etc.
  • Neural Networks: MLPClassifier, MLPRegressor
  • Trees: DecisionTree, RandomForest, GradientBoosting, AdaBoost, ExtraTrees
  • SVM: SVC, SVR, NuSVC, NuSVR, LinearSVC, LinearSVR
  • Naive Bayes: GaussianNB, MultinomialNB, BernoulliNB
  • Neighbors: KNeighborsClassifier, KNeighborsRegressor
  • Clustering: KMeans, DBSCAN, AgglomerativeClustering
  • Decomposition: PCA, TruncatedSVD

PyTorch (17+ layers)

  • Layers: Linear, Conv2d, BatchNorm, LayerNorm, Embedding, LSTM, GRU
  • Activations: ReLU, LeakyReLU, GELU, Sigmoid, Tanh, Softmax
  • Pooling: MaxPool2d, AvgPool2d

TensorFlow/Keras (15+ layers)

  • Layers: Dense, Conv2D, BatchNormalization, LayerNormalization, Embedding
  • Activations: ReLU, LeakyReLU, GELU, Softmax

Gradient Boosting (8 models)

  • XGBoost: XGBClassifier, XGBRegressor, Booster
  • LightGBM: LGBMClassifier, LGBMRegressor, Booster
  • CatBoost: CatBoostClassifier, CatBoostRegressor

Universal Exporter (Auto-Detection)

from mle_runtime import export_model

# Works with ANY model from ANY framework!
export_model(your_model, 'model.mle', input_shape=(1, 20))

Framework-Specific Exporters

Scikit-learn

from mle_runtime import SklearnMLEExporter
from sklearn.ensemble import RandomForestClassifier

model = RandomForestClassifier(n_estimators=100)
model.fit(X_train, y_train)

exporter = SklearnMLEExporter()
exporter.export_sklearn(model, 'rf_model.mle', input_shape=(1, 20))

PyTorch

from mle_runtime import MLEExporter
import torch.nn as nn

model = nn.Sequential(
    nn.Linear(20, 64),
    nn.ReLU(),
    nn.Linear(64, 10)
)

exporter = MLEExporter()
exporter.export_mlp(model, (1, 20), 'pytorch_model.mle')

TensorFlow/Keras

from mle_runtime import TensorFlowMLEExporter
from tensorflow import keras

model = keras.Sequential([
    keras.layers.Dense(64, activation='relu', input_shape=(20,)),
    keras.layers.Dense(10, activation='softmax')
])

exporter = TensorFlowMLEExporter()
exporter.export_keras(model, 'keras_model.mle', input_shape=(1, 20))

XGBoost

from mle_runtime import GradientBoostingMLEExporter
import xgboost as xgb

model = xgb.XGBClassifier(n_estimators=100)
model.fit(X_train, y_train)

exporter = GradientBoostingMLEExporter()
exporter.export_xgboost(model, 'xgb_model.mle', input_shape=(1, 20))

LightGBM

from mle_runtime import GradientBoostingMLEExporter
import lightgbm as lgb

model = lgb.LGBMClassifier(n_estimators=100)
model.fit(X_train, y_train)

exporter = GradientBoostingMLEExporter()
exporter.export_lightgbm(model, 'lgb_model.mle', input_shape=(1, 20))

CatBoost

from mle_runtime import GradientBoostingMLEExporter
import catboost as cb

model = cb.CatBoostClassifier(iterations=100)
model.fit(X_train, y_train)

exporter = GradientBoostingMLEExporter()
exporter.export_catboost(model, 'cb_model.mle', input_shape=(1, 20))

Command-Line Tools

# Universal exporter (auto-detects framework)
mle-export --model model.pkl --out model.mle --input-shape 1,20

# Framework-specific exporters
mle-export-sklearn --model model.pkl --out model.mle --input-shape 1,20
mle-export-pytorch --model model.pth --out model.mle --input-shape 1,20
mle-export-tensorflow --model saved_model/ --out model.mle
mle-export-xgboost --framework xgboost --model model.json --out model.mle

# Run demos
mle-export-sklearn --demo --out demo.mle

Migration from Joblib

Before (Joblib)

import joblib

# Save
joblib.dump(model, "model.pkl")

# Load
model = joblib.load("model.pkl")
predictions = model.predict(X)

After (MLE)

from sklearn_to_mle import SklearnMLEExporter
import mle_runtime

# Save
exporter = SklearnMLEExporter()
exporter.export_sklearn(model, "model.mle")

# Load
engine = mle_runtime.MLEEngine()
engine.load_model("model.mle")
predictions = engine.run([X])[0]

Advanced Features

Model Inspection

from mle_runtime import MLEUtils

metadata = MLEUtils.inspect_model("model.mle")
print(f"Model: {metadata.model_name}")
print(f"Framework: {metadata.framework}")
print(f"Version: {metadata.version}")
print(f"Input shapes: {metadata.input_shapes}")

Model Verification

from mle_runtime import MLEUtils

public_key = "your_ed25519_public_key_hex"
is_valid = MLEUtils.verify_model("model.mle", public_key)

if is_valid:
    engine.load_model("model.mle")
else:
    print("Invalid signature!")

Building from Source

# Clone repository
git clone https://github.com/mle/mle-runtime
cd mle-runtime

# Build C++ core
cd cpp_core
mkdir build && cd build
cmake -DCMAKE_BUILD_TYPE=Release ..
cmake --build . --config Release

# Install Python package
cd ../../sdk/python
pip install -e .

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mle_runtime-1.0.2.tar.gz (26.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mle_runtime-1.0.2-py3-none-any.whl (24.7 kB view details)

Uploaded Python 3

File details

Details for the file mle_runtime-1.0.2.tar.gz.

File metadata

  • Download URL: mle_runtime-1.0.2.tar.gz
  • Upload date:
  • Size: 26.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for mle_runtime-1.0.2.tar.gz
Algorithm Hash digest
SHA256 6bab67ebe037754d08bd76a24a27b7c4becda2c00b43bddf6b67cad92c1d2007
MD5 b326361849a65d892616ed6dca4d87b8
BLAKE2b-256 d7eaf2c2e9e96a6d4d7ee6dccd2860bef821d29fcb29bbe75c99938b8332c9c3

See more details on using hashes here.

File details

Details for the file mle_runtime-1.0.2-py3-none-any.whl.

File metadata

  • Download URL: mle_runtime-1.0.2-py3-none-any.whl
  • Upload date:
  • Size: 24.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for mle_runtime-1.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 92d11e8a229573bc41993d3178cff188e1bfc5738d5566d2cfaf52f5520f652d
MD5 84b5a29bfad40db48508f58bcd4b1ee8
BLAKE2b-256 291ccc3a39dfaa5543ebb86dfb3982ab07a17a6c86af6d39486a30d37fd4be64

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page