Lightweight ML experiment tracker — log, compare and visualize your ML experiments locally
Project description
PyMLens 🧪
A lightweight ML experiment tracking tool that helps data scientists log, compare, and visualize their model experiments — runs fully locally on your machine, no cloud required.
Installation
pip install pymlens
Why PyMLens?
Managing multiple experiments manually becomes chaotic and time-consuming. After running several models, it becomes difficult to track which model performed best and on which problem. PyMLens eliminates this by automatically recording all your experiments in a local SQLite database and giving you a rich Streamlit dashboard to explore results.
Features
- ✅ Minimal code changes — wrap your existing training code in a
withblock - 🔒 Fully local — data stored in
~/.pymlens/experiments.db(SQLite), nothing leaves your machine - 📊 Classification & Regression support
- 🏷️ Custom experiment keywords — label individual runs with
exp_keyword - 📈 Train accuracy tracking — detects overfitting by comparing train vs. validation scores
- 🧮 Confusion matrix — stored and visualized per model for classification
- 🔁 Cross-validation — enabled by default for classification, optional for regression
- 🧬 AI DNA Report — powered by Groq (LLaMA 3.1), gives per-model analysis with a best-model verdict
- 🌈 Dynamic themes — randomize the full dashboard appearance on demand
- 📋 Copy hyperparameters — inspect and copy any model's params directly from the dashboard
- 🌀 Interactive Sunburst explorer — drill down from experiment → model → metric
Supported Metrics
Classification
| Metric | Description |
|---|---|
| Accuracy | Validation accuracy |
| Train Accuracy | Training accuracy (overfitting check) |
| Precision | Weighted precision |
| Recall | Weighted recall |
| F1 Score | Weighted F1 |
| Cross Validation Score | Mean CV score (3-fold, enabled by default) |
| Confusion Matrix | Stored as JSON, visualized as heatmap |
Regression
| Metric | Description |
|---|---|
| MSE | Mean Squared Error |
| MAE | Mean Absolute Error |
| RMSE | Root Mean Squared Error |
| R2 | R² Score |
| Cross Validation Score | Mean CV score using neg_mean_squared_error (opt-in) |
Run Locally
Clone the project:
git clone https://github.com/munishmalhotra6230/model_tracker-MLENS-.git
cd model_tracker-MLENS-
Install dependencies:
pip install -r requirements.txt
Run the dashboard:
python -m pymlens dashboard
# or
pymlens dashboard
🐳 Docker
Build the image
docker build -t pymlens .
Run the dashboard
docker run -p 8501:8501 \
-v pymlens_data:/root/.pymlens \
-e GROQ_API_KEY=your_groq_api_key_here \
pymlens
Then open http://localhost:8501 in your browser.
| Flag | Purpose |
|---|---|
-p 8501:8501 |
Maps container port to your machine |
-v pymlens_data:/root/.pymlens |
Persists your experiment DB across restarts |
-e GROQ_API_KEY=... |
Passes your Groq key securely (for Critics page) |
Leave out
-e GROQ_API_KEYif you don't need the AI Critics page.
Using docker-compose (recommended)
Create a docker-compose.yml in your project:
services:
pymlens:
build: .
ports:
- "8501:8501"
volumes:
- pymlens_data:/root/.pymlens
environment:
- GROQ_API_KEY=${GROQ_API_KEY}
volumes:
pymlens_data:
Then run:
# Start
docker compose up
# Start in background
docker compose up -d
# Stop
docker compose down
Pass your key without hardcoding it:
# Windows PowerShell
$env:GROQ_API_KEY="your_key_here"; docker compose up
# Linux / Mac
GROQ_API_KEY="your_key_here" docker compose up
Usage — Classification
from pymlens import Classification_Experiment
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
x, y = load_iris(return_X_y=True)
xtrain, xval, ytrain, yval = train_test_split(x, y, test_size=0.2, random_state=42)
with Classification_Experiment("Iris_Classification", xtrain, ytrain, xval, yval) as exp:
exp.Start_experiment(model=LogisticRegression(), exp_keyword="Logistic_reg", cross_val=True)
exp.Start_experiment(model=RandomForestClassifier(), exp_keyword="RF_baseline", cross_val=True)
exp.Start_experiment(model=SVC(), exp_keyword="SVM_rbf", cross_val=True)
Note:
cross_val=Trueis the default for classification. Setcross_val=Falseto skip cross-validation and speed up training.
Usage — Regression
from pymlens import Regression_Experiment
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.linear_model import LinearRegression
x, y = fetch_california_housing(return_X_y=True)
xtrain, xval, ytrain, yval = train_test_split(x, y, test_size=0.2, random_state=42)
with Regression_Experiment("House_Price", xtrain, ytrain, xval, yval) as exp:
exp.Start_experiment(model=LinearRegression(), exp_keyword="Linear_baseline")
exp.Start_experiment(model=RandomForestRegressor(), exp_keyword="RF_regressor", cross_val=True)
exp.Start_experiment(model=GradientBoostingRegressor(), exp_keyword="GBR_v1", cross_val=True)
Note:
cross_val=Falseis the default for regression (usesneg_mean_squared_errorscoring when enabled).
Start_experiment Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
model |
sklearn estimator | required | Any scikit-learn compatible model |
exp_keyword |
str |
None |
Custom label for this run (falls back to class name) |
cross_val |
bool |
True (Classification) / False (Regression) |
Enable/disable 3-fold cross-validation |
How Data Is Stored
All results are persisted in a local SQLite database at:
~/.pymlens/experiments.db
Three tables are used:
Experiments— experiment name + timestampScores— classification results per model (accuracy, precision, recall, F1, CV score, confusion matrix, params)Regression_Scores— regression results per model (MSE, MAE, R2, RMSE, CV score, params)
Re-running an experiment with the same name replaces existing results for that model (upsert behavior).
Dashboard Pages
Launch with:
pymlens dashboard
The dashboard has 3 pages:
1. 📊 Model Comparison
- Leaderboard table sorted by F1 Score (classification) or R² (regression)
- Grouped bar chart comparing all metrics across models
- Radar (spider) chart for multi-metric comparison
- Precision vs. Recall scatter plot (classification) / MSE vs. MAE scatter (regression)
- Cross-validation stability bar chart
- Interactive confusion matrix heatmap (classification)
- Copy model hyperparameters as JSON
2. 🌀 Infographics — Sunburst Explorer
- Drill-down from All Experiments → Experiment → Model → Metric
- Hover to see all metric values per model
- Supports both classification and regression data
3. 🧬 Critics — AI DNA Report (Groq-powered)
- Sends model metrics to LLaMA 3.1 (8B) via Groq API
- Generates a structured per-model report:
- Score interpretation, overfitting analysis, CV stability
- One specific improvement recommendation per model
- Final VERDICT — best model with justification
- Requires a Groq API key (see setup below)
AI DNA Report Setup
The Critics page requires a free Groq API key.
Save it using the built-in settings utility:
from pymlens import Pymlens_settings
settings = Pymlens_settings()
settings.add_api_key() # prompts you to paste your Groq key
This saves your key to pymlens/.streamlit/secrets.toml so the dashboard can load it securely.
Managing Your Database
Use Pymlens_settings to manage stored data:
from pymlens import Pymlens_settings
settings = Pymlens_settings()
settings.delete_db() # interactive confirmation + 6-digit code required
⚠️
delete_db()clears all experiments from the database. A random 6-digit confirmation code is required to prevent accidental deletion.
Screenshots
Dependencies
| Package | Purpose |
|---|---|
scikit-learn |
Model training & metric computation |
numpy |
Numerical operations |
pandas |
Data manipulation |
streamlit |
Dashboard UI |
plotly |
Interactive charts |
groq |
AI DNA Report via LLaMA 3.1 |
sqlite3 |
Local experiment storage (stdlib) |
Feedback & Community
Have suggestions or found a bug? Join the Discord:
Links
Author
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pymlens-1.1.6.tar.gz.
File metadata
- Download URL: pymlens-1.1.6.tar.gz
- Upload date:
- Size: 16.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b87840e42849d64b07de4ef8c850122cfe67b1a1eada9a1a60d9e3cfd140d041
|
|
| MD5 |
5339cd832df0d58722cf13df8800ca1c
|
|
| BLAKE2b-256 |
6972a7b022de9b8d81ec09274e2d0f33eec4ea97f2b7a94351fe9889f8856e0e
|
File details
Details for the file pymlens-1.1.6-py3-none-any.whl.
File metadata
- Download URL: pymlens-1.1.6-py3-none-any.whl
- Upload date:
- Size: 13.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d73c63709c01d206b712f675cc02c03b18f2530b2b9144da6e1cb5223cce6fb2
|
|
| MD5 |
b4a905a4b4036ce8a15ec680b28d721f
|
|
| BLAKE2b-256 |
c7d2a854e4e18ac3466f39b8d54468b34a825385a12be1e04cb24de8659ca76a
|