A pure Python/NumPy Machine Learning Library
Project description
simpleml
A lightweight, educational implementation of a scikit-learn-like machine learning library in pure Python and NumPy.
Features
- Classification Models: Logistic Regression, Decision Trees, Random Forests, Naive Bayes, Support Vector Classifiers
- Regression Models: Linear Regression, Decision Tree Regression, Random Forest Regression, Support Vector Regression
- Clustering Algorithms: K-Means, DBSCAN (with working
predict) - Preprocessing: StandardScaler, MinMaxScaler, OneHotEncoder, PolynomialFeatures
- Model Selection: K-Fold Cross-Validation, Grid Search, Cross-Validation
- Metrics: Accuracy, Precision, Recall, F1-Score, Confusion Matrix, ROC-AUC, MSE, MAE, MAPE, R² Score
- Consistent API: Inspired by scikit-learn's fit-predict interface
Requirements
- Python >= 3.9
- NumPy >= 2.0
Note: This library requires NumPy 2.0+. Earlier versions used deprecated APIs (
np.trapz,np.array(copy=False)) that were removed or changed in NumPy 2.0.
Installation
git clone https://github.com/sergeauronss01/simpleml.git
cd simpleml
pip install -e .
The -e flag installs the package in editable mode — changes to the source are
reflected immediately without reinstalling.
Alternatively, if you do not want to install the package, add this to
pyproject.toml so pytest can find the source:
[tool.pytest.ini_options]
pythonpath = ["src"]
Quick Start
Import paths: All models live under
simpleml.models, utilities undersimpleml.utils, and metrics undersimpleml.metrics. The examples below use the correct paths.
Basic Classification
from simpleml.models.linear import LogisticRegression
from simpleml.utils.preprocessing import StandardScaler, train_test_split
from simpleml.metrics.classification import accuracy_score
import numpy as np
# Generate sample data
X = np.random.randn(100, 5)
y = np.random.randint(0, 2, 100)
# Split data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_seed=42)
# Preprocess
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
# Train model
clf = LogisticRegression(n_iter=1000)
clf.fit(X_train, y_train)
# Evaluate
y_pred = clf.predict(X_test)
print(f"Accuracy: {accuracy_score(y_test, y_pred)}")
Linear Regression
from simpleml.models.linear import LinearRegression
from simpleml.metrics.regression import r2_score
reg = LinearRegression(learning_rate=0.01, n_iter=1000, alpha=0.0001)
reg.fit(X_train, y_train)
y_pred = reg.predict(X_test)
print(f"R² Score: {r2_score(y_test, y_pred)}")
Decision Tree Classifier
from simpleml.models.tree import DecisionTreeClassifier
X = np.array([[0, 0], [1, 1], [0, 1], [1, 0]], dtype=float)
y = np.array([0, 1, 1, 0])
clf = DecisionTreeClassifier(max_depth=3, criterion="gini")
clf.fit(X, y)
predictions = clf.predict(X)
Random Forest
from simpleml.models.ensemble import RandomForestClassifier
clf = RandomForestClassifier(n_estimators=10, random_state=42)
clf.fit(X_train, y_train)
predictions = clf.predict(X_test)
Support Vector Machines
from simpleml.models.svm import LinearSVC, SVR
# Classification — predict() returns original class labels (e.g. {0, 1}),
# not the internal {-1, +1} encoding.
svc = LinearSVC(C=1.0, max_iter=1000, random_state=42)
svc.fit(X_train, y_train)
labels = svc.predict(X_test) # returns {0, 1}, not {-1, 1}
# Regression
svr = SVR(C=1.0, epsilon=0.1, max_iter=1000)
svr.fit(X_train, y_train)
values = svr.predict(X_test)
Naive Bayes
from simpleml.models.naive_bayes import GaussianNaiveBayes, MultinomialNaiveBayes
# Gaussian — for continuous features
gnb = GaussianNaiveBayes()
gnb.fit(X_train, y_train)
predictions = gnb.predict(X_test)
# Multinomial — for discrete count features (e.g. word counts)
# Works best when classes differ in *which* features are active,
# not just in their overall magnitude.
mnb = MultinomialNaiveBayes(alpha=1.0) # alpha controls Laplace smoothing
mnb.fit(X_counts_train, y_train)
predictions = mnb.predict(X_counts_test)
K-Means Clustering
from simpleml.models.clustering import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
kmeans.fit(X)
labels = kmeans.predict(X_new)
print(f"Inertia: {kmeans.inertia_}")
DBSCAN Clustering
from simpleml.models.clustering import DBSCAN
db = DBSCAN(eps=0.5, min_samples=5)
db.fit(X)
print(db.labels_) # -1 means noise/outlier
# predict() assigns new points to the nearest core point's cluster
# or returns -1 if no core point is within eps distance.
new_labels = db.predict(X_new)
Preprocessing
from simpleml.utils.preprocessing import (
StandardScaler,
MinMaxScaler,
PolynomialFeatures,
OneHotEncoder,
train_test_split,
)
# Standardise (zero mean, unit variance)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X_train)
# Scale to [0, 1]
mms = MinMaxScaler()
X_minmax = mms.fit_transform(X_train)
# Generate polynomial features
# e.g. [a, b] → [1, a, b, a², ab, b²] for degree=2
pf = PolynomialFeatures(degree=2, include_bias=True)
X_poly = pf.fit_transform(X_train)
# One-hot encode categorical columns
enc = OneHotEncoder()
X_encoded = enc.fit_transform(X_categorical)
# Split — note the parameter is `random_seed`, not `random_state`
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_seed=42
)
Cross-Validation
from simpleml.utils.model_selection import cross_validate, KFold
scores = cross_validate(clf, X, y, cv=5)
print(f"Mean test score: {scores['mean_test_score']:.4f}")
print(f"Std test score: {scores['std_test_score']:.4f}")
print(f"Mean train score: {scores['mean_train_score']:.4f}")
# Custom splitter
kf = KFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_validate(clf, X, y, cv=kf)
Grid Search
from simpleml.utils.model_selection import GridSearchCV
# Note: use `n_iter`, not `n_iterations`
param_grid = {
'learning_rate': [0.01, 0.1],
'n_iter': [100, 500],
}
gs = GridSearchCV(clf, param_grid, cv=5)
gs.fit(X_train, y_train)
print(f"Best parameters: {gs.best_params_}")
print(f"Best CV score: {gs.best_score_:.4f}")
predictions = gs.predict(X_test)
Metrics
from simpleml.metrics.classification import (
accuracy_score,
precision_score,
recall_score,
f1_score,
roc_auc_score,
confusion_matrix,
classification_report,
CrossEntropyLoss,
softmax,
)
from simpleml.metrics.regression import (
mean_squared_error,
mean_absolute_error,
mean_absolute_percentage_error,
r2_score,
)
# Classification
print(accuracy_score(y_true, y_pred))
print(precision_score(y_true, y_pred))
print(recall_score(y_true, y_pred))
print(f1_score(y_true, y_pred))
print(roc_auc_score(y_true, y_probs))
confusion_matrix(y_true, y_pred)
classification_report(y_true, y_pred)
# Regression
print(mean_squared_error(y_true, y_pred))
print(mean_absolute_error(y_true, y_pred))
print(mean_absolute_percentage_error(y_true, y_pred))
print(r2_score(y_true, y_pred))
Module Structure
Scratch-ML-Library/
├── .gitignore
├── LICENSE
├── README.md
├── conftest.py ← shared pytest fixtures and sys.path setup
├── pyproject.toml
├── requirements.txt
├── src/
│ └── simpleml/
│ ├── __init__.py
│ ├── core/
│ │ ├── __init__.py
│ │ └── base.py ← BaseEstimator, RegressorMixin, ClassifierMixin,
│ │ TransformerMixin, ClusterMixin
│ ├── metrics/
│ │ ├── __init__.py
│ │ ├── classification.py ← accuracy, precision, recall, F1, ROC-AUC,
│ │ │ CrossEntropyLoss, softmax, confusion_matrix
│ │ └── regression.py ← MSE, MAE, MAPE, R²
│ ├── models/
│ │ ├── __init__.py
│ │ ├── clustering.py ← KMeans, DBSCAN
│ │ ├── ensemble.py ← RandomForestClassifier, RandomForestRegressor
│ │ ├── linear.py ← LinearRegression, LogisticRegression
│ │ ├── naive_bayes.py ← GaussianNaiveBayes, MultinomialNaiveBayes
│ │ ├── svm.py ← LinearSVC, SVR
│ │ └── tree.py ← DecisionTreeClassifier, DecisionTreeRegressor
│ ├── nn/
│ │ └── __init__.py ← reserved for future neural network support
│ └── utils/
│ ├── __init__.py
│ ├── model_selection.py ← KFold, cross_validate, GridSearchCV
│ ├── preprocessing.py ← StandardScaler, MinMaxScaler,
│ │ PolynomialFeatures, OneHotEncoder,
│ │ train_test_split
│ └── validation.py ← check_array, check_X_y
└── tests/
├── conftest.py ← (generated by pytest from root conftest.py)
├── test_core_base.py
├── test_metrics_classification.py
├── test_metrics_regression.py
├── test_models_clustering.py
├── test_models_ensemble.py
├── test_models_linear.py
├── test_models_naive_bayes.py
├── test_models_svm.py
├── test_models_tree.py
├── test_utils_model_selection.py
├── test_utils_preprocessing.py
└── test_utils_validation.py
API Reference
Parameter names
Some parameter names differ from scikit-learn. The table below lists the ones most likely to cause confusion:
| Class | simpleml parameter | scikit-learn equivalent |
|---|---|---|
LinearRegression |
n_iter |
max_iter |
LogisticRegression |
n_iter |
max_iter |
train_test_split |
random_seed |
random_state |
LinearSVC output labels
LinearSVC.predict() returns the original class labels you trained on
(e.g. {0, 1}). Internally the SVM uses {-1, +1}, but this is mapped back
before returning. You do not need to manually remap predictions.
DBSCAN predict behaviour
DBSCAN.predict() assigns new points to the nearest core point's cluster if
that core point is within eps distance. Points with no core point nearby
are labelled -1 (noise). Core points are identified and stored during fit.
R² on constant targets
r2_score returns 0.0 when all true values are identical (the denominator
SS_tot is zero). This is by convention — do not expect 1.0 for a perfect
predictor on constant data.
DecisionTreeRegressor leaf predictions
With min_samples_split=2 (the default), the tree stops splitting when a
node would produce children with fewer than 2 samples. Leaf nodes predict the
mean of all samples they contain, not individual sample values.
MultinomialNaiveBayes separability
MultinomialNaiveBayes classifies based on relative feature frequency
proportions within each sample. For good accuracy, classes should differ in
which features are active, not just in overall magnitude. Ranges like
[0–2] vs [20–22] may be insufficient; use ranges like [high, ~0] vs
[~0, high] across different feature subsets.
Testing
Ensure the package is installed or src/ is on the path first (see
Installation above), then run:
# All tests, verbose
pytest tests/ -v
# With coverage report
pytest tests/ --cov=simpleml
# Single file
pytest tests/test_models_linear.py -v
# Single test
pytest tests/test_models_linear.py::TestLinearRegressionFit::test_fit_returns_self -v
The test suite contains 339+ tests across 12 files covering:
- Core base classes and mixins
- All model
fit/predict/scoreflows - Edge cases: unfitted models, shape mismatches, single-class folds
- Numerical correctness: convergence, regularisation, impurity metrics
- Preprocessing: scaling, encoding, polynomial expansion, splitting
- Model selection: KFold splits, cross-validation, grid search
Limitations
- No support for neural networks or deep learning (the
nn/module is reserved for future use) - No GPU acceleration
- Limited to basic algorithms — no gradient boosting, stacking, or other advanced ensemble techniques
- Not optimised for large-scale datasets
LinearSVCsupports binary classification only — raisesValueErrorfor more than two classes
Performance
This library is for educational purposes. For production use and performance-critical applications, use scikit-learn.
Contributing
- Fork the repository
- Create a feature branch
- Write tests for new functionality in the
tests/directory - Ensure all tests pass:
pytest tests/ -v - Submit a pull request
License
MIT License
Author
Serge Auronss Gbaguidi
References
Note: This is an educational project created to understand machine learning algorithms and best practices for library design. For production machine learning work, please use scikit-learn.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file simpleml_auronss-0.2.0.tar.gz.
File metadata
- Download URL: simpleml_auronss-0.2.0.tar.gz
- Upload date:
- Size: 41.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ada764ed108ba40e1de563c1b7b5d31e121c6317d3529add42ebb77d20b09d56
|
|
| MD5 |
bb464e2e955a3c78174929d7324b1f63
|
|
| BLAKE2b-256 |
ccafd3a3932a52cbc06cca74154affe4967540f23c8b47fea3bc4e2f4c94f99d
|
File details
Details for the file simpleml_auronss-0.2.0-py3-none-any.whl.
File metadata
- Download URL: simpleml_auronss-0.2.0-py3-none-any.whl
- Upload date:
- Size: 29.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a4b7b434a7bad309d404bb61a0148fc12b443492ce725d12cf4d32983d044719
|
|
| MD5 |
d5b92154d550af14411edcd32fcd02e5
|
|
| BLAKE2b-256 |
26e03ab477c7a2252361f48f0e3efc0e7f0bdd075d69ced868627b94a36576d8
|