glin
CLI tool and Python library that equips AI agents with an instant, statistical "gut feeling" (System 1 thinking).
glin trains an Explainable Boosting Machine on a CSV, then exposes it to LLM agents over MCP — locally over stdio, or remotely over MCP's standard streamable-http transport. Every prediction comes with an exact, zero-approximation breakdown of which features drove it, straight from the model's own additive structure (no SHAP/LIME approximation).
Install
pip install -e .
Train a model
glin train path/to/data.csv --target churn --name churn_v1
Models are saved under ~/.glin/models/<name>/.
glin list
Use it locally (Claude Desktop, Cursor, ...)
Add to your MCP client's config (e.g. claude_desktop_config.json):
{
"mcpServers": {
"glin": {
"command": "glin",
"args": ["serve", "--mode", "stdio"]
}
}
}
Deploy it remotely (e.g. one EC2 box, any MCP-aware agent)
docker build -t glin .
docker run -p 8000:8000 -v ~/.glin:/root/.glin glin
Then point any MCP client at the standard streamable-http endpoint:
claude mcp add --transport http glin http://<host>:8000/mcp
Tools exposed over MCP
list_models()— all trained models available.inspect_model(model_name)— feature schema and target classes.predict(model_name, features, top_n=10)— predicted class and probabilities, plus the full glassbox audit: base rate, every term's contribution (sorted by magnitude), and an explicit additivity check against the model's own predicted probability.
Data requirements
glin train validates your CSV before doing any work — hard problems (e.g. a target column with only one class) stop training with a clear error; soft issues (e.g. a date-like column) print a warning and training proceeds anyway. These rules live in glin/validation.py as a flat, appendable list, so support for a currently-unsupported shape below can be added by adding one rule and one preprocessing case, without touching the rest of the pipeline.
Feature columns — supported today:
| Data shape | What happens |
|---|---|
Numeric (int/float), including NaN |
Passed through as-is; EBM natively bins missing values into their own split. |
Dirty numeric strings (blanks, e.g. " " for a new customer's total charges) |
Coerced to float64 if ≥80% of non-null values parse as numbers; unparseable cells become NaN. |
| Categorical strings, any cardinality up to 250 unique values | Standardized (lowercased, stripped) and one-hot-style binned by EBM; missing/unseen values map to a __missing__ sentinel. |
| Boolean columns | Treated as a 0/1 continuous numeric feature. |
Feature columns — not yet supported (each is flagged by glin train's validator when detected):
| Data shape | What actually happens | Why |
|---|---|---|
| Dates / timestamps | No date features are extracted. The column is either dropped (if high-cardinality) or kept as a meaningless categorical label — no time-based signal survives either way. | Cut from V1 scope; a real dataset need should drive adding cyclical date-feature extraction. |
| Free text / natural language | Dropped once it exceeds 250 unique values; below that threshold it becomes a set of (almost certainly useless) categorical labels. | glin doesn't do NLP; a text column isn't a set of classes. |
Currency symbols / thousands separators ($, €, £, ,) |
Not stripped. "$1,234.56" fails the ≥80% numeric-parse threshold and falls back to categorical (i.e. garbage). Clean these before training. |
Cut from V1 scope under an "assume a healthy dataset" simplification. |
| Lists / dicts / nested JSON in a cell | Not parsed structurally; treated as an opaque string. | No structured extraction implemented. |
| Row identifiers (sequential IDs, UUIDs, hashes) | Dropped intentionally — not a gap, this is by design. | IDs carry no predictive signal. |
Target column requirements:
- At least 2 distinct non-null values (binary or multiclass) — a single-class target is a hard error.
- Rows with a missing target value are dropped automatically (with a warning); the rest are unaffected.
- A numeric target with many distinct values (>20) triggers a warning that it looks like a regression target —
glintrains classifiers, not regressors.
Release files for glin-ml 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| glin_ml-0.1.0.tar.gz | 17.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| glin_ml-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.0 kB
Release files / glin_ml-0.1.0.tar.gz
| Download URL | glin_ml-0.1.0.tar.gz |
|---|---|
| Size | 17.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3de0a0a359a2787c6ea4783fb36a418abedb77d26c9d8b7630622c14ce18fb92
|
|
BLAKE2b-256 checksum How to use checksums |
45623bd17264c8b4602e90b5db695e888620b237b74fb2a27276a3ef7b47d682
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / glin_ml-0.1.0-py3-none-any.whl
| Download URL | glin_ml-0.1.0-py3-none-any.whl |
|---|---|
| Size | 14.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
abcced265d91072cb06c10509d7093913a606c7a7e68af6f603f3d1118a70246
|
|
BLAKE2b-256 checksum How to use checksums |
90dfa7d9e560b5a26570a93fb571ca89dc29c65ba59b6c7f50ccd768b94a6c8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|