servectl
Serve a model file over HTTP in one command, with health and Prometheus metrics built in.
You have a trained model on disk and you want it behind an HTTP endpoint to try
it, wire it into a demo, or scrape its metrics, without writing a FastAPI app
each time. servectl loads the artifact and serves it: a typed /predict, a
/health check, and a Prometheus /metrics endpoint, ready to scrape.
$ servectl serve model.joblib --port 8000
servectl: serving 'model' on http://127.0.0.1:8000
$ curl -s localhost:8000/predict -d '{"instances": [[5.1, 3.5, 1.4, 0.2]]}'
{"predictions": [0]}
Install
$ pip install servectl # from PyPI, once released
$ pip install git+https://github.com/jmweb-org/servectl # latest, available now
Loads any joblib/pickle artifact that exposes a scikit-learn-style predict
(and optionally predict_proba).
Usage
$ servectl serve model.joblib # serve on 127.0.0.1:8000
$ servectl serve model.joblib --host 0.0.0.0 --port 9000
$ servectl info model.joblib # inspect without serving
Endpoints
| Method | Path | Purpose |
|---|---|---|
| POST | /predict |
{"instances": [[...], ...]} to {"predictions": [...]} |
| POST | /predict_proba |
Class probabilities, if the model supports it |
| GET | /health |
Model name, feature count and version |
| GET | /metrics |
Prometheus exposition format |
The request body is validated: instances must be a non-empty list of equal-
length numeric rows. A bad request returns 400 with a message, not a stack
trace.
Metrics
The /metrics endpoint exposes:
servectl_requests_total{endpoint, outcome}— request count per endpoint and ok/error.servectl_predictions_total— total prediction instances served.servectl_predict_seconds— prediction latency histogram.
Each server uses its own registry, so the counters reflect only that process.
Prometheus scrape example
If servectl is running on port 8000, add a scrape job like this to
prometheus.yml:
scrape_configs:
- job_name: "servectl"
metrics_path: "/metrics"
static_configs:
- targets: ["localhost:8000"]
For multiple model servers, give each target a stable label so dashboards can separate them:
scrape_configs:
- job_name: "servectl"
metrics_path: "/metrics"
static_configs:
- targets: ["iris-api:8000"]
labels:
model: "iris"
- targets: ["churn-api:8000"]
labels:
model: "churn"
Grafana panel queries
Use these PromQL queries for a basic dashboard:
| Panel | PromQL |
|---|---|
| Request rate | sum by (endpoint, outcome) (rate(servectl_requests_total[5m])) |
| Prediction throughput | rate(servectl_predictions_total[5m]) |
| p95 prediction latency | histogram_quantile(0.95, sum by (le) (rate(servectl_predict_seconds_bucket[5m]))) |
For per-model latency when you added a model scrape label, group the histogram
by both model and le:
histogram_quantile(
0.95,
sum by (model, le) (rate(servectl_predict_seconds_bucket[5m]))
)
Scope
servectl is for trying a model, demos, and internal services. It does no
authentication, batching, or autoscaling. For a hardened deployment, put it
behind a reverse proxy or use a full serving stack; for a quick, observable
endpoint from a model file, this is one command.
License
MIT. See LICENSE.
Metadata
Release files for servectl 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| servectl-0.2.0.tar.gz | 9.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| servectl-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.0 kB
Release files / servectl-0.2.0.tar.gz
| Download URL | servectl-0.2.0.tar.gz |
|---|---|
| Size | 9.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
17710cc8acafafc4ffb2b527707626c5a331ba3dc5f11b04200f518ab50f8e0a
|
|
BLAKE2b-256 checksum How to use checksums |
c21d2fbca0b900870b7893bc060e53d920d6ce1f786abe884793724d4cd70b19
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.20 {"installer":{"name":"uv","version":"0.11.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / servectl-0.2.0-py3-none-any.whl
| Download URL | servectl-0.2.0-py3-none-any.whl |
|---|---|
| Size | 8.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b678d2d60001491af6615346daad6ce73c44ce9eacf70ec0877e2501c2546aa2
|
|
BLAKE2b-256 checksum How to use checksums |
3705458745461fb79c584c3b98223285483f6f0f7f527f37ab28117a54d40845
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.20 {"installer":{"name":"uv","version":"0.11.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|