laya-mlx-http
HTTP server for laya-mlx typed decision models.
Runs the model once, keeps it in unified memory, and answers POST /v1/predict
requests from anything on your network.
This package is only the server. The API below is plain JSON over HTTP and is
client-agnostic — curl, Python, Node, a sensor, anything. No SDK is required.
The Vercel AI SDK integration is a separate, optional npm package,
@noahwaldner/laya-mlx-http;
it translates AI SDK calls to and from this same API and is irrelevant if you are not
using the AI SDK.
Requires Apple silicon (MLX) and Python 3.11+.
Install
pip install laya-mlx-http # or: uv tool install laya-mlx-http
Run
# localhost only (default, safest)
laya-mlx-http
# expose on your LAN — always set an API key
laya-mlx-http --host 0.0.0.0 --api-key "$(openssl rand -hex 24)"
Then from another machine:
curl http://<server-ip>:8000/health
Endpoints
| Method | Path | Purpose |
|---|---|---|
GET |
/health |
liveness + model readiness (never authenticated) |
POST |
/v1/predict |
{"text": "...", "preset": "triage"} or a full questions map |
GET |
/v1/predict?text=...&preset=triage&flat=true |
same, for simple clients |
If --api-key is set, /v1/* requires X-API-Key: <key> (or Authorization: Bearer <key>).
Predict format
This section is the server's own wire format — what goes over the wire no matter which client sends it. It is not tied to any SDK.
If you are used to the AI SDK adapter, two names differ: its state is text here, and its
boolean is Laya's noul. Both translations happen client-side; the server only ever sees
text/noul. Full mapping:
adapter README.
Request
POST /v1/predict with a JSON body:
| Field | Type | Notes |
|---|---|---|
text |
string | Required. The state to classify (the message, ticket, …). |
questions |
object | Question map (below). Provide this or preset. |
preset |
string | One of the bundled presets; only read when questions is absent. |
flat |
bool | Default false. Adds a flat map of question id → chosen value. |
model |
string | Optional. Echoed back as model in the response instead of the server's model id. |
GET /v1/predict takes the same call as query parameters: text, preset, flat, and
questions as a raw JSON string (?questions={"a":{"type":"noul",...}}, URL-encoded).
You must send questions or preset — otherwise 400.
Question map
{
"department": {
"type": "choice", // choice | score | noul
"instructions": "Which team should handle this?", // required, non-blank
"criteria": { // shape depends on type
"billing": "Charges and refunds",
"support": "Other requests"
}
}
}
instructions— required and must not be blank. A string, or any JSON-serializable object/array (it is serialized to JSON before it reaches the model).type: "choice"—criteriarequired: a non-empty{label: description}object (or an array of unique labels; descriptions may benull). Returns the winning label plus full probabilities.type: "score"—criteriarequired: a non-empty, ordered array of level descriptions (e.g.["calm", "annoyed", "very angry"]). Returns a 0-based expected score.type: "noul"— Laya's boolean.criteriaoptional:{"false": "...", "true": "..."}. Returns a probability in[0, 1]fortrue.
Response
{
"model": "aac6fef/laya-mlx",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.91, "support": 0.09 },
"confidence": 0.91, "answer_confidence": 0.91,
"action": { "act_probability": 0.42 }
},
"frustration": {
"type": "score",
"score": 1.64,
"legend": { "0": "calm", "1": "annoyed", "2": "very angry" },
"probabilities": { "0": 0.1, "1": 0.5, "2": 0.4 },
"confidence": 0.5, "answer_confidence": 0.5,
"action": { "act_probability": 0.42 }
},
"is_urgent": {
"type": "noul",
"noul": 0.42,
"confidence": 0.58, "answer_confidence": 0.58,
"action": { "act_probability": 0.42 }
}
},
"flat": { "department": "billing", "frustration": 1.64, "is_urgent": 0.42 },
"usage": {
"input_tokens": 512, "output_tokens": 0,
"state_tokens": 480, "state_tokens_dropped": 0,
"truncated": false, "truncated_questions": []
}
}
flat only appears with "flat": true. Probabilities, scores and confidences are rounded to
4 decimals. usage.truncated is true when input tokens had to be dropped to fit the model's
context limit; the affected question ids are in truncated_questions.
Errors
| Status | When |
|---|---|
400 |
Neither questions nor preset; unknown preset (the message lists valid names); invalid questions JSON on GET; any model-side validation error (empty instructions, bad criteria, …). |
401 |
--api-key set and the key is missing/wrong. |
503 |
Model not loaded yet (server still starting). |
Example
curl -s http://127.0.0.1:8000/v1/predict \
-H 'content-type: application/json' -H 'X-API-Key: ...' \
-d '{"text": "I was charged twice, please refund", "preset": "triage", "flat": true}'
Presets
preset is a shortcut for a ready-made question map. It is a server-side feature: any
HTTP client can send it (the AI SDK adapter does not expose it — it always sends an explicit
questions map). Bundled names:
| Name | What it answers |
|---|---|
triage |
intent, urgency, frustration, refund request, churn risk (support tickets) |
email |
team/category routing, spam, phishing, urgency, reply expected |
guard |
jailbreak, prompt injection, sensitive data, harm severity, topic |
moderation |
toxicity, harassment, threats, spam, violation severity |
router |
difficulty, domain, needs tools, sensitive (money/legal/medical/safety) |
The definitions live in the laya-mlx package; dump
any of them as JSON:
python -c "import json, laya_mlx; print(json.dumps(laya_mlx.triage_questions(), indent=2))"
The API accepts only the preset names — the definitions themselves are not served over
HTTP. Dump one as above and pass it back as questions if you want to edit it first.
Run at login (launchd, macOS)
laya-mlx-http install --host 0.0.0.0 --api-key "$KEY" # write + load a user agent
laya-mlx-http start | stop | status | logs | uninstall
No sudo, no hand-edited system files; the plist lives in ~/Library/LaunchAgents.
Configuration
Flags only — nothing is read from the environment or a config file:
| Flag | Default |
|---|---|
--host |
127.0.0.1 |
--port |
8000 |
--api-key |
unset (no auth) |
--model |
aac6fef/laya-mlx |
--dtype |
float16 |
Library use
Embed the server in your own Python process. This is the ASGI app, not a client SDK —
from Python you can also skip HTTP entirely and call
laya_mlx directly (agent.predict(text, questions)).
from laya_mlx_http import create_app, Settings
app = create_app(Settings(host="127.0.0.1", port=8000, api_key="…"))
# uvicorn laya_mlx_http:app ← runs with default settings (127.0.0.1:8000)
License
Apache-2.0. See NOTICE in the repository.
Metadata
Release files for laya-mlx-http 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| laya_mlx_http-0.1.1.tar.gz | 15.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| laya_mlx_http-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 30.7 kB
Release files / laya_mlx_http-0.1.1.tar.gz
| Download URL | laya_mlx_http-0.1.1.tar.gz |
|---|---|
| Size | 15.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3f34223c15606db858594a5b7443a3d7d5d5a1ee2bcfe23f08738a75e7256aac
|
|
BLAKE2b-256 checksum How to use checksums |
a9138a95ca730691f7d0dd9edd89283bf45a5a6675422b8a7376d211264ba96a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / laya_mlx_http-0.1.1-py3-none-any.whl
| Download URL | laya_mlx_http-0.1.1-py3-none-any.whl |
|---|---|
| Size | 15.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f21efdf4134fb5b0a35d3b856fb1d55920a4ac2c7bde61fa6622df7b4e024eba
|
|
BLAKE2b-256 checksum How to use checksums |
4d7c437fdcf0aeca3d387784a2f066cabbc7fafdf2434ee2131e3125e218c501
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log