Skip to main content

laya-mlx-http

HTTP server for laya-mlx typed decision models. Runs the model once, keeps it in unified memory, and answers POST /v1/predict requests from anything on your network.

This package is only the server. The API below is plain JSON over HTTP and is client-agnostic — curl, Python, Node, a sensor, anything. No SDK is required.

The Vercel AI SDK integration is a separate, optional npm package, @noahwaldner/laya-mlx-http; it translates AI SDK calls to and from this same API and is irrelevant if you are not using the AI SDK.

Requires Apple silicon (MLX) and Python 3.11+.

Install

pip install laya-mlx-http          # or: uv tool install laya-mlx-http

Run

# localhost only (default, safest)
laya-mlx-http

# expose on your LAN — always set an API key
laya-mlx-http --host 0.0.0.0 --api-key "$(openssl rand -hex 24)"

Then from another machine:

curl http://<server-ip>:8000/health

Endpoints

Method Path Purpose
GET /health liveness + model readiness (never authenticated)
POST /v1/predict {"text": "...", "preset": "triage"} or a full questions map
GET /v1/predict?text=...&preset=triage&flat=true same, for simple clients

If --api-key is set, /v1/* requires X-API-Key: <key> (or Authorization: Bearer <key>).

Predict format

This section is the server's own wire format — what goes over the wire no matter which client sends it. It is not tied to any SDK.

If you are used to the AI SDK adapter, two names differ: its state is text here, and its boolean is Laya's noul. Both translations happen client-side; the server only ever sees text/noul. Full mapping: adapter README.

Request

POST /v1/predict with a JSON body:

Field Type Notes
text string Required. The state to classify (the message, ticket, …).
questions object Question map (below). Provide this or preset.
preset string One of the bundled presets; only read when questions is absent.
flat bool Default false. Adds a flat map of question id → chosen value.
model string Optional. Echoed back as model in the response instead of the server's model id.

GET /v1/predict takes the same call as query parameters: text, preset, flat, and questions as a raw JSON string (?questions={"a":{"type":"noul",...}}, URL-encoded).

You must send questions or preset — otherwise 400.

Question map

{
  "department": {
    "type": "choice",                                  // choice | score | noul
    "instructions": "Which team should handle this?",  // required, non-blank
    "criteria": {                                      // shape depends on type
      "billing": "Charges and refunds",
      "support": "Other requests"
    }
  }
}
  • instructions — required and must not be blank. A string, or any JSON-serializable object/array (it is serialized to JSON before it reaches the model).
  • type: "choice" — criteria required: a non-empty {label: description} object (or an array of unique labels; descriptions may be null). Returns the winning label plus full probabilities.
  • type: "score" — criteria required: a non-empty, ordered array of level descriptions (e.g. ["calm", "annoyed", "very angry"]). Returns a 0-based expected score.
  • type: "noul" — Laya's boolean. criteria optional: {"false": "...", "true": "..."}. Returns a probability in [0, 1] for true.

Response

{
  "model": "aac6fef/laya-mlx",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": { "billing": 0.91, "support": 0.09 },
      "confidence": 0.91, "answer_confidence": 0.91,
      "action": { "act_probability": 0.42 }
    },
    "frustration": {
      "type": "score",
      "score": 1.64,
      "legend": { "0": "calm", "1": "annoyed", "2": "very angry" },
      "probabilities": { "0": 0.1, "1": 0.5, "2": 0.4 },
      "confidence": 0.5, "answer_confidence": 0.5,
      "action": { "act_probability": 0.42 }
    },
    "is_urgent": {
      "type": "noul",
      "noul": 0.42,
      "confidence": 0.58, "answer_confidence": 0.58,
      "action": { "act_probability": 0.42 }
    }
  },
  "flat": { "department": "billing", "frustration": 1.64, "is_urgent": 0.42 },
  "usage": {
    "input_tokens": 512, "output_tokens": 0,
    "state_tokens": 480, "state_tokens_dropped": 0,
    "truncated": false, "truncated_questions": []
  }
}

flat only appears with "flat": true. Probabilities, scores and confidences are rounded to 4 decimals. usage.truncated is true when input tokens had to be dropped to fit the model's context limit; the affected question ids are in truncated_questions.

Errors

Status When
400 Neither questions nor preset; unknown preset (the message lists valid names); invalid questions JSON on GET; any model-side validation error (empty instructions, bad criteria, …).
401 --api-key set and the key is missing/wrong.
503 Model not loaded yet (server still starting).

Example

curl -s http://127.0.0.1:8000/v1/predict \
  -H 'content-type: application/json' -H 'X-API-Key: ...' \
  -d '{"text": "I was charged twice, please refund", "preset": "triage", "flat": true}'

Presets

preset is a shortcut for a ready-made question map. It is a server-side feature: any HTTP client can send it (the AI SDK adapter does not expose it — it always sends an explicit questions map). Bundled names:

Name What it answers
triage intent, urgency, frustration, refund request, churn risk (support tickets)
email team/category routing, spam, phishing, urgency, reply expected
guard jailbreak, prompt injection, sensitive data, harm severity, topic
moderation toxicity, harassment, threats, spam, violation severity
router difficulty, domain, needs tools, sensitive (money/legal/medical/safety)

The definitions live in the laya-mlx package; dump any of them as JSON:

python -c "import json, laya_mlx; print(json.dumps(laya_mlx.triage_questions(), indent=2))"

The API accepts only the preset names — the definitions themselves are not served over HTTP. Dump one as above and pass it back as questions if you want to edit it first.

Run at login (launchd, macOS)

laya-mlx-http install --host 0.0.0.0 --api-key "$KEY"   # write + load a user agent
laya-mlx-http start | stop | status | logs | uninstall

No sudo, no hand-edited system files; the plist lives in ~/Library/LaunchAgents.

Configuration

Flags only — nothing is read from the environment or a config file:

Flag Default
--host 127.0.0.1
--port 8000
--api-key unset (no auth)
--model aac6fef/laya-mlx
--dtype float16

Library use

Embed the server in your own Python process. This is the ASGI app, not a client SDK — from Python you can also skip HTTP entirely and call laya_mlx directly (agent.predict(text, questions)).

from laya_mlx_http import create_app, Settings

app = create_app(Settings(host="127.0.0.1", port=8000, api_key="…"))
# uvicorn laya_mlx_http:app   ← runs with default settings (127.0.0.1:8000)

License

Apache-2.0. See NOTICE in the repository.

Metadata

Release files for laya-mlx-http 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for laya-mlx-http 0.1.1
File Size Uploaded
laya_mlx_http-0.1.1.tar.gz 15.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for laya-mlx-http 0.1.1
File Interpreter ABI Platform
laya_mlx_http-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 30.7 kB

Release files / laya_mlx_http-0.1.1.tar.gz

Download URL laya_mlx_http-0.1.1.tar.gz
Size 15.1 kB
Tags Source
SHA-256 checksum
How to use checksums
3f34223c15606db858594a5b7443a3d7d5d5a1ee2bcfe23f08738a75e7256aac
BLAKE2b-256 checksum
How to use checksums
a9138a95ca730691f7d0dd9edd89283bf45a5a6675422b8a7376d211264ba96a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / laya_mlx_http-0.1.1-py3-none-any.whl

Download URL laya_mlx_http-0.1.1-py3-none-any.whl
Size 15.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f21efdf4134fb5b0a35d3b856fb1d55920a4ac2c7bde61fa6622df7b4e024eba
BLAKE2b-256 checksum
How to use checksums
4d7c437fdcf0aeca3d387784a2f066cabbc7fafdf2434ee2131e3125e218c501
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page