Skip to main content

Deqio

Decisions in. Probabilities out.

Deqio runs fast typed AI decision models behind one consistent local API and a lightweight browser UI. It supports multiple decision engines and lets you switch between installed models without changing your application code.

Public decision endpoints stay the same regardless of the active model:

POST /v1/noul
POST /v1/choice
POST /v1/shared

Supported models

Model MLX MPS CUDA
SemIf / Qwen3.5 4B ✓ ✓ ✓
Kev 0.8B ✓ ✓ ✓
Kev 4B ✓ ✓ ✓
Kev 9B ✓ ✓ ✓
Decider 0.8B — ✓ ✓
Decider 2B — ✓ ✓
Decider 4B — ✓ ✓
Laya English 421M ✓ ✓ ✓
Laya Multilingual 322M ✓ ✓ ✓
Laya Typed Decisions 421M ✓ ✓ ✓
Von — ✓ ✓
Bespoke Nimble 9B ✓ — ✓

1. Installation

The recommended installation is from PyPI with uv tool. Deqio keeps models, isolated runtimes, configuration, and benchmark results in the directory where you use it; the Python package itself stays small.

macOS — Apple Silicon

Deqio supports both MLX and MPS on Apple Silicon.

  1. Install uv if you do not already have it:
curl -LsSf https://astral.sh/uv/install.sh | sh
  1. Install Deqio from PyPI:
uv tool install deqio
  1. Create a workspace and enter it:
mkdir -p deqio-work
cd deqio-work
  1. Choose and install a backend/model:
deqio models setup

On a Mac, the installer offers:

  • mlx — recommended for models with native MLX support
  • mps — PyTorch on Apple Silicon, required by models such as Decider

On first setup Deqio creates editable config.json, models.json, and benchmarks/basic.json files in this workspace. Model runtimes and weights are also kept outside the PyPI package.

  1. Start Deqio:
deqio serve

The first start of some models may download additional weights.


Linux — NVIDIA GPU

Deqio uses the CUDA backend on Linux.

  1. Make sure the NVIDIA driver is working:
nvidia-smi
  1. Install uv and Deqio:
curl -LsSf https://astral.sh/uv/install.sh | sh
uv tool install deqio
  1. Create a workspace:
mkdir -p deqio-work
cd deqio-work
  1. Choose and install a CUDA-compatible model:
deqio models setup
  1. Start Deqio:
deqio serve

Windows — NVIDIA GPU

Deqio uses the CUDA backend on Windows.

  1. Make sure the NVIDIA driver is working in PowerShell:
nvidia-smi
  1. Install uv and Deqio:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
uv tool install deqio
  1. Create a workspace:
New-Item -ItemType Directory -Force deqio-work
Set-Location deqio-work
  1. Choose and install a CUDA-compatible model:
deqio models setup
  1. Start Deqio:
deqio serve

Install from source

For development, clone the repository instead of installing the PyPI tool:

git clone https://github.com/ILuce/deqio.git
cd deqio
uv sync --extra dev
uv run deqio models setup
uv run deqio serve

When running from a source checkout, use uv run deqio ...; when installed from PyPI with uv tool install deqio, use deqio ... directly.

Models you already installed

List locally available models:

deqio models installed

Choose another installed model before starting the server:

deqio models use

Show the current selection:

deqio status

2. Using Deqio from the UI

Start the server:

deqio serve

Open:

http://127.0.0.1:8787

The browser redirects to the Deqio UI.

Select a model

At the top of the UI, choose one of the models already installed on the machine and click Activate. Deqio unloads the previous runtime, loads the selected model, runs a warmup, and keeps the public API on the same address.

Noul — yes/no decision

Use Noul when the result should be a binary decision.

Fill in:

  • State — the context
  • Question — the yes/no question

Click Send to see the selected answer, probabilities, latency, token count, and cache information.

Choice — choose between options

Use Choice when the model should select one option from a list.

Fill in:

  • State
  • Question
  • one or more options with an ID and description

Add or remove options directly in the UI and click Send.

Shared — several decisions over one state

Use Shared when several decisions should be evaluated against the same context.

Enter the shared State, add the decisions and their options, then send them together. This is preferable to several separate requests when an engine can reuse the common state efficiently.

UI:        http://127.0.0.1:8787/ui
API docs:  http://127.0.0.1:8787/docs
Health:    http://127.0.0.1:8787/health

Stop the server with Ctrl+C.

3. Using Deqio as an API

Start Deqio once:

deqio serve

Base URL:

http://127.0.0.1:8787

The UI is not involved when your application calls the API directly.

Noul

curl -s \
  -X POST http://127.0.0.1:8787/v1/noul \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "The patch changed source code and no tests have been run yet.",
    "question": "Should tests be run before considering the task complete?"
  }'

Typical response:

{
  "decision": "yes",
  "probabilities": {
    "yes": 0.97,
    "no": 0.03
  }
}

Choice

curl -s \
  -X POST http://127.0.0.1:8787/v1/choice \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "A customer was charged twice for the same subscription renewal.",
    "question": "Which team should handle this request?",
    "options": [
      {"id": "billing", "description": "Payments, refunds, invoices, and duplicate charges."},
      {"id": "access", "description": "Login and account access problems."},
      {"id": "technical", "description": "Product defects and service failures."}
    ]
  }'

Shared

curl -s \
  -X POST http://127.0.0.1:8787/v1/shared \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "A patch changed an authentication module. Unit tests passed, but integration tests have not been run.",
    "decisions": [
      {
        "id": "validation",
        "question": "What should happen next?",
        "options": [
          {"id": "run_integration_tests", "description": "Run integration tests before proceeding."},
          {"id": "finish", "description": "Finish without more validation."}
        ]
      },
      {
        "id": "release",
        "question": "Is the change ready to release?",
        "options": [
          {"id": "yes", "description": "The change is ready to release."},
          {"id": "no", "description": "More validation is required."}
        ]
      }
    ]
  }'

Check and switch the active model through the API

List installed model profiles:

curl -s http://127.0.0.1:8787/v1/models/installed

Activate an already installed profile:

curl -s \
  -X POST http://127.0.0.1:8787/v1/models/activate \
  -H 'Content-Type: application/json' \
  -d '{
    "model_id": "decider-0.8b",
    "backend": "mps"
  }'

The decision API URL does not change when the active model changes.

Benchmarking installed models

Deqio includes an editable starter suite in benchmarks/basic.json: 50 Noul, 50 Choice, and 50 Shared requests. Stop deqio serve before benchmarking so the benchmark can load each model with the machine's memory available.

Run it interactively and choose all installed models or selected profiles:

deqio benchmark

Run every installed model compatible with the current machine:

deqio benchmark --all

Or select profiles explicitly:

deqio benchmark \
  --model decider-0.8b:mps \
  --model laya-typed-decisions:mlx

The console shows live PASS/FAIL and latency for every request. Full results.jsonl and summary.json files are written under .deqio/benchmarks/<timestamp>/. Add or edit cases in benchmarks/basic.json as the benchmark grows.

Nimble note: Bespoke Nimble 9B follows the upstream MLX/CUDA workflow. Its first installation downloads the adapter and pinned Qwen3.5-9B base, then prepares merged local weights, so it needs substantially more disk/RAM than the smaller models.

For the complete request and response schemas, open:

http://127.0.0.1:8787/docs

License

MIT. See LICENSE.

Metadata

Release files for deqio 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for deqio 0.1.0
File Size Uploaded
deqio-0.1.0.tar.gz 222.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for deqio 0.1.0
File Interpreter ABI Platform
deqio-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 302.3 kB

Release files / deqio-0.1.0.tar.gz

Download URL deqio-0.1.0.tar.gz
Size 222.4 kB
Tags Source
SHA-256 checksum
How to use checksums
aec3903cda31fb6c6155dbc3cd2fbadb7a7ddea22efa2787da29d0dbae6828d0
BLAKE2b-256 checksum
How to use checksums
bd5f34dea2e2e2f96bd613b5da3be215151d80605ee63dadfa3e8604dc0a55e7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / deqio-0.1.0-py3-none-any.whl

Download URL deqio-0.1.0-py3-none-any.whl
Size 79.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b4e3c07756407955c284e7fdde2931493c402d23aa2223b06421a18816f485cc
BLAKE2b-256 checksum
How to use checksums
4c1ceddb3021a5cd6cd7ecb1a07400db7f0d34c2e59fe6fb9cbb7deda248d5e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page