Skip to main content

forge-cloud

OpenAI-compatible reasoning-aware inference proxy for Qwen3.6.

Point your OpenAI client at forge-cloud instead of directly at vLLM/SGLang/Ollama. The proxy routes thinking mode based on query complexity, swaps sampling parameters to match the mode, normalizes backend flags, and tags responses with routing metadata.

What it does

  1. Receives a standard /v1/chat/completions request
  2. Classifies query complexity (simple/moderate/complex)
  3. Decides thinking mode (think vs no_think) with correct sampling params
  4. Normalizes the enable_thinking flag for the target backend (vLLM nested, DashScope top-level, llama.cpp server-side)
  5. Forwards to the user's configured backend
  6. Tags the response with routing metadata and estimated token split (thinking vs response)

The proxy does not run inference. It configures and monitors it.

Install

pip install forge-cloud

Quick start

# Set admin key and backend URL
export FORGE_ADMIN_KEY=my-secret
export FORGE_BACKEND_URL=http://localhost:8000
export FORGE_BACKEND_TYPE=vllm

# Start the proxy
forge-cloud

Create an API key:

curl -X POST http://localhost:8741/v1/keys \
  -H "Authorization: Bearer my-secret" \
  -H "Content-Type: application/json" \
  -d '{"name": "my-app"}'
# Returns: {"key": "fk-...", "name": "my-app", "tier": "free", ...}

Use it like any OpenAI endpoint:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8741/v1",
    api_key="fk-..."  # key from above
)

response = client.chat.completions.create(
    model="Qwen/Qwen3.6-35B-A3B",
    messages=[{"role": "user", "content": "refactor this module"}],
)
print(response.choices[0].message.content)

Response metadata

Every response includes a forge field with routing metadata and estimated token counts:

{
  "id": "chatcmpl-test",
  "choices": ["..."],
  "usage": {"...": 0},
  "forge": {
    "thinking_mode": "think",
    "complexity": "complex",
    "backend": "vllm",
    "sampling_profile": "thinking",
    "thinking_tokens": 450,
    "response_tokens": 120
  }
}

Endpoints

Method Path Description
POST /v1/chat/completions Proxied chat completion with forge routing
POST /v1/keys Create API key (admin auth required)
GET /health Proxy health check

Configuration

All settings are environment variables with FORGE_ prefix:

Variable Default Description
FORGE_HOST 127.0.0.1 Bind address
FORGE_PORT 8741 Port
FORGE_BACKEND_URL http://localhost:8000 Default backend URL
FORGE_BACKEND_TYPE vllm Backend type: vllm, sglang, dashscope, llamacpp
FORGE_FREE_DAILY_LIMIT 1000 Free tier requests per day
FORGE_ADMIN_KEY (empty) Admin key for creating API keys
FORGE_DB_PATH forge.db SQLite database path
FORGE_REQUEST_TIMEOUT 120.0 Backend request timeout (seconds)

Per-key backend override

Each API key can have its own backend URL and type:

curl -X POST http://localhost:8741/v1/keys \
  -H "Authorization: Bearer my-secret" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "sglang-user",
    "tier": "paid",
    "backend_url": "http://sglang-server:30000",
    "backend_type": "sglang"
  }'

Tiers

  • Free: 1,000 requests/day, single backend target
  • Paid: no rate limit, per-key backend routing

Streaming

Streaming is supported. Set stream: true in the request and the proxy forwards the SSE stream from the backend.

Note: Streaming responses do not include forge metadata (thinking_mode, complexity, token counts). The proxy cannot inspect the response until the stream completes, so the forge field is only present in non-streaming responses.

Dependencies

  • qwen-think -- thinking session manager (routing, budget, sampling)
  • FastAPI + uvicorn
  • httpx -- async HTTP client for backend forwarding
  • aiosqlite -- async SQLite for API keys and usage tracking

License

Apache-2.0

Metadata

Release files for forge-infer-cloud 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for forge-infer-cloud 0.1.3
File Size Uploaded
forge_infer_cloud-0.1.3.tar.gz 17.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for forge-infer-cloud 0.1.3
File Interpreter ABI Platform
forge_infer_cloud-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 32.7 kB

Release files / forge_infer_cloud-0.1.3.tar.gz

Download URL forge_infer_cloud-0.1.3.tar.gz
Size 17.1 kB
Tags Source
SHA-256 checksum
How to use checksums
960d150daaea4fa641bcfeeb96fcff645b2956a77c031589ebfc8970991caf13
BLAKE2b-256 checksum
How to use checksums
75e2ff0812ea36e8bd3f6f9d0ca63b523ff44a0954d0b65d90587b326c271c9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 2, 2026.

Transparency log

Release files / forge_infer_cloud-0.1.3-py3-none-any.whl

Download URL forge_infer_cloud-0.1.3-py3-none-any.whl
Size 15.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0fec283546359ebee8ca6239fe0bbd782bac413a82a634af78a071620cc1f883
BLAKE2b-256 checksum
How to use checksums
8b35aee25ef9fdc70109e47c1695c0087d7081196a9e9d92b4e4d0d1c8245930
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page