forge-cloud
OpenAI-compatible reasoning-aware inference proxy for Qwen3.6.
Point your OpenAI client at forge-cloud instead of directly at vLLM/SGLang/Ollama. The proxy routes thinking mode based on query complexity, swaps sampling parameters to match the mode, normalizes backend flags, and tags responses with routing metadata.
What it does
- Receives a standard
/v1/chat/completionsrequest - Classifies query complexity (simple/moderate/complex)
- Decides thinking mode (think vs no_think) with correct sampling params
- Normalizes the
enable_thinkingflag for the target backend (vLLM nested, DashScope top-level, llama.cpp server-side) - Forwards to the user's configured backend
- Tags the response with routing metadata and estimated token split (thinking vs response)
The proxy does not run inference. It configures and monitors it.
Install
pip install forge-cloud
Quick start
# Set admin key and backend URL
export FORGE_ADMIN_KEY=my-secret
export FORGE_BACKEND_URL=http://localhost:8000
export FORGE_BACKEND_TYPE=vllm
# Start the proxy
forge-cloud
Create an API key:
curl -X POST http://localhost:8741/v1/keys \
-H "Authorization: Bearer my-secret" \
-H "Content-Type: application/json" \
-d '{"name": "my-app"}'
# Returns: {"key": "fk-...", "name": "my-app", "tier": "free", ...}
Use it like any OpenAI endpoint:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8741/v1",
api_key="fk-..." # key from above
)
response = client.chat.completions.create(
model="Qwen/Qwen3.6-35B-A3B",
messages=[{"role": "user", "content": "refactor this module"}],
)
print(response.choices[0].message.content)
Response metadata
Every response includes a forge field with routing metadata and estimated token counts:
{
"id": "chatcmpl-test",
"choices": ["..."],
"usage": {"...": 0},
"forge": {
"thinking_mode": "think",
"complexity": "complex",
"backend": "vllm",
"sampling_profile": "thinking",
"thinking_tokens": 450,
"response_tokens": 120
}
}
Endpoints
| Method | Path | Description |
|---|---|---|
| POST | /v1/chat/completions |
Proxied chat completion with forge routing |
| POST | /v1/keys |
Create API key (admin auth required) |
| GET | /health |
Proxy health check |
Configuration
All settings are environment variables with FORGE_ prefix:
| Variable | Default | Description |
|---|---|---|
FORGE_HOST |
127.0.0.1 |
Bind address |
FORGE_PORT |
8741 |
Port |
FORGE_BACKEND_URL |
http://localhost:8000 |
Default backend URL |
FORGE_BACKEND_TYPE |
vllm |
Backend type: vllm, sglang, dashscope, llamacpp |
FORGE_FREE_DAILY_LIMIT |
1000 |
Free tier requests per day |
FORGE_ADMIN_KEY |
(empty) | Admin key for creating API keys |
FORGE_DB_PATH |
forge.db |
SQLite database path |
FORGE_REQUEST_TIMEOUT |
120.0 |
Backend request timeout (seconds) |
Per-key backend override
Each API key can have its own backend URL and type:
curl -X POST http://localhost:8741/v1/keys \
-H "Authorization: Bearer my-secret" \
-H "Content-Type: application/json" \
-d '{
"name": "sglang-user",
"tier": "paid",
"backend_url": "http://sglang-server:30000",
"backend_type": "sglang"
}'
Tiers
- Free: 1,000 requests/day, single backend target
- Paid: no rate limit, per-key backend routing
Streaming
Streaming is supported. Set stream: true in the request and the proxy forwards the SSE stream from the backend.
Note: Streaming responses do not include forge metadata (thinking_mode, complexity, token counts). The proxy cannot inspect the response until the stream completes, so the forge field is only present in non-streaming responses.
Dependencies
- qwen-think -- thinking session manager (routing, budget, sampling)
- FastAPI + uvicorn
- httpx -- async HTTP client for backend forwarding
- aiosqlite -- async SQLite for API keys and usage tracking
License
Apache-2.0
Metadata
Release files for forge-infer-cloud 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| forge_infer_cloud-0.1.3.tar.gz | 17.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| forge_infer_cloud-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.7 kB
Release files / forge_infer_cloud-0.1.3.tar.gz
| Download URL | forge_infer_cloud-0.1.3.tar.gz |
|---|---|
| Size | 17.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
960d150daaea4fa641bcfeeb96fcff645b2956a77c031589ebfc8970991caf13
|
|
BLAKE2b-256 checksum How to use checksums |
75e2ff0812ea36e8bd3f6f9d0ca63b523ff44a0954d0b65d90587b326c271c9a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 2, 2026.
Transparency logRelease files / forge_infer_cloud-0.1.3-py3-none-any.whl
| Download URL | forge_infer_cloud-0.1.3-py3-none-any.whl |
|---|---|
| Size | 15.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0fec283546359ebee8ca6239fe0bbd782bac413a82a634af78a071620cc1f883
|
|
BLAKE2b-256 checksum How to use checksums |
8b35aee25ef9fdc70109e47c1695c0087d7081196a9e9d92b4e4d0d1c8245930
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 2, 2026.
Transparency log