Qwen3-Think
The Thinking Session Manager for Qwen3.6 (or any Qwen3+ model that uses enable_thinking) -- SDK + Router + Bug Fixes
Python library for managing Qwen3.6's thinking state across sessions, backends, and frameworks.
What It Solves
-
Backend Normalization -- Qwen3.6's
enable_thinkingflag has three different invocation patterns across backends. This library normalizes them into a single API. -
Sampling Parameter Swap -- Qwen3.6 requires different sampling parameters for thinking vs. non-thinking mode. A router that flips
enable_thinkingwithout also swapping params produces silently degraded output. -
Context Budget Guard -- Qwen3.6 advises maintaining at least 128K tokens of context to preserve thinking capabilities. This library tracks and guards against silent degradation.
Installation
pip install qwen-think
# With OpenAI client support:
pip install qwen-think[openai]
Quick Start
from openai import OpenAI
from qwen_think import ThinkingSession
client = OpenAI(base_url="http://localhost:8000/v1", api_key="...")
session = ThinkingSession(client, backend="vllm", budget=200_000)
# Auto-routes: detects complexity, sets mode, swaps sampling
response = session.chat("refactor this module", preserve=True)
The Three Backend Patterns
| Backend | Flag Format | Notes |
|---|---|---|
| vLLM / SGLang | extra_body={"chat_template_kwargs": {"enable_thinking": False}} |
Nested |
| DashScope | extra_body={"enable_thinking": False} |
Top-level |
| llama.cpp | --chat-template-kwargs '{"enable_thinking": false}' |
Server-side only |
Bug Fixes Included
vLLM Semantic Router (#858)
When use_reasoning: false is configured, the router removes the field instead of explicitly setting enable_thinking: false. Since Qwen3.6 thinks by default, removing the field has no effect.
Fix: This library always explicitly sets the boolean -- never omits it.
Ray Serve (#52979)
enable_thinking: false in the HTTP body doesn't propagate to the model, and thinking output continues appearing.
Fix: The normalized payload always includes the explicit flag.
Sampling Parameters
Qwen3.6 requires different sampling params depending on thinking mode:
Thinking mode:
temperature=0.6, top_p=0.95, top_k=20, min_p=0.0,
presence_penalty=1.5, repetition_penalty=1.0
Instruct / non-thinking mode:
temperature=0.7, top_p=0.80, top_k=20, min_p=0.0,
presence_penalty=1.5
A router that flips enable_thinking without atomically swapping these params produces silently degraded output -- not incorrect, just suboptimal in ways that don't surface as errors.
Context Budget
Qwen3.6 advises maintaining a context length of at least 128K tokens to preserve thinking capabilities. The BudgetManager tracks usage and:
- Warns when approaching the threshold
- Auto-compresses older messages when running low
- Refuses to continue when below the minimum
session = ThinkingSession(client, budget=200_000, min_context=128_000)
status = session.budget_status
# BudgetStatus(total_tokens=200000, used_tokens=10000, available_tokens=190000,
# min_context=128000, action=BudgetAction.OK,
# message="Available: 190,000 of 200,000 tokens.")
Thinking Preservation
Qwen3.6 introduces preserve_thinking -- a feature that retains thinking context across conversation history, improving reasoning quality for iterative development.
session = ThinkingSession(client, preserve_thinking=True)
# Thinking content is cached and included in subsequent turns
Complexity Router
The router classifies query complexity and selects the appropriate mode:
| Complexity | Mode | Preserved | Use Case |
|---|---|---|---|
| SIMPLE | NO_THINK | No | "What is X?" |
| MODERATE | THINK | No | Multi-sentence reasoning |
| COMPLEX | THINK + preserve | Yes | Coding, debugging, refactoring |
from qwen_think import ComplexityRouter
router = ComplexityRouter()
decision = router.route("implement a REST API with authentication")
# RouterDecision(complexity=MODERATE, mode=THINK, preserve_thinking=False)
API Reference
ThinkingSession
session = ThinkingSession(
client, # OpenAI-compatible client
backend="vllm", # or "sglang", "dashscope", "llamacpp"
model="Qwen/Qwen3.6-35B-A3B",
budget=200_000, # Total context budget
min_context=128_000, # 128K minimum for thinking
preserve_thinking=True,
auto_route=True, # Auto-classify complexity
)
Manual Mode Control
# Force a specific mode
session.chat("quick answer", mode=ThinkingMode.NO_THINK)
# Let the router decide
session.chat("refactor this module")
# Check current state
session.thinking_mode # -> ThinkingMode.THINK
session.budget_status # -> BudgetStatus(...)
License
Apache-2.0
Metadata
Release files for qwen-think 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| qwen_think-0.1.3.tar.gz | 24.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| qwen_think-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 48.8 kB
Release files / qwen_think-0.1.3.tar.gz
| Download URL | qwen_think-0.1.3.tar.gz |
|---|---|
| Size | 24.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d67d9c16792ce960efd324f83ca856740187b99a278fc4f2a8cac7cb70547704
|
|
BLAKE2b-256 checksum How to use checksums |
258166cb14c37aaf4d8a7d8205731b74578f3c1067c2ebdca7c7bd595ef03140
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 2, 2026.
Transparency logRelease files / qwen_think-0.1.3-py3-none-any.whl
| Download URL | qwen_think-0.1.3-py3-none-any.whl |
|---|---|
| Size | 23.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c243b79486e7c19a9eccf4b8c39f7ec8fe02b1d4489d4a3e97fed9679a4fc0aa
|
|
BLAKE2b-256 checksum How to use checksums |
c5b4bdf7af3fd21548744684c0eeabbbf4ec69a6de29687b9882088013b6eee7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 2, 2026.
Transparency log