Inferlab Integration for SGLang
Framework-specific planning and rendering for running SGLang servers through Inferlab. The consuming workspace supplies SGLang and its hardware runtime; this package supplies only the Inferlab integration boundary.
See the Inferlab repository for workspace authoring and supported topology documentation.
When profiling intent is enabled, the integration declares every single,
prefill, and decode model-serving replica as a capture target. Its replica
entry endpoint opens the Nsight Systems range through POST /start_profile
with SGLang's CUDA_PROFILER activity and closes it through
POST /stop_profile. InferLab, not this package, owns the profiler lifecycle,
capture plan, report verification, cleanup, and records.
For direct single, setting enable_metrics = true enables SGLang's native
Prometheus endpoint and declares /metrics to InferLab. Without that effective
setting the integration advertises no server-metrics capability.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file inferlab_integration_sglang-0.6.0-py3-none-any.whl.
File metadata
- Download URL: inferlab_integration_sglang-0.6.0-py3-none-any.whl
- Upload date:
- Size: 8.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
826f2759af5f712fd263e8f1055b41eee2378d68a10c24a7b9584c80f921387f
|
|
| MD5 |
fbfb09ab1198e0866d26b13e561a27d4
|
|
| BLAKE2b-256 |
5222580f5df7470528460f829bae0b0cbd400d823fa75e96070fbc7adb6e3364
|