Skip to main content

Inferlab Integration for Specialized Engines

This package connects Inferlab to a hardware-by-model Specialized Engine that implements the canonical inferlab-token-engine smg-worker command. The Engine accepts prompt token IDs and returns generated token IDs over SMG's worker protocol. TokenSpeed SMG remains responsible for the public HTTP API, tokenization, chat templates, detokenization, and response formatting.

The integration contains no model-, architecture-, hardware-, or Engine-implementation-specific lowering. A downstream workspace supplies the Rust Engine binary, SMG, model intent, source revision, locked environment, and private placement bindings. Consequently, a new conforming Engine does not need another Inferlab integration package.

The supported shape is deliberately closed: one single Engine replica in one process behind one SMG Gateway. The process owns an arbitrary nonzero pure-TP device set; attention, expert, and dense-expert tensor parallelism all equal the outer TP width, while pipeline, data, context, and expert parallelism stay at one. The contract remains serial and has no P/D Router, KV-transfer, batching, or Engine-local profiling surface. InferLab can profile the Engine process tree while TokenSpeed SMG exposes the capture-window POST /start_profile and POST /stop_profile actions on the Gateway.

The Gateway exposes server metrics on its separately allocated prometheus port. The integration selects SMG's single-target least_load policy so its canonical worker monitor polls Engine load fields while retaining the same sole routing target. In the current downstream validation, that endpoint exported SMG-owned metric families but no smg_engine_* families, so Engine-series re-export remains a downstream qualification gap.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file inferlab_integration_specialized_engine-0.4.0-py3-none-any.whl.

File metadata

File hashes

Hashes for inferlab_integration_specialized_engine-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 40e36f5aca159842a3896b2a78e69c5245a09ebca4190228dcb9229989b5ba8c
MD5 d0e7cb3b5fead134bf3ec16b41519cfd
BLAKE2b-256 b3556dfe1fdddbec1c20783876ba9c1dcd687ae0933ceffff95f204e8e64121f

See more details on using hashes here.

Release history Release notifications | RSS feed

0.10.0

1 file

0.9.0

1 file

0.8.0

1 file

0.5.0

1 file

This release

0.4.0 This release

1 file

0.3.1

1 file

0.3.0

1 file

0.2.2

1 file

0.2.1

1 file

0.2.0

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page