Sentence Transformer Server
Lightweight FastAPI wrapper that exposes any sentence-transformers model over HTTP for easy embedding generation.
Quick Start
- Install the package from PyPI (Python 3.13+):
pip install stserve
- Launch the server with the default MiniLM model (choose your preferred runner):
stserve # or via uvx uvx stserve
- Request embeddings from another terminal:
curl -X POST http://127.0.0.1:8501/embed \ -H "Content-Type: application/json" \ -d '{"texts": ["hello world", "how are you?"]}'
Configuration
--model: sentence-transformers model name or local path (defaults toall-MiniLM-L6-v2).--device: target Torch device; auto-detects CUDA.--batch-size: batches encode calls for throughput.--normalize: toggles L2 normalization on embeddings.--show-progress: prints encode progress in the server logs.--host/--port: Uvicorn bind address (defaults127.0.0.1:8501).
Example with GPU and normalization:
uvx stserve --model sentence-transformers/all-MiniLM-L12-v2 --device cuda --normalize
API
POST /embed: Body{ "texts": [str, ...] }→{ "embeddings": [[float, ...], ...] }.GET /health: Returns current configuration (model, device, batch size, etc.).
Development Notes
- The CLI entry point is
stserve.app:main. - Embedding work happens in a thread pool so the event loop stays responsive.
- CUDA usage requires PyTorch with GPU support.
Metadata
Release files for stserve 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| stserve-0.1.0.tar.gz | 52.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| stserve-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 55.4 kB
Release files / stserve-0.1.0.tar.gz
| Download URL | stserve-0.1.0.tar.gz |
|---|---|
| Size | 52.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
da15faedd7d3cfd556ef3b248f55d64864cf6fee034be2a7e955e249407bc9ee
|
|
BLAKE2b-256 checksum How to use checksums |
a2f842c730dfbe14ff4beb676e6d59c3b5a54792f3cd6f2fd339099848ab412f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2025.
Transparency logRelease files / stserve-0.1.0-py3-none-any.whl
| Download URL | stserve-0.1.0-py3-none-any.whl |
|---|---|
| Size | 3.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6bb04c73f0716502e6efb80e27304fbe7b67e412c9942085fc2bfd8ab63e12d2
|
|
BLAKE2b-256 checksum How to use checksums |
a3ca76f1e973244a2d8d98f308b4d8a473bf456da653e69780459db1f9f16670
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2025.
Transparency log