Skip to main content

InferKit - Deploy any AI function in 3 lines

Library for ML / Vision / LLM / Agent. No FastAPI boilerplate needed.

Install

pip install inferkit              # core (no Pillow)
pip install inferkit[vision]      # + Pillow for image in/out
pip install inferkit[torch,transformers]  # heavy ML stacks
pip install -e .[dev]             # local dev

Usage - 3 lines

# my_model.py
from inferkit import infer

@infer
async def run(payload, files=None):
    # payload: {"text": "..."}  files: list[bytes] for images/audio
    return {"output": f"echo: {payload.get('text')}"}

@infer.stream  # optional for LLM streaming
async def run_stream(payload):
    for tok in payload.get("text","").split():
        yield tok + " "

Run:

inferkit serve my_model.py --port 8001
# docs at http://localhost:8001/docs

Endpoints auto created:

  • POST /api/v1/infer (multipart file + json)
  • POST /api/v1/infer/json (json only)
  • POST /api/v1/infer/stream (SSE)
  • WS /ws/infer and WS /api/v1/ws/infer (WebSocket + streaming)

Image output helpers:

from inferkit import image_to_base64, bytes_to_response
from PIL import Image
return image_to_base64(Image.new("RGB",(512,512),"red"))
return bytes_to_response(png_bytes, "image/png")
# also still supported: return {"image_base64": b64} or return png_bytes

Init new project

inferkit init
# creates .env.example, .env, Dockerfile, my_model.py

Deploy (one command, any OS, auto detects Docker)

inferkit deploy
# if docker available -> docker compose/build
# else -> venv + uvicorn on INFERKIT_HOST:INFERKIT_PORT

Config via .env

INFERKIT_APP_NAME=InferKit   # Swagger title - custom service name
INFERKIT_HOST=0.0.0.0
INFERKIT_PORT=8000
INFERKIT_CORS_ORIGINS=["*"]   # or * or http://a.com,http://b.com
INFERKIT_MAX_UPLOAD_MB=50
INFERKIT_RATE_LIMIT=60/minute
INFERKIT_API_KEY=            # if set, require X-API-Key header (also ?api_key=)
INFERKIT_DEBUG=false
# API pruning per project (set false to hide from Swagger):
INFERKIT_ENABLE_MULTIPART=true
INFERKIT_ENABLE_JSON=true
INFERKIT_ENABLE_STREAM=true
INFERKIT_ENABLE_WS=true
# plain HOST/PORT/CORS_ORIGINS also work for backwards compat

Customization

Service name (easiest):

# .env
INFERKIT_APP_NAME=MyService
# code (overrides .env)
from inferkit.server import create_app
app = create_app(title="MyService")

Prune APIs per project:

# only JSON, hide multipart/stream/WS
INFERKIT_ENABLE_MULTIPART=false
INFERKIT_ENABLE_STREAM=false
INFERKIT_ENABLE_WS=false
app = create_app(enable_multipart=False, enable_stream=False, enable_ws=False)
# stream is auto-hidden if no @infer.stream is defined

Tutorial (0 to 100)

Complete guide with training and checkpoint: tutorial/00-100-complete-guide.md

python tutorial/train_example.py      # train and save checkpoints/model.pkl
inferkit serve tutorial/inference_example.py --port 8001  # serve

Programmatic

from inferkit import serve
serve("my_model.py", port=8000)

Documentation

  • docs/index.md - Usage
  • docs/api.md - API Reference
  • docs/vision.md - Vision example
  • docs/tutorial.md - Tutorial index

Release files for inferkit 0.1.11

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for inferkit 0.1.11
File Size Uploaded
inferkit-0.1.11.tar.gz 22.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for inferkit 0.1.11
File Interpreter ABI Platform
inferkit-0.1.11-py3-none-any.whl Python 3 none any Details

Total release size: 36.3 kB

Release files / inferkit-0.1.11.tar.gz

Download URL inferkit-0.1.11.tar.gz
Size 22.6 kB
Tags Source
SHA-256 checksum
How to use checksums
ce54f58978e0457763f64621e786b5d5612f95d4efe007947b2b01d823b68b57
BLAKE2b-256 checksum
How to use checksums
1c537aa85e58a73754473e5f0dcecf642e1e457d47da9d894e5b27614721cc8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / inferkit-0.1.11-py3-none-any.whl

Download URL inferkit-0.1.11-py3-none-any.whl
Size 13.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bdf16d8ce696e5ef052c3bcd0789dd5612f85327260bce0b35855ac1cceba047
BLAKE2b-256 checksum
How to use checksums
957fa53cbae96cdba033d9324854934644e38d6798a1f6a8c0f4d1ca455c5a29
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.11 This release

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page