The Hardware-Agnostic, Adaptive Operating System for Local LLM Inference
Project description
InferenceOS
The Hardware-Agnostic, Adaptive Operating System for Local LLM Inference
PyPI Package • Overview • Why InferenceOS • Architecture • Feature Matrix • Quick Start • CLI Reference • Docs • Roadmap • Contributing
Overview
InferenceOS is an open-source, hardware-aware operating system and control plane designed for local Large Language Model (LLM) inference. Built on an architectural fork of llama.cpp, InferenceOS bridges high-performance C++ tensor execution kernels with intelligent, dynamic Python runtime orchestration.
Install instantly via PyPI:
pip install inferenceos
While traditional inference runtimes rely on static offload configs and fixed microbatch sizes, InferenceOS dynamically adapts execution in real time. It profiles system hardware capabilities (NVIDIA, AMD, Apple Silicon, Intel Arc, and integrated GPUs), optimizes layer placement using integer linear programming cost models, dynamically sizes prefill microbatches to eliminate OOM errors, monitors live VRAM pressure to execute hysteresis-controlled layer migrations, compresses KV caches using attention-sink eviction policies (H2O, StreamingLLM), and records telemetry in a persistent SQLite Runtime Learning database.
Why InferenceOS Exists
Local LLM deployment faces critical hardware boundaries:
- Heterogeneous System Memory: Modern systems mix fast VRAM, system RAM, and unified iGPU memory across PCIe interconnects. Static layer offloading leads to underutilized GPUs or out-of-memory crashes.
- Prefill Ingestion Bottlenecks: Prompt ingestion (prefill) requires large microbatches for throughput, but static batch sizes spike VRAM and trigger fatal OOMs.
- Rigid Resource Allocation: Workloads vary dramatically between single-turn completion and extended multi-user chat sessions. Fixed memory allocations waste VRAM or truncate context windows.
- Lack of Feedback Loops: Runtimes typically treat each execution as stateless, failing to learn optimal configurations across consecutive runs on identical hardware.
InferenceOS solves these challenges by providing an adaptive, self-tuning runtime layer over high-performance GGUF execution engines.
Key Innovations & Features
- Hardware Profiler: Multi-vendor detection engine (NVIDIA
pynvml/nvidia-smi, AMDamdsmi/rocm-smi, Apple Metalsysctl, Intel oneAPI/wmic) generating system capability profiles with automatic backend and quantization recommendations. - Dynamic Microbatch Scheduler: Prefill batch scheduler using multi-heuristic scoring (VRAM headroom, prompt TPS, GPU occupancy, PCIe boundary crossings) to maximize ingestion throughput while maintaining memory safety.
- Live Hysteresis Layer Migration: Real-time layer offload coordinator that migrates layers between dGPU, iGPU, and CPU during active inference under dynamic memory pressure.
- KV Cache Manager: Advanced memory management supporting FP16/Q8_0/Q4_0 cache quantization and attention-sink eviction strategies (H2O Heavy-Hitter Oracle, StreamingLLM, LRU, FIFO).
- Runtime Learning & Knowledge Base: SQLite-backed feedback database recording execution telemetry (TTFT, prompt t/s, eval t/s, OOM history) and predicting optimal settings via ML regression models.
- OpenAI & Ollama API Server: High-throughput FastAPI HTTP server supporting
/v1/chat/completions,/v1/completions,/v1/embeddings, and Ollama/api/*endpoints with streaming Server-Sent Events (SSE). - 21-Command Rich CLI: Interactive terminal suite including
chat,run,serve,benchmark,optimize,doctor,inspect,monitor,hardware, andplacement.
Architecture Overview
InferenceOS separates low-level tensor execution from high-level scheduling, memory policy enforcement, and runtime learning.
graph TD
User["User / Client (CLI, HTTP API, Python SDK)"] --> ControlPlane["InferenceOS Control Plane"]
subgraph ControlPlane ["InferenceOS Control Plane"]
CLI["Rich CLI Suite (21 Subcommands)"]
Server["FastAPI Server (OpenAI & Ollama APIs)"]
Profiler["Hardware Profiler (CPU/GPU/iGPU/RAM)"]
AdaptiveEngine["Adaptive Runtime Engine"]
end
AdaptiveEngine --> SchedulerSystem["Scheduler & Memory Subsystems"]
subgraph SchedulerSystem ["Scheduler & Memory Subsystems"]
Microbatch["Dynamic Microbatch Scheduler"]
ContextSched["Context Window Scaling"]
PlacementEngine["Layer Placement Cost Model"]
MigrationEngine["Live Hysteresis Migration"]
KVCache["KV Cache Manager (H2O, StreamingLLM)"]
MemoryBudget["Unified Memory Budget Planner"]
end
AdaptiveEngine --> LearningSystem["Learning & Intelligence"]
subgraph LearningSystem ["Learning & Intelligence"]
RuntimeLearning["Runtime Learning (SQLite Database)"]
RuntimeKB["Hardware Knowledge Base"]
PerfIntel["Performance Intelligence & Health Sensor"]
Optimizer["5-Step Automated Placement Optimizer"]
end
AdaptiveEngine --> BackendLayer["Execution Backend Layer (llama.cpp Fork)"]
subgraph BackendLayer ["Execution Backend Layer"]
CUDA["CUDA Backend (NVIDIA)"]
ROCm["HIP / ROCm Backend (AMD)"]
Metal["Metal Backend (Apple)"]
Vulkan["Vulkan Backend (Cross-Platform)"]
CPU["CPU Backend (AVX-512 / AMX / NEON)"]
end
Feature Comparison
| Feature / Capability | InferenceOS | llama.cpp | Ollama | vLLM | ONNX Runtime |
|---|---|---|---|---|---|
PyPI Package (pip install) |
Yes (inferenceos) |
No | No | Yes | Yes |
| GGUF Tensor Kernels | Core | Native | Native | No | No |
| Hardware Auto-Profiling | Multi-Vendor | Manual | Basic | Basic | Basic |
| Dynamic Microbatch Prefill | Automated | Fixed (-b) |
Fixed | Paged | Fixed |
| Live Hysteresis Layer Migration | Automated | Static | Static | No | No |
| iGPU & Unified Memory Modeling | Specialized | Basic | No | No | Basic |
| KV Cache Compression (H2O/Sinks) | Dynamic | Basic | No | Paged | No |
| Runtime Learning (SQLite ML) | Built-in | No | No | No | No |
| OpenAI & Ollama Dual API Server | Native | OpenAI | Ollama | OpenAI | No |
| 21-Command Rich Terminal TUI | Built-in | Simple CLI | Simple CLI | No | No |
Quick Start
Option A: Install via PyPI (Recommended)
pip install inferenceos
Option B: Install from Source
git clone https://github.com/rudy-07/InferenceOS.git
cd InferenceOS
python -m venv .venv
# On Windows:
.venv\Scripts\activate
# On Linux/macOS:
# source .venv/bin/activate
pip install -e .
1. Profile System Hardware
inferenceos hardware
2. Single-Shot Prompt Execution
inferenceos run models/llama-3-8b-instruct.Q4_K_M.gguf --prompt "Explain quantum entanglement in 3 sentences."
3. Interactive Terminal Chat Session
inferenceos chat models/llama-3-8b-instruct.Q4_K_M.gguf --theme nord
4. Launch OpenAI & Ollama Compatible Server
inferenceos serve models/llama-3-8b-instruct.Q4_K_M.gguf --port 11434 --backend auto
Call via Standard OpenAI Client in Python:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="inferenceos")
response = client.chat.completions.create(
model="llama-3-8b-instruct",
messages=[{"role": "user", "content": "What makes InferenceOS adaptive?"}],
stream=True
)
for chunk in response:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
CLI Command Reference
InferenceOS includes 21 dedicated subcommands accessible via inferenceos <command>:
| Subcommand | Syntax Example | Description |
|---|---|---|
chat |
inferenceos chat model.gguf --theme nord |
Launch interactive terminal UI chat session with live streaming |
run |
inferenceos run model.gguf -p "Hello" -b vulkan |
Execute single-shot prompt with streaming token output |
serve |
inferenceos serve model.gguf --port 11434 |
Launch local OpenAI (/v1/*) and Ollama (/api/*) compatible server |
benchmark |
inferenceos benchmark model.gguf -t 8 |
Run end-to-end performance benchmarking matrix across batch sizes |
optimize |
inferenceos optimize model.gguf |
Automated 5-step hardware detection, placement generation & caching |
profile |
inferenceos profile model.gguf |
Profile execution and generate interactive flamegraph HTML reports |
hardware |
inferenceos hardware |
Inspect CPU, GPU, iGPU, VRAM, RAM, and interconnect bandwidth |
doctor |
inferenceos doctor |
Run system diagnostic health check across drivers, SDKs, and binaries |
inspect |
inferenceos inspect model.gguf |
Inspect GGUF header metadata, tensor shapes, and layer counts |
models |
inferenceos models list |
Register, list, inspect, tag, and manage local GGUF model library |
config |
inferenceos config --interactive |
View, edit, or interactively configure runtime settings |
plugins |
inferenceos plugins |
List and manage loaded extension plugins |
cache |
inferenceos cache clear |
View or clean layer placement and benchmark optimization caches |
logs |
inferenceos logs -n 50 |
View and tail recent runtime execution logs |
telemetry |
inferenceos telemetry |
Launch terminal metrics & historical telemetry dashboard UI |
monitor |
inferenceos monitor |
Launch HTOP-style live process & GPU hardware resource monitor |
stats |
inferenceos stats |
Display aggregated performance summary statistics |
placement |
inferenceos placement model.gguf |
Compute and visualize layer placement plan across CPU/dGPU/iGPU |
update |
inferenceos update |
Check for engine updates and C++ binary builds |
version |
inferenceos version |
Display InferenceOS release version and system build information |
reset |
inferenceos reset |
Reset configuration, profiles, and runtime caches to default state |
Python API Usage
from pathlib import Path
from profiler.hardware_profiler import get_system_resources
from layer_placement import PlacementEngine, ModelDescriptor
from inference_runtime import InferenceSession, RuntimeConfig
# 1. Profile system resources
sys_res = get_system_resources()
hw_profile = sys_res.to_dict()
# 2. Build model descriptor
descriptor = ModelDescriptor.from_gguf_metadata(
metadata={"arch": "llama", "num_layers": 32, "hidden_size": 4096},
model_size_bytes=4 * 1024**3,
quant_type="Q4_K_M",
model_name="Llama-3-8B",
)
# 3. Compute optimal layer placement plan
engine = PlacementEngine(hw_profile=hw_profile)
plan = engine.generate_placement_plan(descriptor, context_length=4096)
# 4. Launch session
config = RuntimeConfig(threads=8, use_flash_attn=True)
with InferenceSession(
model_path=Path("models/llama-3-8b.gguf"),
plan=plan,
config=config,
hw_profile=hw_profile,
) as session:
result = session.run("Explain dynamic microbatching.")
print(result.generated_text)
print(f"Eval Throughput: {result.stats.eval_tps:.2f} tok/s")
Roadmap
- Phase 1 — Hardware Profiler: Multi-vendor CPU, GPU, iGPU, VRAM, and RAM discovery (
profiler/). - Phase 2 — Dynamic Microbatch & KV Management: Adaptive prefill batching, H2O attention sinks, KV quantization (
scheduler/,kv_manager/). - Phase 3 — Layer Placement & Hysteresis Migration: ILP cost modeling and live layer migration under memory pressure (
layer_placement/,layer_migration/). - Phase 4 — Runtime Learning & HTTP Server: SQLite execution database, prediction engine, OpenAI & Ollama dual REST API server (
runtime_learning/,server/). - Phase 5 — Distributed Multi-Node Engine (In Progress): Pipeline parallel distribution across networked local workstations.
- Phase 6 — Speculative Decoding Pipeline (Planned): Automated draft model speculative decoding for low-latency generation.
Contributing & Governance
We welcome contributions from developers, systems engineers, researchers, and AI enthusiasts!
- Please read our Contributing Guide for details on code style, branch workflows, and developer tutorials.
- Abide by our Code of Conduct in all community interactions.
- Report security vulnerabilities following our Security Policy.
License
InferenceOS is released under the open-source MIT License.
Built for the Local AI Community.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file inferenceos-1.0.0.tar.gz.
File metadata
- Download URL: inferenceos-1.0.0.tar.gz
- Upload date:
- Size: 367.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d5be9f75b4bd0a5910e711c8e350cb385ae1303c2fa7f4f632ea11e6521aebe9
|
|
| MD5 |
fdf6200c77f5d0e86f63d1c0e8e94ead
|
|
| BLAKE2b-256 |
4ff5f0ce18e2d4ed512db45b750adc88182f30543e3f04f0eb6154d0e2b3374f
|
File details
Details for the file inferenceos-1.0.0-py3-none-any.whl.
File metadata
- Download URL: inferenceos-1.0.0-py3-none-any.whl
- Upload date:
- Size: 412.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
32ce834a7abe812f6e5b59edafbe3211e646902adb230509f2e25bd55e175372
|
|
| MD5 |
731c244e3726ab741bcad278d2a4aa45
|
|
| BLAKE2b-256 |
6e16714f046fd862208bacbdd3bfabc218bed44e6db590e8f8dd1d17f5b82884
|