Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

VaporRAM 💨

VaporRAM is a local inference server, CLI and web dashboard for google/gemma-4-E4B-it, packaged for consumer hardware.

Its goal is to run the model under a 1.5 GB RAM ceiling by streaming transformer layers directly from NVMe SSD storage. That streaming engine — unbuffered O_DIRECT reads with kernel prefetch hints and an int8 quantised KV cache — is implemented in pure C under c/.

Current status (alpha): token generation runs through llama.cpp, which memory-maps the full GGUF file. The C layer streamer is built but not yet wired into the token path, so the RAM ceiling is not yet achieved — measured RSS with the Q4_K_M weights is roughly 6 GB. The dashboard reports real measured RSS, so you can see this for yourself. Connecting the streaming engine to generation is the primary remaining work.

PyPI version PyPI Downloads Docs License RAM Ceiling Status CI Pipeline


Key Features

  • Live Memory Telemetry: The dashboard and /v1/stats report the engine's actual measured RSS, host RAM and KV-cache projections — no estimated or placeholder figures.
  • Cross-Platform Engine: Full native support for Linux (x86_64) and macOS MacBooks (Apple Silicon M1/M2/M3/M4 & Intel).
  • Sequential Layer Pipeline (SLP): Zero-copy unbuffered O_DIRECT NVMe SSD layer streaming with asynchronous POSIX kernel prefetching hints (POSIX_FADV_WILLNEED).
  • AVX2 SIMD & ARM NEON Kernels: Hand-written matrix-vector kernels in the C engine, with an OpenMP build (make -C c). Benchmark them on your own hardware with ./vapor bench.
  • int8 Quantized KV Cache: Compresses Key & Value attention states with per-token scale factors (implemented in c/kv_cache.c; used by the C engine path).
  • OpenAI-Compatible API: Built-in HTTP server supporting /v1/chat/completions, /v1/responses, /v1/models, and /health.
  • Web UI & Interactive CLI: Includes an interactive terminal chat mode (vapor chat) and a web dashboard (vapor web).

Hardware & System Requirements

Resource Minimum Requirement Recommended
RAM Ceiling design target: 1.5 GB (not yet met — see status above)
Actual RSS today ~6 GB with Q4_K_M weights via llama.cpp 8 GB+ system RAM
Storage 18 GB NVMe SSD PCIe Gen3 / Gen4 NVMe SSD
Supported OS Linux (x86_64, WSL2), macOS (MacBooks M1–M4 & Intel) Linux (x86_64), macOS (Apple Silicon)
Build Tools gcc / clang / Apple Clang, make, Python 3.9+ GCC 11+ / Apple Clang with OpenMP

Installation

Option 1: Install via PyPI (Recommended)

pip install vapor-ram

Option 2: Prebuilt Release

Download and extract the latest prebuilt binary tarball:

mkdir vapor-ram && cd vapor-ram
tar xzf vapor-ram-v1.0.1-linux-x86_64.tar.gz

Option 3: Build from Source

Clone the repository and compile using make:

git clone https://github.com/sudsarkar13/vapor-ram.git
cd vapor-ram
make -C c

Usage Guide

The project includes a CLI launcher called vapor (./vapor or python3 vapor).

1. System Diagnostics & Resource Planning

Run diagnostics to check system capabilities and memory budget compliance:

# Run hardware diagnostic checks
./vapor doctor

# View RAM ceiling budget breakdown (< 1.5 GB)
./vapor plan

2. Interactive Terminal Chat

Launch an interactive chat session:

./vapor chat --preset coder

3. One-Shot Prompt Generation

Execute a quick single prompt generation from the command line:

./vapor run "Explain quantum computing in simple terms."

4. OpenAI-Compatible API Server

Start an HTTP server supporting OpenAI endpoints (/v1/chat/completions):

./vapor serve --host 0.0.0.0 --port 8000

Query the API using curl:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-E4B-it",
    "messages": [{"role": "user", "content": "Hello! What can you do?"}]
  }'

5. Web Interface

Start the server and automatically launch the Web UI in your default browser:

./vapor web

Configuration & Preset Flags

You can customize execution using presets or flags:

Subcommand / Flag Description
./vapor config Interactive terminal configuration wizard
./vapor profile High-precision RSS memory profiler
./vapor inspect Inspect model weight files and tensor layout
./vapor bench Run AVX2 SIMD core throughput benchmark
./vapor presets List available persona presets (coder, reasoner, concise)

Project Structure

  • c/vapor_engine: Compiled C SIMD inference engine binary.
  • vapor: Main Python CLI frontend launcher.
  • doctor.py: Installation and hardware diagnostic script.
  • openai_server.py: OpenAI-compatible HTTP API server implementation.
  • resource_plan.py: Dynamic memory budget calculation planner.
  • version.py: Engine version information.
  • web/: Frontend dashboard UI static assets.
  • docs/: GitHub Pages documentation website and screenshot guides.

License

This project is licensed under the Apache 2.0 License. See the LICENSE file for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vapor_ram-1.0.7a4.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vapor_ram-1.0.7a4-py3-none-any.whl (1.3 MB view details)

Uploaded Python 3

File details

Details for the file vapor_ram-1.0.7a4.tar.gz.

File metadata

  • Download URL: vapor_ram-1.0.7a4.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vapor_ram-1.0.7a4.tar.gz
Algorithm Hash digest
SHA256 362cb72717a656e3f49efe1d0a19e996435507d0f763ca87869b25db1613fe14
MD5 cf67949227b4727558c5a263a2c143dc
BLAKE2b-256 8c9ab50b1dccd759f714181456507973dde1824dca618df2b4ae7014f3f2a52f

See more details on using hashes here.

Provenance

The following attestation bundles were made for vapor_ram-1.0.7a4.tar.gz:

Publisher: release.yml on sudsarkar13/vapor-ram

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vapor_ram-1.0.7a4-py3-none-any.whl.

File metadata

  • Download URL: vapor_ram-1.0.7a4-py3-none-any.whl
  • Upload date:
  • Size: 1.3 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vapor_ram-1.0.7a4-py3-none-any.whl
Algorithm Hash digest
SHA256 4626b3030ea59f54e658a55227084a850bdbd8ea769b5956af78cbad7a115e9d
MD5 107291ede7a98fb6487ebc270559c6a5
BLAKE2b-256 834b0e74f0ee2a4b2f3062131677919492e9aa768eff74a1dff1856bd11032f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for vapor_ram-1.0.7a4-py3-none-any.whl:

Publisher: release.yml on sudsarkar13/vapor-ram

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page