Compact AI, any device. No API key. No cost.
Project description
🤖 CompactLLM
Compact AI, any device. No API key. No cost.
CompactLLM lets you run powerful AI models locally on any device — from an old laptop to a gaming PC. No API keys. No subscriptions. No internet required after the first download. Just AI that runs anywhere.
from compactllm import Model
model = Model('smollm2')
print(model.ask('What is AI?'))
# → "AI, or Artificial Intelligence, is the simulation of human intelligence..."
That's it. CompactLLM downloads the model on first run and caches it forever. Every run after that is instant.
✨ Why CompactLLM?
| CompactLLM | Other Libraries | |
|---|---|---|
| API key needed | ❌ Never | ✅ Always |
| Works offline | ✅ Yes | ❌ No |
| Auto hardware detection | ✅ Yes | ❌ No |
| Shows RAM requirements | ✅ Yes | ❌ No |
| Beginner-friendly | ✅ Yes | ⚠️ Sometimes |
| Free forever | ✅ Yes | ⚠️ Free tier limits |
🚀 Installation
pip install compactllm
CPU-only install (lighter, no CUDA dependencies):
pip install compactllm
pip install torch --index-url https://download.pytorch.org/whl/cpu
With GPU acceleration:
pip install compactllm[gpu]
Everything included:
pip install compactllm[full]
⚡ Quick Start
1. Check what your machine can run
from compactllm import detect_hardware, list_models
# See your hardware
hw = detect_hardware()
print(hw)
# {
# 'ram_gb': 16.0,
# 'cpu_name': 'Intel Core i7-12700H',
# 'has_gpu': False,
# 'tier': 'mid',
# ...
# }
# List models your machine can run right now
models = list_models()
for m in models:
print(f"{m['name']:<30} ~{m['min_ram_gb']} GB RAM | {m['best_for']}")
2. Load and run a model
from compactllm import Model
model = Model('smollm2')
response = model.ask('Explain machine learning in simple terms.')
print(response)
3. Let CompactLLM pick for you
# Automatically selects the best model your hardware can run
model = Model.auto()
print(model.ask('Hello!'))
# Auto-select for a specific task
model = Model.auto(task='coding')
print(model.ask('Write a Python function to reverse a string.'))
📋 Model Catalog
CompactLLM organises models into 4 tiers based on RAM requirements. Use list_models() to see what's available for your machine.
🥔 Potato Tier — Under 4 GB RAM
Works on old laptops, budget PCs, Raspberry Pi 5
| Model ID | Name | RAM | Best For | License |
|---|---|---|---|---|
smollm2-tiny |
SmolLM2 135M | ~0.3 GB | Ultra-fast, testing | Apache 2.0 |
qwen2.5-0.5b |
Qwen 2.5 0.5B | ~0.6 GB | Tiny multilingual tasks | Apache 2.0 |
tinyllama |
TinyLlama 1.1B | ~1 GB | General chat | Apache 2.0 |
gemma3-1b |
Gemma 3 1B | ~1 GB | 128K context, 140 languages | Gemma ToS |
deepseek-r1-1.5b |
DeepSeek R1 1.5B | ~1.5 GB | Math & reasoning | MIT |
💻 Basic Tier — 4–8 GB RAM
Works on most laptops from 2018+
| Model ID | Name | RAM | Best For | License |
|---|---|---|---|---|
smollm2 ⭐ |
SmolLM2 1.7B | ~2 GB | Best quality <2B. General use | Apache 2.0 |
qwen2.5-1.5b |
Qwen 2.5 1.5B | ~1.8 GB | Multilingual, long docs | Apache 2.0 |
llama3.2-1b |
Llama 3.2 1B | ~1.2 GB | Meta's edge model | Llama 3.2 |
exaone-2.4b |
EXAONE 3.5 2.4B | ~2.5 GB | Reasoning + EN/KO | Apache 2.0 |
phi3.5-mini |
Phi-3.5 Mini | ~3.5 GB | Coding & reasoning | MIT |
🖥️ Mid Tier — 8–16 GB RAM
Works on good laptops and most desktops
| Model ID | Name | RAM | Best For | License |
|---|---|---|---|---|
llama3.2-3b ⭐ |
Llama 3.2 3B | ~3.5 GB | Best 3B overall | Llama 3.2 |
qwen2.5-coder-3b |
Qwen 2.5 Coder 3B | ~3 GB | Code generation | Apache 2.0 |
gemma3-4b |
Gemma 3 4B | ~4 GB | Chat + image understanding | Gemma ToS |
mistral-7b ⭐ |
Mistral 7B | ~5 GB | Fast, general purpose | Apache 2.0 |
🚀 High Tier — 16 GB+ RAM
Workstations and gaming PCs
| Model ID | Name | RAM | Best For | License |
|---|---|---|---|---|
phi4 |
Phi-4 14B | ~9 GB | Best coding + reasoning | MIT |
llama3.1-8b |
Llama 3.1 8B | ~6 GB | High quality chat | Llama 3.1 |
⭐ = Recommended starting point for that tier
💡 Usage Examples
Simple Q&A
from compactllm import Model
model = Model('smollm2')
print(model.ask('What is the difference between RAM and ROM?'))
Streaming Output
model = Model('mistral-7b')
for chunk in model.stream('Write a short story about a robot learning to paint'):
print(chunk, end='', flush=True)
print() # final newline
Custom System Prompt
model = Model('smollm2')
response = model.ask(
'What should I eat for breakfast?',
system='You are a professional nutritionist. Give concise, practical advice.'
)
print(response)
Deterministic Output
# temperature=0.0 gives the same output every time (good for testing)
response = model.ask('Name 3 planets.', temperature=0.0)
Interactive Chat
model = Model('smollm2')
model.chat()
# Starts an interactive terminal chat session
# Type 'help' for commands, 'quit' to exit
Browse and Compare Models
from compactllm import list_models
# Models your machine can run
my_models = list_models()
# Only coding models
coding = list_models(task='coding')
# Only potato-tier models
tiny = list_models(tier='potato')
# Everything available
all_models = list_models(show_all=True)
# Print a summary
for m in my_models:
print(f"{m['id']:<20} {m['params']:<8} ~{m['min_ram_gb']}GB {m['best_for']}")
Get Model Details
from compactllm import get_model_info
info = get_model_info('mistral-7b')
print(info['name']) # Mistral 7B Instruct
print(info['min_ram_gb']) # 5.0
print(info['license']) # Apache 2.0
print(info['tags']) # ['chat', 'general', 'fast', 'recommended']
🖥️ Command Line Interface
After installing CompactLLM, the compact command is available in your terminal.
# Check your hardware
compact hardware
# List models for your machine
compact list
# List all models regardless of hardware
compact list --all
# Filter by tier
compact list --tier basic
compact list --tier potato
# Filter by task
compact list --task coding
compact list --task reasoning
# Ask a one-shot question
compact ask smollm2 "What is recursion?"
compact ask mistral-7b "Explain quantum computing simply"
# Start an interactive chat
compact chat smollm2
compact chat mistral-7b
# Get full details about a model
compact info smollm2
compact info mistral-7b
# Check version
compact version
🔧 How It Works
pip install compactllm
↓
from compactllm import Model
↓
Model('smollm2') ← Resolves model ID, checks your RAM
↓
model.ask('Hello!') ← Downloads model on first call (cached forever)
↓
Auto-detects CPU/GPU ← Picks best device automatically
↓
Returns response string ← Clean, no boilerplate
Model caching: Models are downloaded to ~/.cache/compactllm/ on first use. After that, they load from disk — no internet needed.
Hardware detection: CompactLLM checks your RAM and GPU automatically. If you try to load a model that's too big, it warns you and suggests a better option.
Model registry updates: The model list updates automatically once a week by fetching a lightweight JSON file from GitHub. This means you always see the latest recommended models without updating the package. It falls back gracefully to the built-in list if you're offline.
🗺️ Roadmap
- v0.1 — Python library: core inference, model registry, hardware detection, CLI
- v0.2 — Streaming improvements, GGUF/llama.cpp backend for better CPU performance
- v0.3 — Vision models (image understanding with LLaVA/Gemma Vision)
- v0.4 — Audio (Whisper speech-to-text, Piper text-to-speech — both free and local)
- v0.5 — Embeddings for semantic search and RAG pipelines
- v1.0 — NPM package (
npm install compactllm) for Node.js and browser - v1.1 — Browser-local inference via ONNX/WebAssembly
🤝 Contributing
Contributions are very welcome! Here's how to get involved:
git clone https://github.com/compactllm-dev/compactllm
cd compactllm
pip install -e ".[dev]"
pytest tests/
Most wanted contributions:
- New model entries in
registry.py(especially non-English models) - Better hardware detection (especially for ARM Macs and Windows)
- Tests and documentation improvements
- Real-world usage examples
📜 License
CompactLLM is released under the Apache 2.0 license — free for personal and commercial use.
Individual AI models have their own licenses (all free for personal use). Check the license field from get_model_info() or list_models() before commercial deployment.
🙏 Acknowledgements
Built on top of the incredible work by:
CompactLLM — Compact AI, any device. No API key. No cost.
getcompactllm.com •
GitHub •
PyPI
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file compactllm-0.1.1.tar.gz.
File metadata
- Download URL: compactllm-0.1.1.tar.gz
- Upload date:
- Size: 24.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f2b58eab8d63fb8b63baabba93934f5c6fb751cd48c554d5898328a103e6e7a2
|
|
| MD5 |
1311d8ab1766aa6c42f9caa194b4c5c5
|
|
| BLAKE2b-256 |
48dbc970c4a6ead80d5810da39b5d056707f211f73bf7a01ce7cd47f0cfa18f7
|
File details
Details for the file compactllm-0.1.1-py3-none-any.whl.
File metadata
- Download URL: compactllm-0.1.1-py3-none-any.whl
- Upload date:
- Size: 22.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b6a156940ef1251c82f6c5e2ae5eee4643e0ceb4b3db933128d65b141a863a58
|
|
| MD5 |
08f7f193545f7465d782022072a3c579
|
|
| BLAKE2b-256 |
a4960b54206a1a8dfcf184c74e682e5d33c1b45f55d9adaf5ed16acd903abe72
|