Simple CLI to chat with GGUF models locally (no Ollama/LM Studio required)
Project description
ai-launcher-cli
Simple Python CLI to chat with .gguf models locally using llama-cpp-python. No Ollama, llama.cpp, LM Studio, or online providers required.
Install
pip install ai-launcher-cli
Usage
# Basic usage
ailaunch path/to/model.gguf
# With custom settings
ailaunch model.gguf -c 8192 -t 0.8 --max-tokens 1024
# Disable streaming (wait for full response)
ailaunch model.gguf --no-stream
# Custom system prompt
ailaunch model.gguf --system "You are a coding assistant."
# Use a built-in system prompt template
ailaunch model.gguf --system-template coder
# List available models
ailaunch --list-models
# Auto-select model from common directories
ailaunch auto
# Options:
# -c, --ctx-size Context window size (default: 4096)
# -g, --gpu-layers GPU layers to offload (-1 = all, default: -1)
# -t, --threads CPU threads (0 = auto, default: 0)
# --temperature Sampling temperature (default: 0.7)
# --max-tokens Max tokens to generate (default: 512)
# --no-stream Disable streaming output
# --system Custom system prompt
# --system-template Built-in template (coder, reviewer, teacher, creative, analyst, translator, shell, greyhat)
# --tools Path to JSON file with tool definitions (OpenAI format) or JSON string
# --tool-choice Tool calling behavior: none, auto, required (default: auto)
# --list-models List available GGUF models and exit
# --save-config Save current options as defaults
# --benchmark Run benchmark after loading
# --export Export conversation on exit (markdown/json)
# --export-file File to export conversation to
# --no-history Disable loading/saving chat history
# --clear-history Clear chat history for this model
# -v, --version Show version
Tool Calling
ailaunch supports OpenAI-style function calling. Tools are defined in a JSON file using the standard OpenAI function schema.
Creating a tool definitions file
[
{
"type": "function",
"function": {
"name": "calculator",
"description": "Evaluate a mathematical expression",
"parameters": {
"type": "object",
"properties": {
"expression": {"type": "string", "description": "A math expression to evaluate"}
},
"required": ["expression"]
}
}
}
]
Using tool calling
# Load tools from a JSON file
ailaunch model.gguf --tools tools.json
# Inline JSON string
ailaunch model.gguf --tools '[{"type":"function","function":{"name":"calculator","description":"Math","parameters":{"type":"object","properties":{"expression":{"type":"string"}},"required":["expression"]}}]'
Built-in tools
The following tools are always available as fallbacks when your tool definitions include them:
| Tool | Description | Parameters |
|---|---|---|
calculator |
Evaluate a math expression | expression (string) |
get_time |
Get current date and time | none |
search_files |
Find files matching a pattern | pattern (string), directory (string) |
read_file |
Read a text file (max 10KB) | path (string) |
In-chat commands
| Command | Description |
|---|---|
/tools |
Show loaded tool definitions |
Model support
Tool calling requires a model that supports structured tool calls in chat completions. Not all GGUF models support this feature. Models like Qwen 2.5, Gemma 3, and some fine-tuned models may produce tool_calls in their responses.
| Command | Description |
|---|---|
/help |
Show help |
/save |
Save conversation to history |
/export [fmt] |
Export conversation (markdown/json) |
/clear |
Clear conversation (keep system prompt) |
/system <prompt> |
Change system prompt |
/template <name> |
Use built-in template |
/config |
Show current configuration |
/bench |
Run benchmark |
/models |
List available models |
/switch [path] |
Switch to another model |
/tools |
Show loaded tool definitions |
exit/quit/q |
Exit |
Configuration
Config is saved to ~/.config/ailaunch/config.yaml. Use --save-config to save current options.
Model Auto-Detection
Models are automatically searched in these directories:
~/.lmstudio/models~/.lmstudio/.internal/bundled-models~/.cache/huggingface/hub~/models~/Downloads~/OneDrive/Downloads~/OneDrive/Documents/Downloads
GPU Acceleration
Install with GPU extras for acceleration:
# NVIDIA CUDA
pip install ai-launcher-cli[cuda]
# Apple Metal
pip install ai-launcher-cli[metal]
Then use -g -1 to offload all layers to GPU.
Requirements
- Python 3.8+
llama-cpp-python>=0.3.0(installs automatically)
Exit
Type exit, quit, q or press Ctrl+C to exit.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai_launcher_cli-0.4.0.tar.gz.
File metadata
- Download URL: ai_launcher_cli-0.4.0.tar.gz
- Upload date:
- Size: 12.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7f367496d078fe71904417553b6ce26fb12b996a50aff22e64590d827dd0809d
|
|
| MD5 |
3542b2eb9c0542bee9992d0fa130d2fd
|
|
| BLAKE2b-256 |
0e3848273671f6c817b9631af2856afe84fd3175d67ccaf1f05adb2b656c0587
|
File details
Details for the file ai_launcher_cli-0.4.0-py3-none-any.whl.
File metadata
- Download URL: ai_launcher_cli-0.4.0-py3-none-any.whl
- Upload date:
- Size: 11.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a580c0461d356019427eca480ebf1b013fef1db818d7348fd323307c5c912992
|
|
| MD5 |
3bb01a1324e60d944a05cee25022c4ab
|
|
| BLAKE2b-256 |
914fb7e38a6bd4a307b0623d9e3e8c363ad6e829557bffaf3cde9c02921eebb7
|