ai-launcher-cli
Simple Python CLI to chat with .gguf models locally using llama-cpp-python. No Ollama, llama.cpp, LM Studio, or online providers required.
Install
pip install ai-launcher-cli
Usage
# Basic usage
ailaunch path/to/model.gguf
# With custom settings
ailaunch model.gguf -c 8192 -t 0.8 --max-tokens 1024
# Disable streaming (wait for full response)
ailaunch model.gguf --no-stream
# Custom system prompt
ailaunch model.gguf --system "You are a coding assistant."
# Use a built-in system prompt template
ailaunch model.gguf --system-template coder
# List available models
ailaunch --list-models
# Auto-select model from common directories
ailaunch auto
# Options:
# -c, --ctx-size Context window size (default: 4096)
# -g, --gpu-layers GPU layers to offload (-1 = all, default: -1)
# -t, --threads CPU threads (0 = auto, default: 0)
# --temperature Sampling temperature (default: 0.7)
# --max-tokens Max tokens to generate (default: 512)
# --no-stream Disable streaming output
# --system Custom system prompt
# --system-template Built-in template (coder, reviewer, teacher, creative, analyst, translator, shell, greyhat)
# --tools Path to JSON file with tool definitions (OpenAI format) or JSON string
# --tool-choice Tool calling behavior: none, auto, required (default: auto)
# --list-models List available GGUF models and exit
# --save-config Save current options as defaults
# --benchmark Run benchmark after loading
# --export Export conversation on exit (markdown/json)
# --export-file File to export conversation to
# --no-history Disable loading/saving chat history
# --clear-history Clear chat history for this model
# -v, --version Show version
Tool Calling
ailaunch supports OpenAI-style function calling. Tools are defined in a JSON file using the standard OpenAI function schema.
Creating a tool definitions file
[
{
"type": "function",
"function": {
"name": "calculator",
"description": "Evaluate a mathematical expression",
"parameters": {
"type": "object",
"properties": {
"expression": {"type": "string", "description": "A math expression to evaluate"}
},
"required": ["expression"]
}
}
}
]
Using tool calling
# Load tools from a JSON file
ailaunch model.gguf --tools tools.json
# Inline JSON string
ailaunch model.gguf --tools '[{"type":"function","function":{"name":"calculator","description":"Math","parameters":{"type":"object","properties":{"expression":{"type":"string"}},"required":["expression"]}}]'
Built-in tools
The following tools are always available as fallbacks when your tool definitions include them:
| Tool | Description | Parameters |
|---|---|---|
calculator |
Evaluate a math expression | expression (string) |
get_time |
Get current date and time | none |
search_files |
Find files matching a pattern | pattern (string), directory (string) |
read_file |
Read a text file (max 10KB) | path (string) |
In-chat commands
| Command | Description |
|---|---|
/tools |
Show loaded tool definitions |
Model support
Tool calling requires a model that supports structured tool calls in chat completions. Not all GGUF models support this feature. Models like Qwen 2.5, Gemma 3, and some fine-tuned models may produce tool_calls in their responses.
| Command | Description |
|---|---|
/help |
Show help |
/save |
Save conversation to history |
/export [fmt] |
Export conversation (markdown/json) |
/clear |
Clear conversation (keep system prompt) |
/system <prompt> |
Change system prompt |
/template <name> |
Use built-in template |
/config |
Show current configuration |
/bench |
Run benchmark |
/models |
List available models |
/switch [path] |
Switch to another model |
/tools |
Show loaded tool definitions |
exit/quit/q |
Exit |
Configuration
Config is saved to ~/.config/ailaunch/config.yaml. Use --save-config to save current options.
Model Auto-Detection
Models are automatically searched in these directories:
~/.lmstudio/models~/.lmstudio/.internal/bundled-models~/.cache/huggingface/hub~/models~/Downloads~/OneDrive/Downloads~/OneDrive/Documents/Downloads
GPU Acceleration
Install with GPU extras for acceleration:
# NVIDIA CUDA
pip install ai-launcher-cli[cuda]
# Apple Metal
pip install ai-launcher-cli[metal]
Then use -g -1 to offload all layers to GPU.
Requirements
- Python 3.8+
llama-cpp-python>=0.3.0(installs automatically)
Exit
Type exit, quit, q or press Ctrl+C to exit.
Metadata
Release files for ai-launcher-cli 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ai_launcher_cli-0.4.0.tar.gz | 12.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ai_launcher_cli-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 24.2 kB
Release files / ai_launcher_cli-0.4.0.tar.gz
| Download URL | ai_launcher_cli-0.4.0.tar.gz |
|---|---|
| Size | 12.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7f367496d078fe71904417553b6ce26fb12b996a50aff22e64590d827dd0809d
|
|
BLAKE2b-256 checksum How to use checksums |
0e3848273671f6c817b9631af2856afe84fd3175d67ccaf1f05adb2b656c0587
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.13
|
Release files / ai_launcher_cli-0.4.0-py3-none-any.whl
| Download URL | ai_launcher_cli-0.4.0-py3-none-any.whl |
|---|---|
| Size | 11.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a580c0461d356019427eca480ebf1b013fef1db818d7348fd323307c5c912992
|
|
BLAKE2b-256 checksum How to use checksums |
914fb7e38a6bd4a307b0623d9e3e8c363ad6e829557bffaf3cde9c02921eebb7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.13
|