Skip to main content

make-localllm-easier — run the best local LLM your GPU can handle, in one command

localllm picks, downloads and runs the most accurate local AI model for your PC and your language, chosen from real benchmark measurements, with llama.cpp tuned for AMD, NVIDIA, Intel and Apple GPUs.

pip install make-localllm-easier
localllm

That's it. localllm checks your GPU and RAM, picks the most accurate model we have measured for your language that fits your card, downloads llama.cpp and the model, starts it with settings profiled op by op, and opens the chat page. You also get an OpenAI-compatible API at http://127.0.0.1:8080/v1 for any app that speaks it. Offline, private, free.

localllm doctor     # what this GPU is good for: model sizes, speed, how much text it can hold
localllm list       # every model we have measured, with scores per language
localllm serve      # API only, no browser
localllm eval       # score any running server in English + your language

FAQ

Which local LLM should I run on my GPU? Run localllm doctor. It lists which model sizes fit your card (4B up to 120B MoE), at which quantization, how fast they should run, and the most accurate measured model for your language.

Can a 16 GB GPU run a 27B model? Yes. Qwen3.8-27B at ~3.5 bits (12.2 GB) runs at ~50 tok/s on an RX 9070 XT and keeps 81.8% on English Global-MMLU-Lite. gemma-4-26B-A4B (13.3 GB) runs at ~69 tok/s with similar accuracy.

Is a 2-bit quantized model good enough? Usually not for non-English use: 2-bit costs 8-13 accuracy points, and Hindi, Arabic and Thai lose the most (13 points).

Why is llama.cpp slow on my AMD (or Intel) GPU on Windows? If Resizable BAR is off, llama.cpp's Vulkan backend puts buffers in a 256 MB host-visible heap backed by system RAM and decode drops up to 1.7x. localllm sets GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1 for you (llama.cpp#27097).

Does it work offline? After the first download, yes. Nothing leaves your PC.

Which languages are measured? 23 languages on Global-MMLU-Lite, 44 countries' own exams on INCLUDE, plus Thai (ThaiExam). localllm eval --langs ... measures any of them on your hardware.

What localllm doctor tells you

GPU AMD Radeon RX 9070 XT  15.9 GB  (640 GB/s)    RAM 32 GB    language: th

Model sizes for this PC (whole model on the GPU = fast):
  [OK  ] 4B                       Q8 4.2 GB  ~90 tok/s (est.)
  [OK  ] 8B                       Q8 8.5 GB  ~45 tok/s (est.)
  [OK  ] 14B                      Q6 11.5 GB  ~33 tok/s (est.)
  [OK  ] 24-32B                   Q3 13.2 GB  ~29 tok/s (est.)
  [SLOW] 30B MoE (3B active)      Q4 18.0 GB with experts in RAM - works, ~10-25 tok/s
  [NO  ] 70B                      needs ~42.0 GB - too big for this PC

Best measured model for you: gemma4-26b-a4b-qat  (MoE with ~4B active params: fastest)
What it can do here:
  TH  real local school/licence exams   65.7% correct  <- your language
  holds ~78k tokens at once (~130 pages of text) next to the model
  answers at ~69 tok/s

Speeds marked est. come from your card's memory bandwidth, calibrated on measured runs. Everything else is measured.

Measured results (RX 9070 XT 16 GB, Windows 11, llama.cpp Vulkan)

Accuracy (%) on multiple-choice exams, zero-shot. global = Global-MMLU-Lite: the same 400 questions translated, so languages compare like for like. regional = INCLUDE: real exams written in each country (ThaiExam for Thai).

Qwen3.8-27B Q3 (12.2 GB) gemma-4-26B-A4B QAT Q4 (13.3 GB) Qwen3.8-27B 2-bit (7.8 GB)
English 81.8 82.2 74.2
Chinese 76.2 / 74.7 73.5 / 66.5 67.8 / 67.8
Spanish 80.2 / 76.8 74.5 / 75.2 70.8 / 69.2
Japanese 73.5 / 87.6 74.5 / 81.9 65.8 / 77.9
Arabic 70.8 / 71.2 71.5 / 73.6 60.8 / 57.2
Hindi 69.0 / 74.3 69.5 / 71.0 56.2 / 55.5
Thai – / 67.1 – / 65.7 – / 54.2
decode speed 50 tok/s (MTP) 69 tok/s 40 tok/s

Cells are global / regional. Margins are about ±4 (global) and ±5 (regional) points at 95%, so localllm treats gaps under 2 points as a tie and picks the faster model.

Findings worth knowing

  1. 2-bit costs 8-13 points, and lower-resource languages pay the most. Hindi, Arabic and Thai lose 13; English, Chinese and Spanish about 8-9. A 177B MoE squeezed to 1.6 bits scored below a 27B at 3 bits.
  2. Calibrating the quantization on your language doesn't help at ~3.5 bits. A Thai-text importance matrix scored the same as the stock one in Thai, English and Chinese (64.8 vs 64.6 Thai). At this level the number of bits matters, the calibration text doesn't.
  3. AMD/Intel cards without Resizable BAR lose up to 1.7x decode speed in llama.cpp's Vulkan backend. Hybrid DeltaNet models (Qwen3.5/3.8) suffer most: they rewrite a 3 MB state per layer per token.
  4. Qwen3.8 GGUFs ship a multi-token-prediction head. Drafting 2 tokens with it adds ~40% decode speed for free; drafting 3 is slower.
  5. The first run of a new llama.cpp build is slow while the GPU driver compiles its shaders once (~15 s).

How the benchmark works

localllm eval asks each question with thinking off and reads the log-probability of every answer letter from the first generated token, then takes the most likely one. It's prompt processing only, so a language takes a few minutes, and the result is deterministic. Data is downloaded at eval time from the original Apache-2.0 datasets (Global-MMLU-Lite, INCLUDE, ThaiExam) and never redistributed. It measures knowledge and reasoning in multiple choice, not writing quality.

Contributing

The catalog only grows with measurements. Run localllm eval --langs en,<yours> on your GPU and open a PR with ~/.localllm/results.json and your GPU name. Other languages' local exams are very welcome. See ROADMAP.md for what's next: using less system RAM (0.2), working alongside cloud provider APIs (0.3), and a speed-only release (0.4).

License

MIT. Models keep their own licenses; benchmark data keeps its own (Apache-2.0).

Metadata

Release files for make-localllm-easier 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for make-localllm-easier 0.1.0
File Size Uploaded
make_localllm_easier-0.1.0.tar.gz 20.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for make-localllm-easier 0.1.0
File Interpreter ABI Platform
make_localllm_easier-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 38.8 kB

Release files / make_localllm_easier-0.1.0.tar.gz

Download URL make_localllm_easier-0.1.0.tar.gz
Size 20.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0b4ba54958466a28d075dbbe7c3dd1bf95a47c7f36a8459d05944e30f19a27ba
BLAKE2b-256 checksum
How to use checksums
9ca7147dabb3087d99acf403e244a68d6c2d38e0ba96b1c728291b94a36697c8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / make_localllm_easier-0.1.0-py3-none-any.whl

Download URL make_localllm_easier-0.1.0-py3-none-any.whl
Size 18.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2692d0d3ff840b91930fcf21ba94c712fd33dfc18b9b9de4b2bc1fa08c6ab764
BLAKE2b-256 checksum
How to use checksums
2f59994e48d4d48498c34530c4795ed580e6adb164e10928c9cd1c0ce16b5e7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page