The assistant that runs on your own PC.
Luna is a personal AI built around a local Qwen3 14B. No cloud AI model and no account. Some tools do use the network, such as web search.
Luna’s safety layer, PolicyGuard, checks the evidence before it states anything.
Measured on an RX 9070 XT: Vulkan up to +23% faster generation (43.7 vs 35.5 tokens/s at 30k context), ROCm 13–20% faster prompt reading. See the benchmarks.
See the benchmarksHow it worksLatest tests
- Generation speed, Vulkan vs ROCm, 1k to 30k context: Vulkan +9% to +23%.
- Prompt reading, Vulkan vs ROCm, long prompts: ROCm 13–20% faster.
Luna in action
Asking Luna to scan a folder for viruses. It finds the EICAR test file. The waiting time is sped up.
Private by design.
Many assistants send your requests to a cloud model. Luna's model runs on one GPU at home.
Local
The model runs on my own GPU through llama.cpp, and no cloud AI model is used. Some tools still use the network: web search goes through my self-hosted SearXNG, which queries search engines; weather, mail checks and system updates contact their own servers.
French first
Built for everyday use in French, with answers checked against the real state of the machine.
Measured
Every number on this page comes from a test I ran, with the setup written down.
One model, one server.
Vulkan up to +23% faster generation (+9% at 1k), ROCm 13–20% faster prompt reading
Same card, same model, same prompts. Generation speed in tokens per second, by context length.
Vulkan generates faster, up to +23% at 30k. ROCm reads long prompts 13–20% faster. VRAM use is about the same. Median of 7 runs, fixed seed, 256 tokens, q8_0 KV cache, llama.cpp b11177, ROCm 7.2.4, Mesa RADV 25.2.8. Output quality was not compared.
How Luna picks a tool
Three real requests from Luna's log (7 Oct 2026): a rule, SetFit, then Qwen3 14B. SetFit decides for 10 of its 53 classes; the rest goes to Qwen3.

How I test reliability
A test bench replays real scenarios 5 times each, and code decides pass or fail. Run of 11 Oct 2026: 109 of 110 evaluated attempts passed (99.1 %, 22 scenarios x 5 runs). 15 more attempts (3 scenarios) were skipped because their precondition was not met, so they are not counted.

Run it yourself.
HIP_VISIBLE_DEVICES=0 llama-server \
--model qwen3-14b-Q4_K_M.gguf --alias qwen3:14b \
--ctx-size 32768 --n-gpu-layers 99 --flash-attn on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--host 127.0.0.1 --port 8080
Same protocol and prompts as the ROCm run, with the official llama.cpp b11177 Vulkan build. bench.py is the script behind the numbers above. run_vk.sh is the same run with private paths replaced by variables; it was not re-run after this cleanup.
- vulkan_command.txt
- llama-server command, build and driver details
- run_vk.sh
- runs q8_0 then f16: 1 warm-up and 7 passes per context
- bench.py
- measures prompt speed, generation speed and VRAM
Raw per-pass values were not kept, only medians and min–max.
ROCm reference is a local HIP build, Vulkan is the official b11177 archive.
- GGUF sha256
a8cc1361f3145dc01f6d77c6c82c9116b9ffe3c97b34716fe20418455876c40e- Versions
- llama.cpp b11177 · ROCm 7.2.4 · Mesa RADV 25.2.8
- Hardware
- RX 9070 XT, 16 GB
Built by one person, measured in public.
New tests and demos go on X first. Luna's code is not published; the measurements are.
Follow @g_lejars