A local AI assistant

The assistant that runs on your own PC.

Luna is a personal AI built around a local Qwen3 14B. No cloud AI model and no account. Some tools do use the network, such as web search.

Luna’s safety layer, PolicyGuard, checks the evidence before it states anything.

Measured on an RX 9070 XT: Vulkan up to +23% faster generation (43.7 vs 35.5 tokens/s at 30k context), ROCm 13–20% faster prompt reading. See the benchmarks.

See the benchmarksHow it works
The constellation Orion with Betelgeuse, Rigel and the three belt stars, and Luna as a glowing orb beside it Betelgeuse Rigel Luna
1home GPU runs the AI model
+23%faster generation with Vulkan at 30k
0trackers, cookies or scripts
Latest

Latest tests

  • Generation speed, Vulkan vs ROCm, 1k to 30k context: Vulkan +9% to +23%.
  • Prompt reading, Vulkan vs ROCm, long prompts: ROCm 13–20% faster.

Atom feed

Demo

Luna in action

Asking Luna to scan a folder for viruses. It finds the EICAR test file. The waiting time is sped up.

Why Luna

Private by design.

Many assistants send your requests to a cloud model. Luna's model runs on one GPU at home.

01

Local

The model runs on my own GPU through llama.cpp, and no cloud AI model is used. Some tools still use the network: web search goes through my self-hosted SearXNG, which queries search engines; weather, mail checks and system updates contact their own servers.

02

French first

Built for everyday use in French, with answers checked against the real state of the machine.

03

Measured

Every number on this page comes from a test I ran, with the setup written down.

How it works

One model, one server.

01Youquestion
02Lunaassistant logic
03llama-serverlocal API · 32k
04Qwen3 14BQ4_K_M
05GPU16 GB VRAM
Benchmarks · RX 9070 XT

Vulkan up to +23% faster generation (+9% at 1k), ROCm 13–20% faster prompt reading

Same card, same model, same prompts. Generation speed in tokens per second, by context length.

Generation speed, ROCm versus Vulkan, by context length Tokens per second on an RX 9070 XT. 1k: ROCm 53.5, Vulkan 58.3; 4k: ROCm 50.2, Vulkan 56.1; 8k: ROCm 47.3, Vulkan 53.9; 16k: ROCm 42.2, Vulkan 49.8; 30k: ROCm 35.5, Vulkan 43.7. 0 20 40 60 tokens / second ROCmVulkan 53.5 58.3 1k+9% 50.2 56.1 4k+12% 47.3 53.9 8k+14% 42.2 49.8 16k+18% 35.5 43.7 30k+23%

Vulkan generates faster, up to +23% at 30k. ROCm reads long prompts 13–20% faster. VRAM use is about the same. Median of 7 runs, fixed seed, 256 tokens, q8_0 KV cache, llama.cpp b11177, ROCm 7.2.4, Mesa RADV 25.2.8. Output quality was not compared.

Routing

How Luna picks a tool

Three real requests from Luna's log (7 Oct 2026): a rule, SetFit, then Qwen3 14B. SetFit decides for 10 of its 53 classes; the rest goes to Qwen3.

Animated explainer of how Luna picks a tool, built from three real French requests in its log of 7 Oct 2026. Each request goes through three steps: rules, then SetFit, then Qwen3 14B. 'What time is it' is settled by a rule. 'And what is my motherboard' is decided by SetFit with a score of 0.710, above the 0.6 threshold. 'Is my server running well' matches no rule, SetFit abstains with a best score of 0.33, and Qwen3 14B picks docker_status. It ends on: SetFit decides for 10 of its 53 classes, everything else falls to Qwen3 14B.
Reliability

How I test reliability

A test bench replays real scenarios 5 times each, and code decides pass or fail. Run of 11 Oct 2026: 109 of 110 evaluated attempts passed (99.1 %, 22 scenarios x 5 runs). 15 more attempts (3 scenarios) were skipped because their precondition was not met, so they are not counted.

Animated explainer of Luna's reliability bench. A real user message is replayed 5 times by the local model, and code checks the result. Three example scenarios are shown, then results by category: security, web search and accuracy mostly pass, conversation robustness is skipped. One miss and three skipped scenarios are listed, with the limits of the test, ending on: measured, not claimed.
Reproduce

Run it yourself.

ROCm run
HIP_VISIBLE_DEVICES=0 llama-server \
  --model qwen3-14b-Q4_K_M.gguf --alias qwen3:14b \
  --ctx-size 32768 --n-gpu-layers 99 --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --host 127.0.0.1 --port 8080
Vulkan run

Same protocol and prompts as the ROCm run, with the official llama.cpp b11177 Vulkan build. bench.py is the script behind the numbers above. run_vk.sh is the same run with private paths replaced by variables; it was not re-run after this cleanup.

vulkan_command.txt
llama-server command, build and driver details
run_vk.sh
runs q8_0 then f16: 1 warm-up and 7 passes per context
bench.py
measures prompt speed, generation speed and VRAM

Raw per-pass values were not kept, only medians and min–max.
ROCm reference is a local HIP build, Vulkan is the official b11177 archive.

GGUF sha256
a8cc1361f3145dc01f6d77c6c82c9116b9ffe3c97b34716fe20418455876c40e
Versions
llama.cpp b11177 · ROCm 7.2.4 · Mesa RADV 25.2.8
Hardware
RX 9070 XT, 16 GB
Follow along

Built by one person, measured in public.

New tests and demos go on X first. Luna's code is not published; the measurements are.

Follow @g_lejars