AI Model Workstation — llama.cpp GUI for Windows
A powerful, open-source GUI for running local AI models via llama.cpp — built for users who want full control over their hardware without the bloat.
I built this because I didn't want to run llama.cpp from the command line every time I wanted to chat with a local model. I just wanted something simple that worked.
100% vibe coded — I'm not a developer by trade. Every line of this app was written with AI assistance (Claude, mostly). If you're a programmer and something makes you cringe, please be kind — PRs are very welcome.
![]() Dark mode |
![]() Light mode |
![]() Real-time VRAM meter — dual GPU |
![]() Hugging Face model browser |

Local model browser with quantization info
Most llama.cpp frontends sacrifice control for simplicity or the other way around. LlamaStation gives you both — a clean interface with full access to every parameter, designed for power users who want to get the most out of their hardware.
- No telemetry. No accounts. No subscriptions. 100% local, 100% yours.
- Built specifically for multi-GPU setups with per-GPU VRAM monitoring
- Supports both official llama.cpp and TurboQuant fork (asymmetric KV cache quantization)
- OpenAI-compatible API — drop-in replacement for any app that uses OpenAI
- Dark / Light mode
- English / Spanish UI (more languages easy to add)
- Collapsible sidebars
- Chat history with session management
- Thinking mode controls — two independent toggles:
- 👁 Reasoning — show/hide the reasoning block in the chat UI
- 🧠 Think ON / ⚡ Think OFF — enable or disable model thinking entirely via
chat_template_kwargs. Works for all connected clients (chat, OpenAI-compatible API, external agents). Automatically restarts the server when toggled. Compatible with TurboQuant fork.
- Web search & Deep Research via a self-hosted SearXNG instance (multi-round research pipeline, no API key needed — see Web Search Setup below)
- Vision support — attach images to chat (multimodal models)
- File attachment — send code files directly to the model
Talk to your local model out loud and have it talk back, fully offline.
- 🎙 Push-to-talk and 🔊 Always-listening modes integrated directly into the main chat
- Voice cloning via XTTS v2 — record 12 seconds of any voice (or import a clip) and the model will reply in that voice
- Speech recognition powered by faster-whisper (tiny / base / small / medium / large-v3)
- 14 languages supported (es, en, fr, de, it, pt, pl, nl, cs, ar, zh-cn, ja, hu, ko)
- Adjustable speech speed (0.5x – 2.0x)
- CUDA or CPU for XTTS — your choice
- Voice replies only when you spoke — typed messages stay text-only, like ChatGPT voice mode
- Automatic cleanup of
<think>and<|channel>tags before TTS so the model never reads internal reasoning out loud - Optional auto-load of voice models at app startup
- Dedicated Voice tab with full configuration and a separate test chat
- Load any
.ggufmodel with a full parameter modal - Per-model profile saving — settings remembered automatically
- Auto-detection of mmproj files for vision models
- Local model browser with quantization badges
- Download models directly from Hugging Face with capability tags (Vision, Tools, Thinking, Code...)
- Start / stop llama.cpp server with one click
- Real-time VRAM meter for each GPU — color coded (green/yellow/red)
- Multi-GPU support: layer split, row split, tensor split ratio
- Continuous batching, Flash Attention, KV cache offload
- Asymmetric KV cache (different types for K and V — TurboQuant)
- Auto-update llama.cpp from GitHub
- 🛡️ Server Watchdog — detects unexpected crashes (OOM, etc.) and optionally auto-relaunches the server after 5 seconds. Configurable via toggle in the Server tab.
- 🧹 Backup cleanup — each llama.cpp update creates a backup of the previous installation. The Server tab shows total backup size and lets you delete them all at once.
- Built-in API Docs tab with copy-ready curl and Python examples
- Headless mode — run without GUI, uses saved model profiles:
# Interactive model selector
python llamastation.py --no-gui
# Direct launch
python llamastation.py --no-gui --model C:\models\qwen3.gguf --port 8080- Windows 10/11
- Python 3.11+
- llama.cpp compiled for your hardware (CUDA, Vulkan, CPU)
- NVIDIA GPU recommended (AMD/Intel via Vulkan also supported)
pip install customtkinter requests PillowOptional — for Voice Mode:
pip install coqui-tts faster-whisper sounddevice soundfile scipy pydubWhat gets downloaded: The first time you click Load models in the Voice tab, XTTS v2 (~1.8 GB) and the selected Whisper model (tiny ~75 MB / base ~145 MB / small ~465 MB) are downloaded automatically to your HuggingFace cache. This only happens once.
For CUDA acceleration on XTTS (recommended if you have an NVIDIA GPU — significantly faster TTS):
# Requires torch 2.1.x with cuDNN 8 — do NOT use a newer torch version
pip install torch==2.1.2 torchvision==0.16.2 torchaudio==2.1.2 --index-url https://download.pytorch.org/whl/cu121
⚠️ CUDA fix for Windows —cudnn64_8.dll not found: If XTTS crashes on load with a cuDNN error, install the cuDNN wheel:pip install nvidia-cudnn-cu11LlamaStation adds it to the DLL search path automatically — no manual PATH editing needed.
Speech recognition (faster-whisper) always runs on CPU so it doesn't compete with your GPU during LLM inference. XTTS can run on CPU or CUDA — selectable in the Voice tab.
Option A — Pre-built .exe (recommended if you just want to run it)
No Python required. No command line. Just download and run.
- Go to the Releases page and download the latest
LlamaStation-vX.X.zip - Extract the entire folder somewhere permanent (e.g.
C:\LlamaStation\) — don't just extract the.exe, the whole folder is needed - Open the extracted folder and double-click
LlamaStation.exe - On first launch, click ⬆ Update / Download backend in the sidebar — this downloads and installs llama.cpp automatically (only needed once)
- Click Download models, search for a model, download it, and you're ready to chat
⚠️ Windows SmartScreen may show a warning the first time — this is normal for any unsigned app. Click "More info" → "Run anyway". The entire source code is open and auditable right here on GitHub — every single line. If you still want extra peace of mind, feel free to scan the .exe on VirusTotal before running it.
💡 Voice Mode is not included in the
.exerelease — it requires PyTorch and XTTS (~3 GB of dependencies). If you want voice, use Option B below.
Option B — Run from source
- Download the repo — click the green Code button → Download ZIP, extract it wherever you want
(or
git clone https://github.com/vico-png/llamastationif you know git) - Open the extracted folder and double-click
iniciar_llamastation.bat - The first time it will install all dependencies automatically — this takes a few minutes, just wait
- LlamaStation will launch automatically when it's done
- Next time, just double-click the same
.batfile to launch
⚠️ Python 3.11 is required. If the .bat shows an error about Python not found, download and install it from python.org first.
⏳ First launch takes a while — the .bat installs all dependencies including voice mode (XTTS v2, Whisper, torch). This only happens once and can take 5-10 minutes depending on your connection. Subsequent launches are instant.
- Launch LlamaStation
- Click the ⬆ Update / Download backend button at the bottom of the left sidebar — this downloads and installs the inference engine automatically (only needed once)
- Click Download models → search for a model (e.g.
gemma-4) → download it - Click My models → select your
.gguffile → configure and load - Hit ▶ Start server — the VRAM bars will fill up as the model loads
- Start chatting
LlamaStation has first-class multi-GPU support. In the model load dialog:
- Split Mode:
layer(recommended) splits the model evenly across GPUs by layers - Tensor Split: set a ratio like
1,1for 50/50 or3,1for 75/25 - Watch the VRAM meter in real time to verify the split is working
LlamaStation's Web search and 🔬 Deep Research modes query a self-hosted SearXNG instance instead of relying on a third-party API. This keeps search fully local/private and avoids API keys or rate limits. You need a running SearXNG container with JSON output enabled before these features will work.
docker run -d \
--name searxng \
-p 8888:8080 \
-v "${PWD}/searxng:/etc/searxng" \
--restart unless-stopped \
searxng/searxng:latestOr with docker-compose:
services:
searxng:
image: searxng/searxng:latest
container_name: searxng
ports:
- "8888:8080"
volumes:
- ./searxng:/etc/searxng
environment:
- SEARXNG_BASE_URL=http://localhost:8888/
restart: unless-stoppedBy default SearXNG only serves HTML. LlamaStation needs the JSON API, so edit the generated searxng/settings.yml (created after the first run) and make sure json is listed under search.formats:
search:
formats:
- html
- jsonRestart the container after saving:
docker restart searxngIn the Server tab, set the SearXNG URL field to http://localhost:8888 (default). If SearXNG is running on another machine or a different port, update the URL accordingly. Once set, both the 🔬 Deep toggle in chat and the regular web search button will use it automatically.
💡 Deep Research mode runs multiple rounds of SearXNG queries (reformulating the search based on prior results) before summarizing, so expect it to take noticeably longer than a single web search.
The Telegram bridge (llamastation_telegram.py) reuses the same SearXNG instance for web search inside chats, so no separate setup is needed — just make sure SearXNG is running before starting the bot.
- Set the same
SEARXNG_URL(defaulthttp://localhost:8888) in the bot's config so it can reach the container. - Web search is triggered automatically when the model decides a query needs current information, or manually via the bot's search command.
- If SearXNG is unreachable, the bot falls back to DuckDuckGo so search still works, just without the SearXNG-specific features (JSON parsing, multi-round research).
- Run the bot on the same host/network as the SearXNG container (or make sure the port is reachable) — if you're using Docker networks, use the container name (e.g.
http://searxng:8080) instead oflocalhost.
| Backend | Description |
|---|---|
| ⚡ Official llama.cpp | Standard build — CUDA, Vulkan, CPU. Supports MTP natively since May 2026 |
| 🔬 TurboQuant (TheTom fork) | Asymmetric KV cache quantization — run 200k+ context on 24GB VRAM |
| ⚛️ AtomicChat (TurboQuant + MTP) | TurboQuant + MTP combined — ~22 tok/s on dual RTX 3060 with 177k context |
| 🐝 BeeLlama (DFlash + TurboQuant) | DFlash speculative decoding with TurboQuant — experimental |
Adding your own fork is two lines of code — edit the BACKENDS dict in llamastation.py and point it to your llama-server.exe. Any fork that compiles from llama.cpp works out of the box.
BACKENDS = {
"⚡ Official (llama.cpp)": r"C:\llama.cpp\llama-server.exe",
"🔬 TurboQuant (TheTom fork)": r"C:\llama-turboquant\llama-server.exe",
"🧪 My Custom Fork": r"C:\llama-myfork\llama-server.exe", # ← add yours
}LlamaStation exposes an OpenAI-compatible API at http://localhost:8080/v1. Any app that supports OpenAI works out of the box:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="no-key")
response = client.chat.completions.create(
model="local",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)See the API Docs tab inside the app for more examples.
llamastation/
├── llamastation.py # Main application
├── llamastation_downloader.py # HuggingFace model downloader
├── llamastation_voice.py # Voice mode (Whisper + XTTS v2)
├── llamastation_i18n.py # Translations (ES/EN)
├── llamastation_icon.ico # App icon
├── iniciar_llamastation.bat # Windows launcher
├── llamastation_profiles.json # Per-model settings (auto-generated)
└── llamastation_settings.json # App settings (auto-generated)
Open llamastation_i18n.py, copy the "en" block, rename it to your language code (e.g. "fr"), translate the values, and open a PR. That's it — no other files need to change.
PRs welcome. The codebase is intentionally kept in a small number of files to make it easy to understand and modify. If you add features, please update both "es" and "en" blocks in llamastation_i18n.py.
LlamaStation is a GUI frontend. The actual inference is powered by these open-source projects:
| Project | Author | License | What it does |
|---|---|---|---|
| llama.cpp | Georgi Gerganov / ggml-org | MIT | Core LLM inference engine (official backend) |
| llama-cpp-turboquant | TheTom | MIT | llama.cpp fork with TurboQuant KV cache compression (turbo2/3/4) |
| atomic-llama-cpp-turboquant | AtomicBot-ai | MIT | llama.cpp fork with TurboQuant + MTP |
| coqui-tts (XTTS v2) | Coqui | CPML | Voice cloning and text-to-speech for the Voice tab |
| faster-whisper | SYSTRAN | MIT | Fast speech-to-text using CTranslate2 |
TurboQuant is based on the paper TurboQuant (arXiv:2504.19874, ICLR 2026) by Zirlin et al.
Both backends are MIT licensed — Copyright © 2023-2026 The ggml authors. Full license text is included in the ⚖️ Acerca de tab inside the app.
MIT — do whatever you want with it.
Made with ❤️ for the local AI community
Not affiliated with any AI company




0 comments
log in to comment.