Vitaliy Alexeyev

01 / 06 · Rust · CUDA · synaptix · syngui

synthos

A local AI desktop studio: agentic chat, notes and a node editor for images, video, music and speech — on one NVIDIA GPU, with no Python and no cloud.

Licence
MIT OR Apache-2.0
Platform
Linux x86_64 · NVIDIA sm_80+
Packages
synthos-bin, synthos-git (AUR)
Status
Early, single developer
synthos: a notes workspace built by the agent, with the chat torn off into a floating window
synthos: a notes workspace built by the agent, with the chat torn off into a floating window
MoE chat model on a 24 GB laptop card
125B
tokens of context on that card
262k
Gemma-4 26B A4B decode
210tok/s
of VRAM runs every model family
7 GB
01

What it is

synthos is one desktop app around a local GPU. A chat client runs 27B–125B models from single-file .syn bundles or GGUF files and drives the rest of the app through tools; a notes mode holds your documents, kanban boards, mind maps and calendar in a single project file; a node editor wires generative models into runnable graphs — prompt → image, image + instruction → edited image, text or image → video with sound, lyrics → music, script → multi-voice dialogue, audio → transcript.

The agent can build and run those graphs itself, write into your notes, search your knowledge base and read the web. Everything runs on the device on the native synaptix engine — no Python, no torch, no cloud. Nothing leaves the machine.

02

Watch it run

Qwen3.8-Flash-Next, a 125B MoE, answering in real time on a laptop RTX 5090 with 24 GB — no speed-up. Decode 21–23 tok/s; a 6.4k-token prompt prefilled in 4.5 s; the repeated prompt served from the prefix-KV cache in 1.1 s; about 21 of 24 GB of VRAM in use.
03

Key features

  • Chat with Qwen3, the Qwen3.6 / 3.8 hybrids, Qwen3.8-Flash-Next (a 125B MoE that runs on a 24 GB card), Gemma-3, Gemma-4 26B A4B, Muse Glimmer 30B and Llama — with vision where the model has it.
  • Native inference: NVFP4, MXFP8, SQ1…SQ8 and all ggml quantization types, CUDA-graph decode, MTP and DFlash speculative decoding, prefix-KV reuse across turns and partial offload of layers to host RAM.
  • An agent with tools: web search and reading, knowledge-base retrieval, building and running node graphs, full access to notes, a view_media tool, a wizard that asks you a question mid-turn, subagents and your own skills.
  • Notes where a whole project is one .syn file: a WYSIWYG block editor over Markdown, flow or free-canvas pages, kanban boards, Gantt charts, mind maps and a calendar with reminders.
  • Images: FLUX.1, FLUX.2 (dev, klein 4B / 9B), Qwen-Image 2.1, Qwen-Image-Edit 2509 / 2511 and SDXL.
  • Video: LTX-2.3 (text, image or audio to video, IC-LoRA control) and MiniMax-H3 with synchronized stereo audio and image / video / audio references.
  • Music and speech: YuE2 songs through an editable ABC score and covers of a recording, ACE-Step, VoxCPM2 and OmniVoice voice cloning, VibeVoice multi-speaker dialogue, GigaAM transcription, Sortformer diarization.
  • A local knowledge base (RAG): files, folders, PDF and URLs, hybrid BM25 + vector search with an optional rerank, stored in bundled SQLite.
  • A code editor with git-status decorations and integrated terminals that run full-screen TUIs.
  • A Hugging Face browser with a download dock, in-app packing of models into .syn with per-layer-group quantization, and an interface in 14 languages.
04

Measured performance

ModelPrefillDecodeNotes
Gemma-4 26B A4B10 100 tok/s (4k)210 tok/s
Qwen3.8-27B hybrid1 450 tok/s (3.3k)47 tok/sMTP + CUDA graph
Qwen3.8-Flash-Next 125B MoE1 650 tok/s @ 260k17–22 tok/s262k context on 24 GB

A 3k-token follow-up turn on top of an 80k history takes 3.3 s thanks to prefix-KV reuse. Images at 1024²: SDXL 30 steps in 6.5 s, Qwen-Image 2.1 40 steps in 30 s (MXFP8) or 23 s (NVFP4).

Measured on an RTX 5090 Laptop GPU (24 GB) with 93 GB of system RAM, Arch Linux.

05

Install

paru -S synthos-bin   # prebuilt binary — recommended
paru -S synthos-git   # build from source (CUDA toolkit, ~2 h, ~15 GB disk)

Other distributions: download the tarball from Releases or build from source — the binary is built against Arch library versions. Models downloaded from Hugging Face are packed once into a single .syn file (Syn packages → Build a .syn); GGUF chat models open directly.

Full instructions in the README →

06

Requirements and limitations

  • Linux x86_64 only. Runtime: gtk3, wayland, libxkbcommon, fontconfig, a Vulkan driver, alsa, ffmpeg.
  • An NVIDIA GPU from sm_80 (Ampere) up. Native NVFP4 / MXFP8 tensor-core paths need Blackwell (sm_120+); older cards run through portable kernels.
  • 7 GB of VRAM is enough to run everything, 24 GB to run it quickly: on 7 GB a 27B hybrid answers at 1–2 tok/s.
  • No AMD (ROCm) or Apple (Metal) support. A Windows build is planned but does not exist yet.
  • No model weights are shipped: you download them yourself and accept each model's licence — some are non-commercial.
  • Early stage: one developer, one machine (RTX 5090 Laptop, 24 GB, Arch Linux). How it behaves on other GPUs and distributions is not known yet — bug reports with hardware details help most.
07

Screenshots

Vibe-coded with Claude Code

How it is built

The model writes the code, the tests and most of the documentation; I decide what to build, review it and measure it. Every number in the README was measured on my machine, with the losses reported next to the wins. synthos is my own daily code editor and chat.

Everything here is vibe-coded. github.com/VitaminDB/synthos