Vitaliy Alexeyev

02 / 06 · Rust · CUDA · NVRTC

synaptix

A native Rust engine for running and training neural networks — hand-written CUDA kernels compiled at runtime through NVRTC, with no PyTorch, no libtorch and no Python runtime.

Licence
MIT OR Apache-2.0
Platform
CUDA sm_80+ · CPU fallback
Packages
Cargo
Status
Young, API not stable
synaptix at work inside synthos: tok/s, prefill time and the prefix-KV share of a chat turn
synaptix at work inside synthos: tok/s, prefill time and the prefix-KV share of a chat turn
lines of Rust
~245k
.cu kernel files, JIT-compiled
73
tests
2 000+
ggml block types executed as they are
27
01

What it is

An alternative to the Python ML stack. Everything from the tensor API and CUDA kernels up to full model ports, a tokenizer, an inference engine with paged KV caches and a training stack is written in Rust.

It runs on any NVIDIA GPU from sm_80 (Ampere) up: kernels are JIT-compiled for the card, with native NVFP4 / MXFP8 block-scale tensor-core paths on Blackwell and portable kernels everywhere else. Models load from single-file .syn bundles or directly from GGUF. Correctness is held to bit-exact parity with PyTorch and NeMo reference implementations — per row, not by a global cosine similarity that hides local errors.

02

Key features

  • Hand-written CUDA kernels, JIT-compiled through NVRTC for the card's compute capability; any sm_80+ GPU, native block-scale mma.sync on Blackwell (sm_120+).
  • Model ports checked against upstream: Qwen3 (dense and MoE), the Qwen3-Next hybrids, Qwen4Exp 125B MoE, Llama, Gemma-3, Gemma-4, Muse Glimmer; FLUX.1, FLUX.2, Qwen-Image 2.1, Qwen-Image-Edit, SDXL; LTX-2.3 and MiniMax-H3 video.
  • Speech and music: Whisper, GigaAM, Sortformer; VoxCPM, OmniVoice, VibeVoice; YuE2, SheetSage2, ACE-Step; BGE-M3 embeddings and reranker.
  • Single-file .syn bundles: mmapped zero-copy, quantized while packing, precision chosen per layer group.
  • SQ1…SQ8, a portable block format with a GPU encoder; GGUF runs directly with all 27 ggml block types; quantized bundles can be transcoded on load for older cards.
  • MXFP8 KV cache by default, paged KV, CUDA-graph decode and speculative decoding.
  • MoE experts in pinned host RAM streaming at ~39 GB/s, plus partial block offload: a 125B MoE runs on a 24 GB card and every model family was measured on 7 GB.
  • Prefix-KV sessions that survive between turns for every architecture, including prompts with images.
  • A CLI for everything: inspect, convert, quantize, run, chat, bench, imagine, video, music, speak, podcast, transcribe.
03

Measured performance

ModelPrefillDecodeNotes
Gemma-4 26B A4B10 100 tok/s @ 4k210 tok/sCUDA-graph decode captures the MoE
Qwen3.8-27B hybrid1 450 tok/s @ 3.3k47 tok/sMTP speculative decode
Qwen3.8-Flash-Next 125B MoE1 650 tok/s @ 260k17–22 tok/s262k context on 24 GB

Gemma-4 decode went 35 → 210 tok/s and prefill 957 → 10 100 tok/s over a week of kernel work. Diffusion at 1024²: SDXL 30 steps in 6.5 s (8.8 GB peak), Qwen-Image-Edit-2511 40 steps in 155 s with NVFP4 (13.2 GB peak).

Measured on an RTX 5090 Laptop GPU (24 GB) with 93 GB of system RAM, Arch Linux.

04

Install

cargo build --release -p synaptix-cli
synaptix convert model.gguf model.syn
synaptix run model.gguf "Explain NVFP4" --max-tokens 256
synaptix bench model.syn --n-tokens 128

Needs the CUDA toolkit at build time (nvcc pins the CUDA version); the driver is loaded dynamically at runtime. The reference tensors for the bit-exact tests are not committed — regenerate them with scripts/reference/.

Full instructions in the README →

05

Requirements and limitations

  • CUDA (primary) and CPU. Any sm_80+ card; native NVFP4 / MXFP8 block-scale MMA needs Blackwell (sm_120+), older cards run the same weights through dequantizing kernels.
  • Young and single-author: the API is not stable, expect breaking changes.
  • Inference is what is production-ready — it powers synthos daily. Training has working autograd, optimizers and checkpointing, but the RLHF, distillation and self-play modules are scaffolding.
  • Some model directories are stubs; only the models listed in the README actually run.
  • Weaker paths are documented, not hidden: bf16 GEMM lands at 0.82–1.16× of cuBLAS depending on shape, behind on large-M tails.
06

Screenshots

Vibe-coded with Claude Code

How it is built

The model writes the code, the tests and most of the documentation. Every number in the README was measured on my machine, and correctness is checked against reference implementations layer by layer rather than taken on the model's word.

Everything here is vibe-coded. github.com/VitaminDB/synaptix