02 / 06 · Rust · CUDA · NVRTC
synaptix
A native Rust engine for running and training neural networks — hand-written CUDA kernels compiled at runtime through NVRTC, with no PyTorch, no libtorch and no Python runtime.
- Licence
- MIT OR Apache-2.0
- Platform
- CUDA sm_80+ · CPU fallback
- Packages
- Cargo
- Status
- Young, API not stable

What it is
An alternative to the Python ML stack. Everything from the tensor API and CUDA kernels up to full model ports, a tokenizer, an inference engine with paged KV caches and a training stack is written in Rust.
It runs on any NVIDIA GPU from sm_80 (Ampere) up: kernels are JIT-compiled for the card, with native NVFP4 / MXFP8 block-scale tensor-core paths on Blackwell and portable kernels everywhere else. Models load from single-file .syn bundles or directly from GGUF. Correctness is held to bit-exact parity with PyTorch and NeMo reference implementations — per row, not by a global cosine similarity that hides local errors.
Key features
- Hand-written CUDA kernels, JIT-compiled through NVRTC for the card's compute capability; any sm_80+ GPU, native block-scale mma.sync on Blackwell (sm_120+).
- Model ports checked against upstream: Qwen3 (dense and MoE), the Qwen3-Next hybrids, Qwen4Exp 125B MoE, Llama, Gemma-3, Gemma-4, Muse Glimmer; FLUX.1, FLUX.2, Qwen-Image 2.1, Qwen-Image-Edit, SDXL; LTX-2.3 and MiniMax-H3 video.
- Speech and music: Whisper, GigaAM, Sortformer; VoxCPM, OmniVoice, VibeVoice; YuE2, SheetSage2, ACE-Step; BGE-M3 embeddings and reranker.
- Single-file .syn bundles: mmapped zero-copy, quantized while packing, precision chosen per layer group.
- SQ1…SQ8, a portable block format with a GPU encoder; GGUF runs directly with all 27 ggml block types; quantized bundles can be transcoded on load for older cards.
- MXFP8 KV cache by default, paged KV, CUDA-graph decode and speculative decoding.
- MoE experts in pinned host RAM streaming at ~39 GB/s, plus partial block offload: a 125B MoE runs on a 24 GB card and every model family was measured on 7 GB.
- Prefix-KV sessions that survive between turns for every architecture, including prompts with images.
- A CLI for everything: inspect, convert, quantize, run, chat, bench, imagine, video, music, speak, podcast, transcribe.
Measured performance
| Model | Prefill | Decode | Notes |
|---|---|---|---|
| Gemma-4 26B A4B | 10 100 tok/s @ 4k | 210 tok/s | CUDA-graph decode captures the MoE |
| Qwen3.8-27B hybrid | 1 450 tok/s @ 3.3k | 47 tok/s | MTP speculative decode |
| Qwen3.8-Flash-Next 125B MoE | 1 650 tok/s @ 260k | 17–22 tok/s | 262k context on 24 GB |
Gemma-4 decode went 35 → 210 tok/s and prefill 957 → 10 100 tok/s over a week of kernel work. Diffusion at 1024²: SDXL 30 steps in 6.5 s (8.8 GB peak), Qwen-Image-Edit-2511 40 steps in 155 s with NVFP4 (13.2 GB peak).
Measured on an RTX 5090 Laptop GPU (24 GB) with 93 GB of system RAM, Arch Linux.
Install
cargo build --release -p synaptix-cli
synaptix convert model.gguf model.syn
synaptix run model.gguf "Explain NVFP4" --max-tokens 256
synaptix bench model.syn --n-tokens 128Needs the CUDA toolkit at build time (nvcc pins the CUDA version); the driver is loaded dynamically at runtime. The reference tensors for the bit-exact tests are not committed — regenerate them with scripts/reference/.
Requirements and limitations
- CUDA (primary) and CPU. Any sm_80+ card; native NVFP4 / MXFP8 block-scale MMA needs Blackwell (sm_120+), older cards run the same weights through dequantizing kernels.
- Young and single-author: the API is not stable, expect breaking changes.
- Inference is what is production-ready — it powers synthos daily. Training has working autograd, optimizers and checkpointing, but the RLHF, distillation and self-play modules are scaffolding.
- Some model directories are stubs; only the models listed in the README actually run.
- Weaker paths are documented, not hidden: bf16 GEMM lands at 0.82–1.16× of cuBLAS depending on shape, behind on large-M tails.
Screenshots
- (open full size)

Packing a .syn bundle: components, auxiliary files and precision per layer group (synthos UI) - (open full size)

ACE-Step text → music graph running on synaptix (synthos node editor) - (open full size)

The agent checks VRAM, frees it and starts a video run (synthos)
Vibe-coded with Claude Code
How it is built
The model writes the code, the tests and most of the documentation. Every number in the README was measured on my machine, and correctness is checked against reference implementations layer by layer rather than taken on the model's word.
Everything here is vibe-coded. github.com/VitaminDB/synaptix