Mushishi Sovereign AI Stack shipped
Self-hosted multimodal LLM infrastructure where 'client data never touches a cloud API' is enforced by routing, not policy.
One RTX 5090, 32GB of VRAM, and three fallback chains with three different guarantees — including a client profile that refuses to fall back, because for paid work, failing loudly beats degrading silently.
Measured: 275 tok/s on Nemotron NVFP4 via vLLM, 180K context at FP8 KV cache. The repo carries the full Decision Log, including the six hours of TensorRT-LLM debugging that ended in choosing vLLM — the dead ends are documented because they were the expensive part.
Posts about this project
- A GTK switch for one GPU shared by seven AI stacks
- I built a three-stack AI system without writing code
- Sovereignty as routing, not policy
- What I designed in May vs what shipped in June
- Moving Docker's data root doesn't move containerd
- Six hours in TensorRT-LLM so you don't have to
- I had my KV-cache math 14× wrong (I treated my Mamba-hybrid like a transformer)
- A 32GB GPU is a budget, not a suggestion
- Your firewall isn't protecting your Docker containers
- My backup failed silently for 17 days