Self-Hosted Multi-Model LLM Inference Platform
A dedicated GPU inference server, bare hardware to production workload. Runs quantized LLMs locally — a 27B reasoning/coding model and a 35B vision-capable model — with hot-swap routing via llama-swap. Tuned dual-GPU tensor distribution on ik_llama.cpp, diagnosed utilization imbalance, and caught a precision setting corrupting mixture-of-experts activations.