Case studies
How Cekura cut p95 latency in half on Gemma 4 26B with Autoloops
Cekura moved Gemma 4 26B voice-agent testing to Autoloops for 2.3× faster p50 and 2.5× faster p95 than DeepInfra Turbo.
Read article →Meet Hanoi, our ultra-fast inference engine
How we specialized Gemma 4 26B inference for voice workloads and beat tuned vLLM on TTFT at 400 concurrent calls.
Read article →