From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Megakernel serving engine achieves 1.25× to 1.41× faster end-to-end throughput than vLLM
Cohere released a serving engine for North Mini Code built around a decode megakernel, achieving 1.25×–1.41× faster end-to-end performance than vLLM on a single H100 with BF16.
The system supports continuous batching, paged attention, ragged sequence lengths, and an OpenAI-compatible endpoint with tool calling.
From the source
This post presents what we believe is the first fully fledged serving system built around a decode megakernel.
cohere.com