# Cohere — Inside the megakernel serving engine for North Mini Code

- Company: Cohere (cohere.com)
- Announced: 2026-09-08
- Category: infrastructure-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://cohere.com/blog/megakernels
- Record: https://forck.live/items/9285-inside-the-megakernel-serving-engine-for-north-mini-code
- Subject: Command
- Models affected: North Mini Code

Cohere released a serving engine for North Mini Code built around a decode megakernel, achieving 1.25×–1.41× faster end-to-end performance than vLLM on a single H100 with BF16. The system supports continuous batching, paged attention, ragged sequence lengths, and an OpenAI-compatible endpoint with tool calling.

## Evidence

Verbatim from https://cohere.com/blog/megakernels:

> This post presents what we believe is the first fully fledged serving system built around a decode megakernel.

---

Record: https://forck.live/items/9285-inside-the-megakernel-serving-engine-for-north-mini-code
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
