Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Together AI discusses its approach to efficient inference at scale, highlighting research contributions like FlashAttention-4, ThunderKittens, and Aurora, and its full-stack hardware optimization on NVIDIA Blackwell hardware. The post covers the economics of inference, but does not announce a new product, model, or specific pricing change; it is a general thought-leadership piece about inference infrastructure.
From the source
For Together AI, none of this is new. The inference imperative is what we’ve been building for.
together.ai