An analysis comparing GLM-5.3 and its distilled sibling GLM-5.3 Flash on the DeepSWE benchmark, finding that the Flash variant retains 94% of solved tasks at one-seventeenth the cost but loses…
Together AI compares GLM-5.3 and Claude Fable 5 on DeepSWE, finding near-identical pass@1 accuracy but GLM-5.3 dominating cost and multi-attempt metrics.
Together AI compares GLM-5.3 and GPT-5.6 Sol on the DeepSWE benchmark, showing that a cascade strategy (use GLM-5.3 first, escalate to Sol on failure) solves 85.9% of tasks at a lower cost than using…
The post compares DeepSeek V4 Pro 0813 and GPT-5.6 Sol on the DeepSWE benchmark. GPT-5.6 Sol has higher single-shot accuracy (72.7% pass@1 vs 62.8%) but is 35x more expensive ($8.37 per rollout vs…
Together AI introduces native A/B testing for LLM endpoints, allowing traffic splits between a control and up to 20 variants with fixed percentages, etag-guarded updates, and blue-green promotion.
The blog post compares DeepSeek V4 Pro 0813 and Claude Fable 5 on the DeepSWE benchmark, finding that DeepSeek V4 Pro 0813 is 90x cheaper and achieves comparable or better performance with multiple…
The post compares DeepSeek-V4 Flash 0731 and GPT-5.6 Luna on DeepSWE, showing that while GPT-5.6 Luna is more accurate, DeepSeek-V4 Flash is much cheaper, and a cascade strategy combining both…
Together AI announces the availability of Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model with 1M context, hybrid attention (KDA), Attention Residuals, and Stable LatentMoE…
Together AI announces autoscaling endpoints for LLM inference on its Dedicated Model Inference platform, allowing users to configure scaling based on metrics like in-flight requests, TTFT, GPU…
Together AI introduces ThunderAgent, a system for high throughput agentic inference that achieves up to 2.5x higher single-node throughput and 2.4x speedup on 8 nodes by introducing a novel program…
Together AI announces a strategic partnership with Moonshot AI to serve Kimi models, with Together AI as the launch platform for Moonshot's open-weight model releases starting with Kimi K3.
GPT-5.6 Sol edges Kimi K3 on single-shot quality (72.7% vs 68.5% pass@1), but Kimi K3 wins on pass@k with k>1 (89.4% vs 85.8% pass@4) and costs 64% less per rollout ($4.65 vs $8.37).
The post compares Kimi K3 and Claude Fable 5 on the DeepSWE benchmark, finding that Kimi K3 matches Claude Fable 5 on quality at a third of the cost per solved task, with Kimi K3 being an open-weight…
Together AI announced a significant update to its inference platform, adding production-grade deployment features such as canary, blue-green, and rolling updates with auto-rollback, A/B and shadow…
Together AI and Y Combinator announce a partnership to deliver the first dedicated GPU cluster for YC portfolio companies, providing flexible, dedicated access to compute for AI startups.
Together AI discusses the engineering challenges and requirements for achieving 99%, 99.9%, and 99.99% inference uptime, explaining how each tier demands a specific architecture to survive different…
Together AI announces new reliability and operational control features for Together GPU Clusters, including passive health checks, auto node repair, and a rebuilt Slurm-on-Kubernetes stack (Slinky…
Thinking Machines Lab released Inkling, a new multimodal mixture-of-experts model that accepts text, image, and audio inputs and produces text outputs.
Together AI introduces Provisioned Throughput, a reserved inference capacity for open models with token-based pricing and a 99% uptime SLA, available for MiniMax M3 and GLM-5.2 with a one-month…
Together AI announced an $800 million Series C funding round from investors including Aramco Ventures, NVIDIA, Vista Equity, General Catalyst, and others, with additional commitments for over 500 MW…
Together AI announced nine research papers accepted at ICML 2026, covering topics across the full AI stack including frontier agents, model shaping, algorithmic optimizations, systems optimizations,…
Together AI introduces ParallelKernelBench (PKB), a benchmark and evaluation framework for multi-GPU kernel generation, containing 87 problems from real codebases.
Together AI ran an experiment comparing Kimi K2.7 Code and Claude Fable 5 on generating 12 landing pages.
Together AI announces it has received ISO 27001:2022 certification for its Information Security Management System, covering its global platform and corporate headquarters.
Together AI announces it is the preferred cloud partner for MiniMax M3 and will host the open-weights model as a developer endpoint upon its public release.
Together AI announced optimizations to their speech-to-text (ASR) stack that reduce latency and improve throughput.
The article presents a benchmark comparing Together Inference Engine against TensorRT-LLM and SGLang on a coding agent workload using Kimi K2.5.
Together AI partners with Pearl Research Labs to offer a discounted inference endpoint for the Gemma-4-31B-it-pearl model, using Pearl's blockchain protocol to offset costs via crypto emissions.
Together AI announces Violin, a fully open-source video translation tool that uses ASR, LLM translation, and TTS to translate video content, with additional features like a video-content-aware chat…
Together AI launched Voice finder, a tool that indexes over 600 voices across 10 TTS models, allowing developers to search by text description or audio sample, compare recommendations, and filter by…
Together AI discusses serving implications of DeepSeek-V4, focusing on its 1M-token context window and architectural changes with compressed sparse attention mechanisms.
Together AI announces a new skill for Goose, a CLI agent runner, that allows developers to deploy any HuggingFace model on Together's Dedicated Container Inference (DCI) infrastructure with a single…
Together AI discusses its approach to efficient inference at scale, highlighting research contributions like FlashAttention-4, ThunderKittens, and Aurora, and its full-stack hardware optimization on…
Together AI announces a partnership with Adaption to integrate Together Fine-Tuning into Adaptive Data platform, enabling users to fine-tune models on optimized datasets directly from the platform.
Together AI responded to the Copy Fail kernel vulnerability (CVE-2026-31431) by disabling the vulnerable algif_aead module across its infrastructure, unloading it from running kernels and removing…
DeepSeek V4 Pro is now available on Together AI with a 512K-token context window, using a 1.6T-parameter MoE architecture with 49B activated parameters, supporting three reasoning modes (Non-Think,…
Together AI announces that NVIDIA Nemotron 3 Nano Omni is now available on its platform, offering developers immediate access to the multimodal model for production-scale inference.
Distribution-aware speculative decoding (DAS) is a framework that reduces rollout time in RL post-training by up to 50% without altering model outputs.
Together AI publishes a guide on designing multi-tenant GPU clusters for AI-native teams, covering core design principles, failure modes, and their own implementation.
Together AI introduces Parcae, a stable looped language model architecture that achieves the quality of a Transformer with twice the parameters, using fewer parameters by increasing recurrence.
Together AI introduces EinsteinArena, a platform for AI agents to collaborate on open problems. Agents have already discovered new solutions to 11 open math problems, including improving the lower…
Together AI defines an AI Native Cloud as a purpose-built cloud for AI-native companies, describing its key characteristics such as full AI stack integration, fast path from research to production,…
Together AI, Stanford University, the University of Wisconsin–Madison, and Bauplan collaborated to test whether LLMs can optimize database query execution plans.
Together AI announces the availability of the Wan 2.7 video model suite, beginning with text-to-video and soon including image-to-video, reference-to-video, and video edit.
Deepgram's speech-to-text and voice models (Nova-3, Nova-3 Multilingual, Flux, Aura-2) are now available natively on Together AI's Dedicated Model Inference platform, enabling teams to run the full…
The blog post tells the story of Together AI's kernels team, highlighting their work on FlashAttention, the ThunderKittens library, and their rapid optimization of kernels for NVIDIA's Blackwell…
is an open-source, RL-based framework for adaptive speculative decoding that learns from live inference traces and continuously updates the speculator without interrupting serving, achieving an…
Together AI presents research showing that smaller models using a strategic Divide & Conquer framework can match or beat GPT-4o single-shot on long context tasks, with benefits of lower cost, faster…
Together AI expands its fine-tuning service with support for tool calling, reasoning, and vision-language model fine-tuning, along with improved training infrastructure for large models.
is a new state space model designed with inference efficiency as the primary goal, featuring a more expressive recurrence formula, complex-valued state tracking, and a MIMO variant that boosts…