ByteDance has released SeedRealtime, a native audio-visual full-duplex LLM designed for omni-modal natural interaction, featuring joint audio-visual understanding, proactive interaction, and natural…
ByteDance announces Seedance 2.5, a video creation model with up to 30-second single-pass generation, multimodal reference input, and timestamp-level editing.
ByteDance announces Seed Audio 1.0, an audio creation model that generates speech, sound effects, ambience, and other scene-level audio elements within one unified framework.
ByteDance officially launched Seedream 5.0 Pro, a multimodal image creation model with across-the-board improvements and four core capability breakthroughs: complex information visualization,…
ByteDance released EdgeBench, an ultra-long-horizon benchmark for measuring how agents learn from real-world environments, comprising 134 tasks across six domains.
ByteDance announces the official release of the Seed2.1 model family, a new generation of agent-capable models with enhanced general agent capabilities, end-to-end coding delivery, and stronger…
ByteDance released Seed3D 2.0, a next-generation 3D generative model with improved geometric precision and material quality, available via API on Volcano Engine.
ByteDance introduces Seeduplex, a native full-duplex speech LLM that enables simultaneous listening and speaking, achieving breakthroughs in interference suppression and adaptive endpoint detection.
ByteDance Seed announces campus recruitment for 2027 graduates and interns, focusing on foundation models, visual intelligence, speech intelligence, machine learning systems, LLM applications, and…
ByteDance announces the Seed2.0 series of large language models, optimized for large-scale production deployment and featuring improved multimodal understanding, instruction following, and inference…
ByteDance introduces Seedream 5.0 Lite, an image generation model with improved understanding, reasoning, generation, real-time search, and world knowledge.
ByteDance released Seedance 2.0, a video generation model with multimodal input (text, image, audio, video) and enhanced controllability and motion realism.