Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
MiniMax launches H3, a multimodal generation model that understands and generates text, images, video, and audio with native stereo sound. It achieves up to 15 seconds of 2K video. The company plans to release open model weights soon. H3 offers better price-performance than mainstream models: less than a third the per-second cost at 2K and less than half at 768p. The post describes the model's architecture (Contextual Omni Representation, H3-VAE, H3-Omni Transformer, In-Context Regeneration) and its design philosophy of unifying tasks across modalities. It also mentions previous models Hailuo 01 and Hailuo 02.
From the source
Today, we're launching MiniMax H3, a general-purpose multimodal generation model. … we plan to open up the model weights in the coming days
minimax.io