From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
9.3B parameter open-weight text-to-image model with structured JSON prompting
Ideogram released Ideogram 4.0, a 9.3B parameter open-weight text-to-image model.
The model uses a 34-layer Diffusion Transformer with a Qwen3-VL-8B-Instruct text encoder and is trained exclusively on structured JSON captions with per-element styling, optional bounding boxes, and color palettes.
The reference inference pipeline validates prompts against the JSON schema before generation.
From the source
Ideogram 4.0 is a 9.3B parameter open-weight text-to-image model. Recent open-weight releases have converged on a single self-attention sequence over text and image tokens [1] [2] [3] , and Ideogram 4.0 follows the same pattern: text and image tokens share the same projections at every layer of a 34-layer DiT.
ideogram.ai