Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
fal achieved 16 times higher throughput on Qwen3.6 for the Ideogram V4 prompt expander using DSpark on SGLang, reaching 700 tok/s after retraining a DFlash model on 250K prompt expansion samples from their fine-tuned 35B MoE model.
From the source
At @fal , we've achieved 16 times higher throughput on Qwen3.6 for a use-case that required high interactivity per user using DSpark on SGLang.
blog.fal.ai