From the source
Odyssey introduced PROWL-2, a framework for multi-agent learning from imagined experience that jointly trains a team of agents and its world model using separate curricula within a single training loop.
It combines task-policy curriculum learning, world-model curriculum learning, and a fidelity gate to filter hallucinated rollouts.
In evaluations on SMACv2 and MQE, PROWL-2 achieved the highest mean performance in all nine SMACv2 scenarios and improved success on the hardest MQE tasks from 29.8% to 70.4% on Gate-3 and from 7.3% to 28.2% on Shepherd-Hard.




