From the source
Apple researchers propose RISED, a method for training a single LLM agent across diverse interactive environments.
An LLM judge tags rollouts using a predefined rubric vocabulary, guiding data selection and providing token-level supervision via self-distillation.
RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment.