From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Guide for GRPO post-training on SageMaker HyperPod
Amazon published a guide showing how to use SageMaker HyperPod with the open-source SkyRL framework to run GRPO post-training on a Qwen3-VL-8B vision-language model for visual maze navigation.
The training improved the maze solve rate from 43.75% to over 95% on a fixed 64-maze evaluation set.
The guide covers cluster setup, Ray integration, and monitoring via Amazon Managed Grafana.
From the source
Starting from the VisGym SFT checkpoint, a supervised fine-tuning (SFT) starting point, GRPO post-training on HyperPod improves the maze solve rate from 43.75% to more than 95% on a fixed 64-maze evaluation set.
aws.amazon.com