# Amazon — Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

- Company: Amazon (amazon.com)
- Announced: 2026-09-25T16:18:07+00:00
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://aws.amazon.com/blogs/machine-learning/accelerate-multimodal-rl-training-with-skyrl-on-amazon-sagemaker-hyperpod/
- Record: https://forck.live/items/13892-accelerate-multimodal-rl-training-with-skyrl-on-amazon-sagemaker-hyperpod
- Subject: Bedrock / Nova
- Models affected: Qwen3-VL-8B

Amazon published a guide showing how to use SageMaker HyperPod with the open-source SkyRL framework to run GRPO post-training on a Qwen3-VL-8B vision-language model for visual maze navigation. The training improved the maze solve rate from 43.75% to over 95% on a fixed 64-maze evaluation set. The guide covers cluster setup, Ray integration, and monitoring via Amazon Managed Grafana.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/accelerate-multimodal-rl-training-with-skyrl-on-amazon-sagemaker-hyperpod/:

> Starting from the VisGym SFT checkpoint, a supervised fine-tuning (SFT) starting point, GRPO post-training on HyperPod improves the maze solve rate from 43.75% to more than 95% on a fixed 64-maze evaluation set.

---

Record: https://forck.live/items/13892-accelerate-multimodal-rl-training-with-skyrl-on-amazon-sagemaker-hyperpod
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
