From the source
Decision Transformer (DT) and Generalized DT, which reformulate reinforcement learning as a sequence modeling problem using a transformer architecture, enabling stable policy learning via supervised learning without temporal difference learning.




