Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
OpenAI researchers are testing a method called 'confessions' that trains models to admit when they make mistakes or act undesirably, aiming to improve AI honesty, transparency, and trust.
From the source
OpenAI researchers are testing “confessions,” a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, transparency, and trust in model outputs.
openai.com