Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post describes an experiment testing how well LLMs can fix their mistakes when given feedback in plain English, using a simple calendar API scenario. The author built a chatbot arena using Keras, JAX, and TPUs to interact with multiple LLMs simultaneously.
From the source
I decided to run a little test with today's LLMs. A super-simplified one, to see how effectively LLMs fix their mistakes when you point them out to them.
huggingface.co