Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Google Research introduces a systematic evaluation framework that converts psychological questionnaires into situational judgment tests to measure how closely LLMs' behavioral dispositions align with human consensus. Analysis of 25 LLMs reveals gaps where model behavior deviates from human agreement or fails to capture the range of human opinions.
From the source
we introduce a systematic evaluation framework that transforms established assessments into large-scale situational judgment tests for large language models.
research.google