Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
ServiceNow AI introduces EVA, an end-to-end evaluation framework for conversational voice agents that jointly scores task accuracy (EVA-A) and conversational experience (EVA-X). They release an initial airline dataset of 50 scenarios and provide benchmark results for 20 systems, finding a consistent Accuracy-Experience tradeoff.
From the source
We introduce EVA, an end-to-end evaluation framework for conversational voice agents that evaluates complete, multi-turn spoken conversations using a realistic bot-to-bot architecture. EVA produces two high-level scores, EVA-A (Accuracy) and EVA-X (Experience), and is designed to surface failures along each dimension.
huggingface.co