From the source
Session duration, tool calls executed, completion rates.
But they rarely tell you how it felt.
A user might complete a task successfully in twenty minutes, but did they spend eighteen of those minutes fighting the tool?
We built Signals to answer that question.
Signals uses LLMs as judges to analyze Factory sessions at scale, identifying moments of friction and delight that metrics alone would miss.
More importantly, it does this without anyone ever reading user conversations.
And when friction crosses a threshold, Droid fixes itself.
This is recursive self-improvement: the agent analyzing its own behavior and evolving autonomously.
58 % 83 % 1.3 1.4 Toggle between aggregate view (average sentiment with confidence band) and individual sessions.
Click a session line to see friction and delight moments.
Hover over points to reveal citations.
The Problem with Metrics Consider a session where a developer asks Droid to refactor a module.
The metrics look fine: forty-five tool calls, twelve-minute session, task completed.
…





