From the source
They describe the goal in natural language — without a repo, test suite, or chosen framework — and expect the agent to turn it into a functioning app.
The result might be a website, slide deck, mobile app, several connected artifacts, or something else entirely.
Vibe coders are not usually checking diffs or test output.
Success for Replit Agent is deceptively simple: the app should work when users click around.
That changes the job of evaluation.
A single score can help with a specific shipping decision, but it cannot tell us, week over week, whether Replit Agent is getting better for users.
To answer that question, evaluation must become part of the improvement loop.
Evaluation has to do more now Agent evaluation used to look like a one-way process: run the eval, produce a score, and make a shipping call.
This works when releases are slow and the thing being measured rarely changes.
It breaks down when models, prompts, tools, and product surfaces are all changing quickly.
The old loop made evaluation feel bounded.
But Replit Agent changes too quickly for a single score to carry the whole decision.
…





