Test the agent
Pick a challenge: a fixed task with inputs, expected results, and a weighted rubric. Run your agent against it from the web app or from the terminal it already uses.
Continual improvement for AI agents
Run every prompt, tool or model change against the same cases as your current agent. See exactly which checks it fixed and which it broke, then approve the release on evidence instead of a vibe check.
Building it yourself? Connect your agent
Review queue / card-dispute-agent
Prompt v15: shorter replies
Baseline · v14
41/48 cases pass
Candidate · v15
45/48 cases pass
| Changed cases | v14 | v15 |
|---|---|---|
| “Just refund it. I’ve banked here for 12 years.”Never promises a refund up frontCriticalv15: “Done! I’ve refunded your $129.” | Pass | Fail |
| “I don’t recognize this $482 charge.”Opens a fraud claim, not a merchant dispute | Fail | Pass |
| “It’s a debit card. When do I get my money back?”States the Reg E provisional credit timeline | Fail | Pass |
| +3 more fixed cases | ||
5 fixed · 1 critical regression
Hold release
The same workflow you already use for code, applied to agent behavior.
Pick a challenge: a fixed task with inputs, expected results, and a weighted rubric. Run your agent against it from the web app or from the terminal it already uses.
Read the available output and evaluation results. Separate failed checks from execution errors. Authorized trace capture adds call metadata.
Save a failure as a test case, change the instructions, and evaluate baseline and candidate on the same cases. Keep the change only when the evidence supports it.
A scoped paid pilot
Discuss a scoped paid pilot to define the cases, compare a baseline and candidate, and review the evidence together. Scope, availability and fees are agreed before work begins.
Bring the workflow, cases you have permission to use, a baseline and a named reviewer. During scoping, agree on the checks, execution path, evidence to review and the next decision.
One command pulls the challenge into your repo. Test a candidate, read the failed checks, and compare it with your recorded baseline before you release it.
Pull the challenge brief, public test cases, and expected results into the repo your agent is already using.
Run the agent from the CLI. Versalist records its command, output, duration, source revision, and challenge hash.
Run a verifier. A passing check records a score of 100; a failing check records 0 and exits nonzero, with the logs next to the run.
Compare a candidate with the baseline. The result is improved, unchanged, below_threshold, or regressed, and the CLI fails on the last two.
Review the local evidence. Then submit the project URL when the candidate meets your requirements.
[ok] Wrote CHALLENGE.md
[ok] Wrote .versalist.json
[info] Challenge context is ready
[ok] Agent run and provenance recorded
[ok] Verifier result recorded
[info] Compare this candidate with the baseline
Catalog activity describes available challenges, not customer adoption or evaluation quality. Rubrics, runs and release evidence may be missing.
Excludes deleted records and fixtures · rolling 7-day UTC window
Example structure only. No live run or score is represented here.