Agent evaluation evidence
See what an agent score leaves out.
Review task-specific agent evaluations with sample deduplication, version boundaries and uncertainty intervals.
Use my own recordsLocal processing · No account · No wallet connection
What would you like to check?
Choose a fictional scenario to run it immediately, then change the assumptions.
Make it your own
Edit the form or import the example-shaped JSON. Input reference
Edit the inputs
Inputs stay in this tab and clear on reload. No automatic upload or wallet access. Maximum 128 KiB.
Reports include your record labels. Review them before sharing.
Report text for manual copy
From your records to a reviewable result
- Use a fixed set of tasks with explicit acceptance rules. Export actual outcomes.
- Declare reviewer relationships and dates. Keep failures and disagreements.
- Review sample coverage and uncertainty before comparing suppliers.
Where this tool fits
A compact review of evidence quality: what was tested, which version, what was excluded and how little a small sample tells you.
Scope: No public reputation ranking, identity verification, calibrated prediction or Sybil detection.
Agent 任务证据评测:可直接使用表单,也可导入 JSON。先运行虚构示例理解结果,再切换到自己的材料。输入仅在当前页面处理。查看中文说明。
Compare with established tools
Use Promptfoo to run evaluations and a registry explorer to discover agents. This tool reviews an imported evidence set; it does not run agents or establish trusted identities.
Official product descriptions reviewed 19 September 2026. These are alternatives, not partners or endorsements.
What does Agent Evidence do?
Agent Evidence reviews imported task outcomes within one task, version and date window. It deduplicates sample labels and shows failures, exclusions and uncertainty. It helps review a supplier comparison; it does not run evaluations or verify reviewer identities.
By the AGI Scorecard team. Method and sources reviewed . Read the method.
Questions before you start
Can I compare agents using a single success percentage?
Only after checking task coverage, versions, dates and the denominator. Duplicate samples, self-reviews and selection bias can make the same percentage describe very different evidence. A small sample should remain visibly uncertain.
Is Agent Evidence free, and do I need a wallet?
All current functions are free beta. No account, wallet connection, subscription or model API key is needed. No public reputation ranking, identity verification, calibrated prediction or Sybil detection.
Can I use my own records and keep them private?
Yes. Enter records, import JSON or paste CSV into record groups. Inputs and comparison snapshots stay in this browser tab. Share-example links contain only a public scenario name. Review downloaded reports before sharing your records.
Use this workflow in your AI assistant
Connect the MCP server to read sources and run these calculations from a supported client. Remote calls send parameters to the server; the browser worksheet remains local.
Get the MCP connection and citation guide