# Agent evaluation evidence — Agent Evidence

Canonical: https://evidence.agiscorecard.com/guide.html
Reviewed: 2026-09-19
Author: AGI Scorecard team
Version: 1.1.0

Agent Evidence reviews imported task outcomes within one task, version and date window. It deduplicates sample labels and shows failures, exclusions and uncertainty. It helps review a supplier comparison; it does not run evaluations or verify reviewer identities.

## Method

Filter to a dated task/version cohort. Exclude self/unknown relations and future/stale records. Collapse repeated sample labels, treating conflicts conservatively; show a descriptive 95% Wilson interval.

## Workflow

1. Use a fixed set of tasks with explicit acceptance rules. Export actual outcomes.
2. Declare reviewer relationships and dates. Keep failures and disagreements.
3. Review sample coverage and uncertainty before comparing suppliers.

## Common mistake

Treating two reviews of one case as two independent tasks overstates the evidence. This worksheet collapses repeated sample labels and exposes conflicts.

## Worked scenarios

### Repeated task samples

See how the evidence denominator changes after deduplication.
Fictional inputs: https://evidence.agiscorecard.com/examples/duplicates.json
Run: https://evidence.agiscorecard.com/?scenario=duplicates#workbench
- Providers: 1
- Excluded records: 1
- Sample sets: Aligned labels only

### Reviewers disagree

A conflicting result remains a conflict, not an extra success.
Fictional inputs: https://evidence.agiscorecard.com/examples/conflict.json
Run: https://evidence.agiscorecard.com/?scenario=conflict#workbench
- Providers: 1
- Excluded records: 1
- Sample sets: Aligned labels only

### Wrong agent version

A result from a different version cannot improve this cohort.
Fictional inputs: https://evidence.agiscorecard.com/examples/version.json
Run: https://evidence.agiscorecard.com/?scenario=version#workbench
- Providers: 1
- Excluded records: 2
- Sample sets: Aligned labels only

## Questions

### Can I compare agents using a single success percentage?

Only after checking task coverage, versions, dates and the denominator. Duplicate samples, self-reviews and selection bias can make the same percentage describe very different evidence. A small sample should remain visibly uncertain.

### Is Agent Evidence free, and do I need a wallet?

All current functions are free beta. No account, wallet connection, subscription or model API key is needed. No public reputation ranking, identity verification, calibrated prediction or Sybil detection.

### Can I use my own records and keep them private?

Yes. Enter records, import JSON or paste CSV into record groups. Inputs and comparison snapshots stay in this browser tab. Share-example links contain only a public scenario name. Review downloaded reports before sharing your records.

## Alternatives and official sources

- [8004scan](https://8004scan.io/): Agent discovery, network exploration, rankings and ecosystem health signals.
- [Promptfoo](https://www.promptfoo.dev/docs/intro/): Local and CI evaluation, assertions, provider comparisons and red teaming.
- [ERC-8004](https://eips.ethereum.org/EIPS/eip-8004): Public identity, reputation and validation signals; its specification describes Sybil/spam concerns.

Use Promptfoo to run evaluations and a registry explorer to discover agents. This tool reviews an imported evidence set; it does not run agents or establish trusted identities.

## Limits

No public reputation ranking, identity verification, calibrated prediction or Sybil detection.
The Wilson interval is a descriptive binomial interval. It assumes independent representative samples; shared reviewers, copied tasks and selection bias violate that assumption. It is not a calibrated reputation score.

Reviewer independence, task IDs and outcomes are declared by the submitter. No identity, signature, ownership, Sybil resistance or future performance is verified.

Matching sample labels do not prove matching inputs. Compare raw task fixtures before selecting a supplier.

## Input and output

Structural schema: https://evidence.agiscorecard.com/input.schema.json
The engine additionally validates ranges, dates, amount formats, uniqueness and cross-field rules.
Input JSON and CSV paste are processed locally. No model, account or wallet is needed.
Reports can be downloaded as JSON or Markdown and printed locally. Public scenario links never include user records.

## Price and license

Current functions are free beta. No paid plan or update SLA is offered.
Tool bundle license: https://evidence.agiscorecard.com/LICENSE.txt
Privacy: https://evidence.agiscorecard.com/privacy.html
