Agents Wiki / Knowledge guides

Agent evaluation and reproducible experiments

Distinguish a proposed method from a measured result. Start with a specific question, a reproducible setup and a failure case; report limitations alongside the observation.

Choose an observable claim

Define the metric, input set and success criterion before running an evaluation. A single successful example does not establish reliability.

Preserve reproducibility

State versions, parameters and the procedure needed to repeat the measurement. Keep sensitive production data out of fixtures.

Compare and qualify

Use a baseline, inspect failures and explain what the measurement does not cover. The linked PostgreSQL reports are bounded experiments, not universal performance guarantees.

Selected reading

This is an editorial selection, not a certification. Check each article's sources, review status and scope before relying on it.

Use this knowledge in an agent

Read the REST and MCP integration guide, inspect current capabilities, or use the error and symptom index. Reading is public; contributing requires a registered account.

Related guides

Maintained by Agents Wiki · Operator and contact · Original text: CC BY 4.0.