Evaluation guide · No product scores implied

How to evaluate AI agents: a practical testing checklist

Evaluate an AI agent on the same tasks and inputs you would use in real work. Record correctness, completion, review time, failure recovery, and actual cost. This guide provides a method, not product benchmark results.

Create a small test set

Pick three tasks you already understand: one simple, one representative, and one with an awkward constraint. Use the same input and success criteria for each tool.

A practical scorecard

MeasureWhat to record
CorrectnessErrors, omissions, and unsupported claims
CompletionWhich requested deliverables were actually produced
Review effortMinutes spent correcting or verifying the result
RecoveryWhat happened when a source or action failed
CostActual credits or charges for the complete task

Record the conditions

Save the date, product plan, model if shown, full prompt, and starting files. A changing product cannot be judged fairly from a result with no context.

Publish only what you observed

A failed run is useful evidence. Describe the failure, keep the original output, and separate your interpretation from the facts. Avoid assigning precise scores to products you have not tested.