Topic
AI Agent Evaluation & Benchmarking
13 posts on ai agent evaluation & benchmarking — part of ai agents & automation on the n4n AI blog.
Why pass@1 is the wrong metric for agent evaluation
The pass@1 metric agent evaluation standard hides retry behavior and variance. This analysis argues for trajectory-aware scoring in production LLM agents.
WebArena vs WebVoyager: two browsing agent benchmarks
WebArena vs WebVoyager: a practitioner's head-to-head on capabilities, cost, latency, ergonomics, and limits for evaluating web browsing agents.
Tau-bench explained: benchmarking customer service agents
Tau-bench customer service agents benchmark defined: a tool for evaluating LLM agents on realistic multi-turn support tasks with tool use and scoring.
SWE-bench vs SWE-bench Verified: what's the difference
A practitioner's breakdown of SWE-bench vs SWE-bench Verified: dataset construction, eval rigor, cost, and which to use for coding agent benchmarks.
Measuring AI agent reliability across repeated task runs
A practical guide to AI agent reliability measurement across repeated runs: define tasks, instrument traces, score outcomes, and track regressions.
GAIA benchmark explained: testing general AI agents
GAIA benchmark AI agents evaluate real-world assistant capabilities through tool-use tasks. Learn how it works, why it matters, and common myths.
Cost per successful task: a better AI agent benchmark
Stop measuring AI agent quality with abstract benchmarks. Cost per successful task AI agent is the metric that maps to production reality and budgets.
AgentBench: evaluating LLMs as agents across environments
AgentBench LLM agents are evaluated across OS, database, and web environments. This guide explains the benchmark's design, scoring, and practical use.
Why AI agent benchmarks don't predict production results
Benchmarks like SWE-bench ignore latency, fallback, and cost. This analysis explains why AI agent benchmarks production performance fails to predict real deployments.
How to evaluate multi-step agent workflows before shipping
A practical how-to for engineering teams to evaluate multi-step agent workflows pre-production: tracing, replay harnesses, scoring, and CI gates.
How to build a custom eval suite for your AI agent
A practical guide to building a custom eval suite for AI agents: define tasks, run traces, score outputs, and integrate continuous evaluation into CI.
How to benchmark AI agents on real-world tasks
Practical steps to build a reproducible harness to benchmark AI agents real-world tasks: task contracts, sandboxing, scoring, and regression gates.
Claude Opus 4.5 vs GPT-5.1 on SWE-bench Verified
Engineering comparison of Claude Opus 4.5 vs GPT-5.1 coding benchmark on SWE-bench Verified: diff generation, cost, latency, ergonomics, and verdict.
More topics in ai agents & automation
- Function Calling Fundamentals27
- Autonomous Coding Agents: Claude Code, Devin, Cursor15
- Model Context Protocol (MCP) Deep Dives15
- Multi-Agent Orchestration Patterns15
- Agentic RAG14
- AI Agent Cost & Latency Optimization14
- AI Agent Framework Comparison14
- AI Agent Security & Prompt Injection Defense14
- AI Agent Tool Use Design Patterns14
- AI Agents in Customer Support14
- LangGraph for Agent Workflows14
- LLM Workflow Automation: n8n, Zapier, Make14