Skip to main content

Judges Overview

AgentEval uses LLM-as-judge for all non-deterministic metrics. The judge evaluates agent responses using research-backed prompt templates based on the G-Eval framework.

Supported Providers

ProviderAvailable SinceNotes
OpenAIP0GPT-4o, GPT-4o-mini, GPT-4.1
AnthropicP0Claude Sonnet, Claude Haiku
GoogleP1Gemini 2.0 Flash, Gemini 2.5 Pro
OllamaP1Any locally hosted model — fully private
Azure OpenAIP1Enterprise OpenAI endpoint
Amazon BedrockP1AWS-native model access
CustomP0Implement JudgeModel interface

Quick Configuration

// In AgentEvalConfig
AgentEvalConfig config = AgentEvalConfig.builder()
.judgeModel(JudgeModels.openai("gpt-4o-mini"))
.build();

AgentEval.configure(config);

Or via environment variable (no code change needed):

export AGENTEVAL_JUDGE_PROVIDER=anthropic
export AGENTEVAL_JUDGE_MODEL=claude-haiku-4-5-20251001
export ANTHROPIC_API_KEY=sk-ant-...

Judge Features

Prompt transparency

All judge prompt templates are stored as classpath resources and can be inspected or overridden:

agenteval-judge/src/main/resources/prompts/
answer-relevancy.txt
faithfulness.txt
hallucination.txt
...

Token tracking

Every LLM judge call tracks token usage:

EvalResults results = AgentEval.evaluate(dataset, metrics);
results.judgeTokenUsage(); // total tokens used by judge
results.estimatedCost(); // estimated USD cost

Result caching

Cache judge responses to avoid redundant LLM calls across runs:

AgentEvalConfig.builder()
.cacheResults(true)
.cacheDirectory(".agenteval-cache")
.build();

Cost budget

Abort the evaluation run if costs exceed a threshold:

AgentEvalConfig.builder()
.costBudget(BigDecimal.valueOf(2.00)) // max $2 per run
.build();

Retry on rate limit

AgentEvalConfig.builder()
.retryOnRateLimit(true, 3) // retry up to 3 times on 429
.build();

Parallel Judge Calls

AgentEval evaluates test cases concurrently using virtual threads:

AgentEvalConfig.builder()
.maxConcurrentJudgeCalls(8) // default: available processors
.build();