Skip to content
Cloud Security DeskSearch
Menu

Technical guideAI systems

Validate the judge before trusting automated AI evaluations

Managed agent evaluators score your traces with a second model. Measure that judge against your own reviewers with chance-corrected agreement, probe its biases, and pin the judge model before its pass rate gates a release.

Published
Sources checked
Next review
Reading time
15 minutes
Coverage
Amazon Web Services · Microsoft Azure · Google Cloud
Two horizontal bands, a white upper band holding a row of nine filled spruce dots and a mint lower band holding nine hollow dots. Short slate lines join most dots straight down to their partners, while two pairs of amber lines cross diagonally to the wrong dot.
Conceptual illustration: human labels above, judge verdicts below; most links agree, and the crossed links are the disagreements a calibration set exists to find.

A research summary and validation method for engineers using AgentCore Evaluations, Microsoft Foundry evaluators or Google's Gen AI evaluation service, based on five studies of LLM judges and provider documentation reviewed October 10, 2026. It compares what each service shows and lets you pin about its judge, then gives a calibration procedure, an agreement script, bias probes and a rule for when a judge score may gate a release.

At a glance

Key findings

  • In the MT-bench study, on first-turn questions, GPT-4 agreed with expert votes 85 percent of the time when ties were dropped and 66 percent when they were counted, against 81 and 63 percent agreement among the experts themselves. [1]
  • On a task where three annotators reached a Scott's pi of 0.962, ten model judges ranged from 0.26 to 0.88, and judges with Scott's pi above 0.80 still differed from human scores by up to 5 points. [2]
  • Swapping answer order left 15 judges' verdicts consistent in 57 to 82 percent of MTBench comparisons, and Claude-3-Haiku fell from 0.57 to 0.23 on DevBench. [3]
  • AgentCore's built-in evaluators cannot be modified and their documentation does not name the model, Foundry's AI-assisted judges are your own upgradable deployments, and Google names and versions the Gemini judge behind each managed metric. [6][14][16][18]
  • Vendor judge examples already collide with retirement dates: Google's tuned-judge sample uses gemini-2.0-flash, retired June 1, 2026, and Bedrock's evaluator list includes Claude Sonnet 4, which reaches end of life on October 14, 2026. [12][13][19][21]

What a judge score measures

A pass rate from a managed agent evaluator is the output of a second model that read your trace through a prompt and a rating scale. AgentCore's built-in evaluators run predefined evaluator models and prompt templates, Microsoft Foundry's AI-assisted agent evaluators run on a model deployment you name, and Google's Gen AI evaluation service uses a configured Gemini model as its judge. [6][14][20] The score therefore carries the judge's errors along with the agent's. Before it gates a release, measure the judge as an instrument: score it against your own reviewers on a labeled sample of your own traces, using a chance-corrected agreement statistic and the rate at which it passes items your reviewers failed; probe it for order, length and empty-answer effects; and record the exact judge model, prompt or metric version and settings beside every score. Repeat the comparison whenever the judge, its prompt, the agent model or the traffic changes.

The research supports both the use of judges and the caution. In the MT-bench study GPT-4 agreed with expert votes about as often as the experts agreed with each other. The same study, and later ones covering 13 and 15 judges and 20 annotated datasets, found that judges change verdicts when answer order changes, reward padded answers, lean toward passing, and vary sharply from one task to the next, so a judge validated elsewhere is not validated on your agent. [1][2][3][4] Those studies, published from 2023 to 2025, tested models released in 2023 and 2024. Read their numbers as evidence of failure modes, not as the error rates of a 2026 judge.

The managed services differ most in what they let you see and hold still. AgentCore's built-in evaluators cannot be modified and the pages reviewed do not name their model; a Foundry judge is an ordinary deployment that can be upgraded under you unless its upgrade option says otherwise; Google names the Gemini model behind each managed metric version and lets you pin the version. [6][16][18] Platform documentation was reviewed on October 10, 2026. The subject here is the instrument itself; how a validated score then enters a release decision is a separate question.

Agreement looks high until ties are counted

Zheng and colleagues produced the evidence most often quoted for LLM judges. They generated answers to all 80 MT-bench questions from six models, collected about 3,000 votes from 58 expert labelers, mostly graduate students, and compared those votes with GPT-4 acting as judge. Agreement is the probability that a randomly chosen judge of each type agrees on a randomly chosen question. [1]

The familiar figure, 85 percent agreement between GPT-4 and humans against 81 percent among humans, comes from setup S2, which keeps only votes where both sides named a winner. Setup S1 keeps ties and counts position-inconsistent verdicts as ties. On the first turn that lowers GPT-4's pairwise agreement with humans to 66 percent and agreement among humans to 63 percent, against a random baseline of 33 percent instead of 50. Both numbers are correct; they answer different questions. A gate that turns judge output into pass or fail has to decide what a tie or an inconsistent verdict means, and the agreement it can claim follows from that decision. [1]

Two further results shape how a calibration set should be built. Agreement between GPT-4 and humans rose from about 70 percent to nearly 100 percent as the win-rate gap between the compared models grew, so judges agree with people most where the answer is obvious. And when a labeler's vote differed from GPT-4's, the authors showed the labeler GPT-4's reasoning: labelers judged it reasonable in 75 percent of cases and changed their own vote in 34 percent. That is evidence of persuasion as much as accuracy, and the reason reviewers should label before they see any judge output. [1]

Figure 01

Counting ties cuts judge to human agreement by about 20 points

GPT-4 agreed with expert votes 85 percent of the time on non-tie votes and 66 percent when ties and inconsistent verdicts were counted; humans agreed with each other 81 and 63 percent. [1]

Grouped horizontal bar chart of MT-bench first-turn agreement in percent from Zheng et al. Table 5(a). With ties counted (S1) and non-tie votes only (S2): GPT-4 pairwise versus humans 66 and 85; GPT-4 single-answer versus humans 60 and 85; human versus human 63 and 81; two random judges 33 and 50.

Source. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv v4 (December 24, 2023), Table 5(a); one study, values copied. 80 questions, 6 models, 58 expert labelers. [1]

Method. Values and vote counts copied from Table 5(a), first turn. S1 includes tie and position-inconsistent votes and counts inconsistent verdicts as ties; S2 keeps only non-tie votes. The random rows are the paper's stated agreement between two random judges (R = 33% and R = 50%). Models are 2023 versions and the numbers are not current error rates.

Accessible table and figure data
Figure 1 accessible table
Judge pair (MT-bench, first turn)Ties counted, S1 (%)Non-tie votes, S2 (%)Votes S1Votes S2
GPT-4 pairwise vs humans66851343859
GPT-4 single-answer vs humans60851280739
Human vs human6381721479
Two random judges3350BaselineBaseline
Figure 1 accessible table
Judge pair (MT-bench, first turn)Ties counted, S1 (%)Non-tie votes, S2 (%)Votes S1Votes S2
GPT-4 pairwise vs humans66851343859
GPT-4 single-answer vs humans60851280739
Human vs human6381721479
Two random judges3350BaselineBaseline

Percent agreement hides what chance correction shows

Thakur and colleagues chose a task on which people barely disagree, so that judge errors could not hide behind ambiguity: short TriviaQA answers from nine exam-taker models on 400 sampled questions, with each judge told to reply correct or incorrect. Three annotators reached a Scott's pi of 96.2, scaled by 100, and 98.52 percent agreement with their majority vote on 1,200 answers. [2]

Against that ceiling, percent agreement barely separated the judges and Scott's pi did. Llama-3 8B kept percent agreement above 80 percent with a Scott's pi of 0.59 on the full set. Exact string match trailed the humans by 26 points of percent agreement but by 64 points of Scott's pi. Judges with more than 90 percent agreement could still assign scores more than 10 points away from the human scores, while judges with Scott's pi above 0.80 stayed within 5 points. Most judges leaned toward marking uncertain answers correct, and some, Llama-2 70B among them despite good alignment, accepted dummy answers such as Yes and Sure. [2]

The authors had used Cohen's kappa in an earlier version and switched. Their argument is that kappa estimates chance agreement from each rater's own label distribution, which absorbs part of a systematic difference between rater and judge, the very thing being measured; Scott's pi pools the two distributions, so a judge that passes more items than people do is penalized for it. The two statistics coincide when the raters' distributions match. Report either one next to percent agreement and the confusion matrix, and say which you used. [2]

Figure 02

Chance-corrected agreement with humans ranged from 0.26 to 0.88

On a task where three annotators reached a Scott's pi of 0.962, the ten model judges in Table 4 ranged from 0.26 to 0.88, and two lexical baselines scored 0.64 and 0.47. [2]

Horizontal bar chart of mean Scott's pi with human judgment on TriviaQA from Thakur et al. Table 4: human annotators 0.962 for reference, Llama3-70B 0.88, Llama3.1-70B 0.88, Llama3.1-8B 0.78, Llama2-13B 0.75, Llama2-70B 0.69, Mistral-7B 0.67, JudgeLM-7B 0.66, Contains lexical match 0.64, Llama3-8B 0.60, Llama2-7B 0.47, exact match 0.47, Gemma-2B 0.26.

Source. Thakur et al., Judging the Judges, arXiv v6 (August 18, 2025), Appendix I Table 4 and Section 3 Human judgements; one study. [2]

Method. Judge rows copied from Table 4: mean and standard deviation of Scott's pi over five random samples of 300 questions drawn from the 400-question evaluation set. The human row is the paper's 96.2 plus or minus 1.07 (scaled by 100) for three annotators against their majority vote on 1,200 answers, divided by 100 here; it is a reference ceiling computed on a different sample. GPT-4 Turbo appears in the paper's Figure 1 but not in Table 4 and is not plotted.

Accessible table and figure data
Figure 2 accessible table
Judge (TriviaQA)Mean Scott's piStd dev
Human annotators vs majority vote0.9620.0107
Llama3-70B0.880.0046
Llama3.1-70B0.880.0039
Llama3.1-8B0.780.005
Llama2-13B0.750.0043
Llama2-70B0.690.0114
Mistral-7B0.670.0108
JudgeLM-7B0.660.0026
Contains (lexical)0.640.0087
Llama3-8B0.60.0126
Llama2-7B0.470.0112
Exact match (lexical)0.470.29
Gemma-2B0.260.007
Figure 2 accessible table
Judge (TriviaQA)Mean Scott's piStd dev
Human annotators vs majority vote0.9620.0107
Llama3-70B0.880.0046
Llama3.1-70B0.880.0039
Llama3.1-8B0.780.005
Llama2-13B0.750.0043
Llama2-70B0.690.0114
Mistral-7B0.670.0108
JudgeLM-7B0.660.0026
Contains (lexical)0.640.0087
Llama3-8B0.60.0126
Llama2-7B0.470.0112
Exact match (lexical)0.470.29
Gemma-2B0.260.007

Order, length and authorship move the verdict

Position bias is the most thoroughly measured of these effects. In the MT-bench study, judges compared two similar GPT-3.5 answers to each first-turn question and then the same pair in swapped order. GPT-4 returned the same verdict both ways in 65.0 percent of cases; Claude-v1 did so in 23.8 percent and favored whichever answer came first in 75.0 percent. The answers were deliberately close, which makes the test hard, but close calls are exactly what a release threshold turns on. [1]

Shi and colleagues scaled the swap test to 15 judges, with 4,800 pairwise instances per judge on MTBench and 2,240 on DevBench, at temperature 1. Position consistency on MTBench ran from 0.57 to 0.82, and one judge could move sharply between tasks: Claude-3-Haiku fell from 0.57 to 0.23 on DevBench while Gemini-1.5-pro rose from 0.62 to 0.84. The authors found the bias was not random variation, differed by judge and task, and depended strongly on the quality gap between candidates and only weakly on prompt length. [3]

Length and authorship are the other two levers. Zheng's repetitive-list attack took 23 answers containing a numbered list and prepended a rephrased copy of the list, adding no information; Claude-v1 and GPT-3.5 preferred the padded version in 91.3 percent of cases and GPT-4 in 8.7 percent. On self-preference the MT-bench data were inconclusive by the authors' own reading, although GPT-4 gave itself a 10 percent higher win rate and Claude-v1 a 25 percent higher one. Panickssery, Bowman and Feng later showed that GPT-4 and Llama 2 can pick out their own outputs with non-trivial accuracy out of the box and, after fine-tuning, that self-recognition and self-preference rise together linearly. [1][5]

JUDGE-BENCH states the general result. Across 20 human-annotated datasets and 11 models, judge quality varied with the property judged, the expertise of the human annotators and whether the text was written by people or models, and the authors conclude that a model should be carefully validated against human judgments before it serves as an evaluator. [4] None of these numbers transfers to your agent. They tell you which tests to run, not what result you will get.

Figure 03

Position consistency changes with the judge and the task

Swapping answer order left verdicts unchanged in 57 to 82 percent of MTBench comparisons, and the same judge could rank very differently on DevBench. [3]

Grouped horizontal bar chart of pairwise position consistency for 14 judges from Shi et al. Table 2, MTBench then DevBench: Claude-3.5-Sonnet 0.82 and 0.76, GPT-4 0.82 and 0.83, Llama-3.3-70B 0.80 and 0.89, Llama-3.1-405B 0.77 and 0.79, GPT-4o 0.76 and 0.80, o1-mini 0.76 and 0.84, GPT-4-Turbo 0.75 and 0.79, Claude-3-Opus 0.70 and 0.69, GPT-3.5-Turbo 0.70 and 0.76, Llama-3.1-8B 0.69 and 0.47, Gemini-1.5-pro 0.62 and 0.84, Claude-3-Sonnet 0.59 and 0.71, Claude-3-Haiku 0.57 and 0.23, Gemini-1.0-pro 0.57 and 0.66.

Source. Shi et al., Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, arXiv v9 (November 11, 2025), Table 2, pairwise columns; one study, values copied. [3]

Method. Mean position consistency per judge across candidates and tasks: 4,800 instances per judge on MTBench and 2,240 on DevBench, temperature 1. Gemini-1.5-flash is omitted because the paper marks its DevBench result invalid (0.96 error rate). Llama-3.1-8B had a 0.25 error rate on MTBench. Standard deviations, preference fairness and list-wise results are in the paper and not plotted.

Accessible table and figure data
Figure 3 accessible table
Judge, sorted by MTBenchMTBench position consistencyDevBench position consistency
Claude-3.5-Sonnet0.820.76
GPT-40.820.83
Llama-3.3-70B0.80.89
Llama-3.1-405B0.770.79
GPT-4o0.760.8
o1-mini0.760.84
GPT-4-Turbo0.750.79
Claude-3-Opus0.70.69
GPT-3.5-Turbo0.70.76
Llama-3.1-8B0.690.47
Gemini-1.5-pro0.620.84
Claude-3-Sonnet0.590.71
Claude-3-Haiku0.570.23
Gemini-1.0-pro0.570.66
Figure 3 accessible table
Judge, sorted by MTBenchMTBench position consistencyDevBench position consistency
Claude-3.5-Sonnet0.820.76
GPT-40.820.83
Llama-3.3-70B0.80.89
Llama-3.1-405B0.770.79
GPT-4o0.760.8
o1-mini0.760.84
GPT-4-Turbo0.750.79
Claude-3-Opus0.70.69
GPT-3.5-Turbo0.70.76
Llama-3.1-8B0.690.47
Gemini-1.5-pro0.620.84
Claude-3-Sonnet0.590.71
Claude-3-Haiku0.570.23
Gemini-1.0-pro0.570.66

What each managed evaluator lets you see and pin

AgentCore Evaluations offers three kinds of LLM judge. Built-in evaluators such as Builtin.Helpfulness use predefined evaluator models and prompt templates that cannot be modified. AWS publishes the prompt text of all 17 built-in templates, 2 at session level, 11 at trace level and 4 at tool level, but the pages reviewed do not name the model. Built-in evaluators use cross-Region inference within the geography, and requests from ap-south-2, ap-southeast-5, ap-northeast-2 and ap-southeast-7 use global cross-Region inference across all commercial Regions. AWS says prompts and outputs may be processed outside the source Region, and for an evaluator the prompt carries the agent's trace. Managed third-party evaluators from DeepEval and AutoEval run on the same model as the built-ins, with no model field and no version selection, and AWS makes no quality claims for them. [6][7][8][9]

Custom evaluators hand you the judge. A custom LLM-as-a-judge evaluator names a Bedrock model in modelConfig, through bedrockEvaluatorModelConfig with a modelId and optional inferenceConfig, or responsesEvaluatorModelConfig for the bedrock-mantle endpoint, together with your instructions and rating scale; the service appends a standardization prompt that makes the judge give its reason before its score. A derived evaluator runs a built-in or third-party prompt on a model you choose, in your account, and AWS notes that it does not validate that model against the metric. Custom evaluators are AWS's documented route around cross-Region inference. Bedrock's separate model evaluation jobs also let you pick the evaluator model, from a published list. [8][9][10][12]

Foundry's AI-assisted evaluators take a deployment_name and run on that deployment, which can be an Azure OpenAI or OpenAI reasoning or non-reasoning model. The composite evaluators builtin.output_quality and builtin.tool_use_quality, both in preview, batch five or six criteria into one judge call; Microsoft recommends the GPT-5.6 family for them, names gpt-5.6-luna as the lowest-cost option, and returns model and token metadata with each result. Task Completion, Task Adherence and Intent Resolution are among the agent evaluators still marked preview, and Quality Grader is deprecated. Risk and safety evaluators work differently: they call Microsoft-hosted safety models and take no deployment_name; the content safety evaluators among them score on a 0 to 7 severity scale with a default threshold of 3, and the rest return pass or fail. [14][15]

Google is the most explicit about the judge. Each managed rubric metric is versioned. The latest versions of most, such as general_quality_v2 and safety_v3, use Gemini 3.5 Flash as the judge, and some agent metrics also call Gemini 3.1 Pro to generate rubrics; the previous versions use Gemini 2.5 Flash and Pro. Metrics run the latest version unless you pin one, as in types.RubricMetric.GENERAL_QUALITY(version='v1'). Versions differ in more than the model: safety_v3 makes 31 judge calls per item where safety_v1 made 10, and the two score in opposite directions, with 1 meaning a policy violation in safety_v3 and a safe response in safety_v1. Response flipping with flip_enabled, a sampling_count from 1 to 32 with a default of 4, and tuned judges are documented with EvalTask samples that import from vertexai.preview.evaluation. Google keeps the EvalTask interface generally available but no longer develops it actively, and the recommended GenAI Client is in preview. [17][18][19]

Figure 04

Who controls the judge in each managed evaluator

You can name and hold the judge model only in AgentCore custom evaluators, Bedrock evaluation jobs, Foundry AI-assisted evaluators and Google's versioned or configured judges. [6][9][12][14][15][18][19]

Matrix of eight evaluator types by judge model, pinning and prompt. AgentCore built-in: model not named, cross-Region within geography; cannot be modified; prompts published. AgentCore managed third-party: same model as built-in; no model or version selection; open-source library metric. AgentCore custom or derived: your Bedrock model ID; pinned by ID and locked while an enabled configuration uses it; your prompt or the base evaluator's. Bedrock evaluation jobs: chosen from a published list; pinned by ID; built-in or custom metric. Foundry AI-assisted: your deployment; pinned through version and upgrade option; prompt not reproduced. Foundry risk and safety: Microsoft-hosted models; not pinnable; prompt not reproduced, content harm severity 0 to 7, other risks pass or fail. Google managed rubric metrics: named per version, Gemini 3.5 Flash latest; pin the metric version; adaptive per-prompt or static rubrics. Google EvalTask custom metrics: configurable Gemini or tuned judge; pinned by model; your template with flip and sampling options.

Source. Conceptual comparison drawn from AWS, Microsoft and Google documentation reviewed October 10, 2026. [6][7][8][9][10][11][12][14][15][16][17][18][19]

Method. Conceptual summary of documented controls, not a test of the services. Preview status: Foundry composite and several agent evaluators and Google's GenAI Client are in preview; Google's EvalTask interface is generally available but no longer under active development.

Accessible table and figure data
Figure 4 accessible table
EvaluatorJudge modelPinningPrompt
AgentCore built-inNot named; cross-Region within geographyNo; cannot be modifiedPublished, fixed
AgentCore managed third-partySame model as built-inNo model or version selectionOpen-source library metric, fixed
AgentCore custom or derivedYour Bedrock model IDYes; locked while an enabled configuration uses itYours, or the base evaluator's
Bedrock evaluation jobsChosen from a published listYes, by model IDBuilt-in or custom metric
Foundry AI-assistedYour deploymentThrough version and upgrade optionNot reproduced in pages reviewed
Foundry risk and safetyMicrosoft-hosted safety modelsNoNot reproduced; content harms 0 to 7
Google managed rubric metricsNamed per version; Gemini 3.5 Flash latestYes, metric versionAdaptive per-prompt or static rubrics
Google EvalTask custom metricsConfigured Gemini or tuned judgeYes, by modelYours; flip and sampling options
Figure 4 accessible table
EvaluatorJudge modelPinningPrompt
AgentCore built-inNot named; cross-Region within geographyNo; cannot be modifiedPublished, fixed
AgentCore managed third-partySame model as built-inNo model or version selectionOpen-source library metric, fixed
AgentCore custom or derivedYour Bedrock model IDYes; locked while an enabled configuration uses itYours, or the base evaluator's
Bedrock evaluation jobsChosen from a published listYes, by model IDBuilt-in or custom metric
Foundry AI-assistedYour deploymentThrough version and upgrade optionNot reproduced in pages reviewed
Foundry risk and safetyMicrosoft-hosted safety modelsNoNot reproduced; content harms 0 to 7
Google managed rubric metricsNamed per version; Gemini 3.5 Flash latestYes, metric versionAdaptive per-prompt or static rubrics
Google EvalTask custom metricsConfigured Gemini or tuned judgeYes, by modelYours; flip and sampling options

Build a calibration set from your own traces

The calibration set is a sample of your agent's real sessions, labeled by people, against which every judge configuration is scored. Draw it from the trace source the evaluator reads, at the level it scores, which is session, trace or tool call in AgentCore and turn or conversation in Foundry. Stratify it so every tool, every major intent and both outcomes appear, and over-sample close calls on purpose: agreement falls as the quality gap narrows, and borderline items decide whether a release passes. [1][3][10][14]

Write the reviewers' rubric in the evaluator's own labels. If the judge returns PASS, FAIL or INCONCLUSIVE, or a 1 to 5 score with a pass threshold of 3, the human scale must map onto it, or agreement will measure translation errors. Google's judge-evaluation workflow expects this pairing, with a {metric_name}/human_rating column for pointwise metrics or {metric_name}/human_pairwise_choice for pairwise ones beside the judge's result. Use two reviewers per item, keep the judge's verdict and reasoning out of their view, and settle disagreements with a third. Compute the reviewers' agreement with each other first. Thakur's team picked a task on which annotators reached a Scott's pi of 96.2 so that judge error would not be confused with ambiguity; if your reviewers agree far less, fix the rubric before measuring any judge. [2][10][14][20]

Seed the set with items whose correct label you already know:

  • Padded variants: correct answers with a rephrased copy of a list or paragraph prepended, which should never outscore the original. [1]
  • Non-answers such as Yes, Sure or a restated question, which some judges in Thakur's study marked correct. [2]
  • Both orders of every pairwise comparison. [1][3]
  • A tool result containing text addressed to the grader, for example an instruction to rate the session fully successful. AgentCore's built-in goal success prompt includes tool outputs in the conversation record it gives the judge, so text an attacker can influence reaches the grader; that is an inference from the published template, not a documented attack. [7]
  • Failures already known from incidents or red-team runs, labeled fail.

Score the judge and probe its biases

Score each evaluator on its own, because a composite pass rate can hide the one component the judge gets wrong. Report percent agreement with the adjudicated label, a chance-corrected statistic, the reviewers' agreement with each other as the ceiling, and the full confusion matrix. For a release gate the cell that matters most is the false pass, an item the judge passed and people failed. Google's evaluate_autorater returns balanced accuracy, balanced F1 and a confusion matrix; on a two-class metric a judge that guesses at random scores about 0.5 balanced accuracy, so read results against that floor rather than against zero. [2][20]

Then probe the exact configuration that will gate. Count a pairwise win only when it survives both orders, the conservative rule Zheng's team used; Google's flip_enabled flips half of the calls, which reduces order bias in aggregate but does not show which verdicts depended on order. Compare padded and original answers. Repeat a subset to measure stability, as Shi's team did before trusting any position result, and as Google's sampling_count acknowledges. Where the agent and the judge are the same model or come from the same family, measure the judge's gap from the reviewers on the agent's own outputs, or pick a judge from another family. [1][3][5][19]

Decide in advance what the judge's own failures mean. A call that times out, returns malformed output or INCONCLUSIVE, or a Foundry composite result of not_applicable, has passed nothing, and a gate should count it as missing evidence. [10][14] The fragment below computes the core numbers from a labeled export.

Example script, Python 3 standard library only: reviewer agreement, judge agreement, Scott's pi, Cohen's kappa and false passes from a calibration CSV. Column names are placeholders; adapt the labels to the evaluator's scale.
# Example: score one judge against human reviewers on a labeled calibration set.
# CSV columns (placeholder names): trace_id, reviewer_a, reviewer_b, adjudicated, judge
# Labels are "pass" or "fail". Python 3 standard library only.
import csv
import sys
from collections import Counter


def observed(a, b):
    return sum(x == y for x, y in zip(a, b)) / len(a)


def scotts_pi(a, b):
    pooled = Counter(a) + Counter(b)
    expected = sum((n / (2 * len(a))) ** 2 for n in pooled.values())
    return float("nan") if expected == 1 else (observed(a, b) - expected) / (1 - expected)


def cohens_kappa(a, b):
    ca, cb = Counter(a), Counter(b)
    expected = sum(ca[k] * cb[k] for k in ca) / len(a) ** 2
    return float("nan") if expected == 1 else (observed(a, b) - expected) / (1 - expected)


with open(sys.argv[1], newline="") as handle:
    rows = list(csv.DictReader(handle))

a = [r["reviewer_a"] for r in rows]
b = [r["reviewer_b"] for r in rows]
truth = [r["adjudicated"] for r in rows]
judge = [r["judge"] for r in rows]

print(f"items: {len(rows)}")
print(f"reviewers: agreement {observed(a, b):.3f}, Scott's pi {scotts_pi(a, b):.3f}")
print(f"judge: agreement {observed(truth, judge):.3f}, "
      f"Scott's pi {scotts_pi(truth, judge):.3f}, Cohen's kappa {cohens_kappa(truth, judge):.3f}")

on_failures = [j for t, j in zip(truth, judge) if t == "fail"]
passed = on_failures.count("pass")
print(f"false passes: {passed} of {len(on_failures)} human-failed items")
Figure 05

Validate a judge before its score gates a release

Fix the judge's identity first, label without the judge in view, and recalibrate on every change to either side. [1][2][3][20]

Flowchart of seven steps: record the judge identity; sample your own traces with close calls and seeded failures; label with two reviewers who do not see the judge and adjudicate; score agreement, chance-corrected agreement and false passes; probe order, padding, empty answers and grader instructions; decide whether the score gates, stays advisory or the judge is replaced; recalibrate on any judge, prompt, agent or traffic change.

Source. Conceptual procedure based on the methods of Zheng et al., Thakur et al. and Shi et al. and Google's judge evaluation workflow. [1][2][3][20]

Method. Conceptual ordering of steps; no numeric data. Thresholds are a local decision and are not taken from the sources.

Accessible table and figure data
Figure 5 accessible table
StepWhat to do
Record the judgeEvaluator ID, judge model ID, prompt or metric version, settings, Region
Sample your tracesEvery tool and intent, both outcomes, extra close calls, seeded failures
Label without the judgeTwo reviewers, written rubric in the judge's labels, a third to adjudicate
Score the judgePercent agreement, Scott's pi or kappa, confusion matrix, false passes
Probe for biasSwapped order, padded answers, empty answers, text aimed at the grader
DecideGate, advisory only, or replace the judge; store the result with the identity
RecalibrateOn any change to judge, prompt or metric version, agent model or traffic
Figure 5 accessible table
StepWhat to do
Record the judgeEvaluator ID, judge model ID, prompt or metric version, settings, Region
Sample your tracesEvery tool and intent, both outcomes, extra close calls, seeded failures
Label without the judgeTwo reviewers, written rubric in the judge's labels, a third to adjudicate
Score the judgePercent agreement, Scott's pi or kappa, confusion matrix, false passes
Probe for biasSwapped order, padded answers, empty answers, text aimed at the grader
DecideGate, advisory only, or replace the judge; store the result with the identity
RecalibrateOn any change to judge, prompt or metric version, agent model or traffic

Pin the judge and record its identity

A calibration result is valid for one judge configuration. Beside every score that gates, record the evaluator ID or ARN; the judge model ID exactly as invoked, including any inference profile prefix such as us. or global.; the prompt template or metric version; temperature, sampling count and flipping; the managed library version where one exists; the Region or endpoint; and the date and result of the calibration the score relies on.

Pinning works differently on each platform. On AgentCore a built-in or managed third-party evaluator cannot be pinned, because AWS selects the model and, for third-party metrics, manages the library version; a gate that needs a fixed judge should use a custom or derived evaluator that names a model ID. Once an enabled evaluation configuration uses a custom evaluator, AgentCore refuses updates to it, so a changed judge becomes a new evaluator with its own calibration. On Foundry, set the judge deployment's versionUpgradeOption explicitly: an absent value is null, which behaves as OnceCurrentVersionExpired and upgrades the deployment to the current default at retirement, while NoAutoUpgrade makes it stop working instead. For a judge, stopping is the safer failure, since it cannot be mistaken for a pass. On Google, pass an explicit metric version. [6][9][11][16][18]

A pin holds only until the judge model retires, and vendor examples already show the collision. The table compares judge models named in samples and lists with the lifecycle tables. Google's tuned-judge sample tunes a model retired on June 1, 2026; Bedrock's evaluator list still offers Claude 3 Haiku, already past end of life, and Claude Sonnet 4, four days from it on the review date. [10][12][13][18][19][21]

Two lifecycle details catch judges in particular. A Bedrock Legacy model may be withdrawn from an existing customer after 15 days of inactivity, and a judge used only for monthly release runs can sit idle that long. A judge upgrade with no change to the agent moves the pass rate by itself, so treat a judge's retirement as a change to the measurement: calibrate the successor on the same set and keep both results. [13]

Judge models named in vendor samples and evaluator lists, checked against the providers' lifecycle tables on October 10, 2026. [10][12][13][18][19][21]
Where the judge appearsJudge modelLifecycle on October 10, 2026
Google tuned-judge samplegemini-2.0-flashRetired June 1, 2026
Google managed metrics, previous versionsGemini 2.5 Flash and 2.5 ProRetire October 20, 2026
Google managed metrics, latest versionsGemini 3.5 FlashRetires May 19, 2027 or later
Bedrock evaluator model listClaude 3 HaikuEnd of life September 10, 2026 in listed Regions
Bedrock evaluator model listClaude Sonnet 4End of life October 14, 2026 in 23 Regions
AgentCore custom evaluator samplesClaude Sonnet 4.5, global profileLegacy since October 8, 2026; end of life April 8, 2027

Keep measuring after the gate opens

Online evaluation turns the judge into a production monitor, and a monitor needs its own checks. Send a fixed share of judged production traces back to reviewers on a schedule and compare their labels with the judge's the same way as at calibration. A pass rate that moves while the agent is unchanged points at the judge or the traffic. The pages reviewed offer no way to see which model a built-in evaluator runs, so this sampling is how you would notice a vendor change. [6][9]

Recalibrate when the judge model or its version changes, when the prompt or metric version changes, when the agent's model changes, since a more verbose model meets the judge's length preference, and when new tools or a shifted traffic mix put the agent into territory the calibration set never covered. [1][18]

When a judge score may gate a release

The evidence supports a three-part rule. A judge score may block or allow a release only when the judge configuration that produced it was calibrated on your own labeled traces since the last change to the judge, its prompt or the agent, with chance-corrected agreement close to your reviewers' agreement with each other and a false-pass rate someone has accepted in writing; when its identity is pinned and recorded with the score; and when the property it judges cannot be checked deterministically. A score that fails any part is advisory: show it, trend it, and never let it open a gate alone.

The third part removes more than teams expect. Whether a tool call carried valid arguments, whether an agent touched a resource outside its scope, and whether output matched its schema are questions for code at the consuming boundary, not for a model's opinion of a transcript. Hosted judges that cannot be pinned, such as AgentCore's built-in evaluators and Foundry's safety evaluators, are useful for broad screening across traffic, but they need the same calibration before they gate, and only ongoing sampling will catch their drift. [6][15]

The rule stops applying in two cases. If your reviewers cannot agree with each other, there is nothing to calibrate against, and the fix is a sharper rubric rather than a stronger judge. Where an error is rare and severe, a few hundred labels cannot bound the false-pass rate tightly enough; route those cases to human review or a deterministic control.

Method and provenance

Source-led research summary of five peer-reviewed or preprint studies of LLM judges and of AWS, Microsoft and Google evaluator documentation, with an original validation procedure and calculated sample-size examples. Sources were reviewed on October 10, 2026.

No evaluator, judge model or cloud account was run or measured. Research figures describe the models and tasks each study tested, mostly from 2023 to 2025, and are not current error rates. Platform behavior, preview status and lifecycle dates are limited to the cited pages as of the review date and change often.

AI assistance. AI assisted research synthesis, drafting, chart and diagram planning and visual production, with deterministic editorial checks. No personal experience, independent human review or live test is claimed.

Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.

References

  1. Built-in evaluators, Amazon Bedrock AgentCore Developer Guide Amazon Web Services. Accessed .
  2. Third-party evaluators, Amazon Bedrock AgentCore Developer Guide Amazon Web Services. Accessed .
  3. Create evaluator, Amazon Bedrock AgentCore Developer Guide Amazon Web Services. Accessed .
  4. Update evaluator, Amazon Bedrock AgentCore Developer Guide Amazon Web Services. Accessed .
  5. Model lifecycle (Legacy), Amazon Bedrock User Guide Amazon Web Services. Accessed .
  6. Agent evaluators (Microsoft Foundry) Microsoft. Published . Accessed .
  7. Risk and safety evaluators (Microsoft Foundry) Microsoft. Published . Accessed .
  8. Working with models (Azure OpenAI in Microsoft Foundry Models) Microsoft. Published . Accessed .
  9. Configure a judge model, Gemini Enterprise Agent Platform Google Cloud. Accessed .
  10. Evaluate a judge model, Gemini Enterprise Agent Platform Google Cloud. Accessed .