# LLM Judges vs. Human Raters: Kappa Bands and Self-Preference

Dr. Samuel Ortiz · August 23, 2026

> Humans rate open-ended quality at just 0.6 kappa — below the 0.75 bar set for LLM judges. How calibration closes the gap and self-preference skews scores.

```html

| Takeaway | Detail |
| --- | --- |
| Human panels no longer clear the consistency bar machines are held to. | Open-ended quality ratings converge near 0.6 inter-rater kappa — under the 0.75 threshold governance councils demand of judges — and reliability guidance ties low agreement to guideline gaps and thin rater training, which a 5% budget carve-out for calibration attacks at the root. |
| The calibrated judge is often the more consistent rater. | Against the same gold set, a tuned judge clears 0.8 agreement at roughly a hundredth of the unit price of a human panel; within a fixed rating budget, a 9.75% reserve funds continuous gold-set refresh without adding headcount. |
| Dataset-level reliability metrics fit production annotation better than pairwise-only kappa. | Krippendorff's alpha handles missing ratings, ordinal scales, and uneven rater coverage where Cohen's and Fleiss' kappa strain; Label Studio Enterprise reports agreement via pairwise methodology across the Data Manager, data-quality dashboard, and Members dashboard — the surfaces where a pre-committed 5% drop in judge-versus-gold agreement fires the flip-back review. |
| The scarce asset is a fresh, adversarially hard human gold set — not labeling labor. | With 90%-plus of rating dollars still flowing to panels hovering near 0.6 kappa, inverting the split so 9.75% of spend maintains a rotating adversarial gold set preserves the audit trail while cutting per-judgment cost by orders of magnitude. |

90%-plus of 2026 rating budgets still flow to human panels whose inter-rater agreement on open-ended quality hovers near 0.6 kappa — while the LLM judge those panels are 'double-checking' clears 0.8 against the same gold set at roughly a hundredth of the unit price. On open-ended output, the old assumption that humans are ground truth has quietly inverted: the calibrated machine is now often the more consistent rater in the room.

The gap is measurable with standard reliability tooling. Krippendorff's alpha absorbs the messiness of real annotation — missing ratings, ordinal scales, uneven rater coverage — where Cohen's and Fleiss' kappa strain, and Label Studio Enterprise exposes task agreement, top confusion pairs, and member-and-model matrices computed with pairwise methodology. Reliability guidance treats persistently low agreement as a symptom of unclear guidelines or thin rater training: a fixable process defect, not proof that human judgment is worthless.

Every certified judgment in the cascade is the same small artifact: a system prompt carrying the scoring rubric — a 5-point anchored scale that spells out what a 2 looks like versus a 4 — plus the user query and one candidate response, sent to a pinned judge snapshot from the 2026 frontier generation, Claude Sonnet-class or GPT-5-class. A typical judgment consumes a modest number of tokens and returns a score with a short rationale; store that rationale with the snapshot ID attached, because it is the audit trail governance councils ask for. Pinning is not bureaucracy — according to GovTech's AI Practice blog (July 2025), small changes in the system prompt can dramatically change a judge's response, so an unpinned judge rewrites its own experiment mid-run.

![LLM Judges vs. Human Raters](https://static.mm-ais.com/article-images-ai/llm-judges-vs-human-raters-kappa-bands-a-ai-6a95e1d0.jpg)

## Anatomy of a Penny Judgment

Pairwise mode needs one structural fix: run both orders. Present A left of B, then B left of A, and count a win only when the verdict survives the swap — position-swap consistency. Judges favor whichever response appears first or last; GovTech's AI Practice blog documents the positional bias, and single-order judging lets it flip a material share of verdicts, enough to reorder a leaderboard decided on close calls. The swap doubles inference cost yet remains the cheapest structural bias fix available — no retraining, no extra model, no new pipeline stage.

Routing demands a confidence value on every item, and the score alone cannot supply one. Where the endpoint exposes token-level logprobs on the verdict token, read them directly — though several reasoning-style APIs withhold logprobs, so verify first. Otherwise sample k=5 verdicts at temperature 0.7 and take the majority: 5-for-5 unanimity signals high confidence, a 3–2 split signals trouble. Items falling below your confidence floor enter the human queue. That signal — not the score — is what makes the cascade architecturally possible; without per-item confidence, "escalate the hard ones" is a slogan, not a router.

Then comes the failure mode that survives calibration: self-preference. A judge shares distributional fingerprints with outputs from its own model family — phrasing habits, formatting tics — and systematically over-scores them, as Panickssery et al. showed in their 2024 paper "LLM Evaluators Recognize and Favor Their Own Generations." The standing mitigation is a cross-family rule: the judge may never share lineage with either candidate being scored. This is where the myth that an aggregate agreement score confers universal trustworthiness dies — the bias concentrates where same-family candidates compete, so the pooled average looks clean while the contested slice skews. Practically, a Claude-family candidate against a GPT-family candidate disqualifies both Sonnet-class and GPT-5-class judges; reach for a third family or an open judge.

Underneath everything sits the gold set. Stratify your production items across rubric lines and difficulty bands, then double-label a subset of items with two human raters to establish your own human-human kappa baseline, the ceiling your judge gets certified against. Freeze the item set and the label versions together: reword an anchor descriptor mid-year and every later certification measures a different instrument. Quarterly comparability holds only because nothing underneath moves.

You already run fragments of this stack. G-Eval's chain-of-thought form-filling is the default rubric-scoring pattern — criteria, reasoning, then score. Prometheus 2 covers on-prem and residency-constrained deployments a frontier API judge cannot reach, and doubles as the neutral third family the cross-family rule sometimes forces. Braintrust, LangSmith, and W&B Weave supply the operational layer — judge configuration, snapshot versioning, per-run agreement metrics — the exact fields an auditor requests.

Before your next pilot readout: confirm the snapshot ID is logged on every stored verdict, replay your last hundred pairwise verdicts through the swap, and grep the candidate roster for lineage collisions with the judge. Treat any collision as a finding, not a footnote.

Zheng et al.'s MT-Bench and Chatbot Arena paper (NeurIPS 2023) is still the anchor result: GPT-4 acting as judge agreed with human preferences on roughly 85% of MT-Bench comparisons — at or above the ~81% human-human agreement the authors measured on the same material. Read that precisely: the judge matched pooled humans, noise included. On general chat quality, "human ground truth" was already a consensus of fallible readers, and a frontier model cleared it.

| Component | Spec | Failure it prevents | Cost note |
| --- | --- | --- | --- |
| Pointwise call | 5-point anchored rubric; pinned 2026-frontier snapshot | Prompt drift, unreproducible scores | Modest token footprint per judgment |
| Swapped pairwise | Verdict counts only if it survives the A/B order swap | First/last positional bias | Doubles inference cost |
| Confidence extraction | Verdict-token logprobs, or k=5 samples at temperature 0.7 | Silent errors on ambiguous items | Roughly 5× the single-call envelope on sampled items |
| Cross-family rule | Judge shares no lineage with either candidate | Self-preference toward its own family | Selection-time only; no added inference |
| Frozen gold set | Fixed item pool; double-labeled subset; labels versioned | Undetected drift between certifications | Human labeling effort; certification runs quarterly |

![Anatomy of a Penny Judgment — LLM Judges vs. Human Raters](https://static.mm-ais.com/article-images-ai/llm-judges-vs-human-raters-kappa-bands-a-ai-a026e534.jpg)

## The Scoreboard

The same paper quantified the two biases every current judge configuration must actively cancel. Swapping the order of the two answers flipped roughly a fifth to a quarter of single-order GPT-4 verdicts — position bias large enough to reorder rankings on its own. GPT-4 also showed measurable self-enhancement, favoring GPT-4-family outputs beyond what human raters did. Neither bias is exotic; both are cancellable with position ensembling and cross-family judge assignment, provided you measure them instead of assuming them away.

The human baseline deserves equal scrutiny. According to Ouyang et al.'s InstructGPT paper (2022), agreement between trained commercial labelers and the task authors who wrote the comparison tasks ran near 72% — placing "human ground truth" on open-ended quality at roughly 0.6 kappa. Two consequences follow. First, pooled-human equivalence is a modest, reachable target, not a heroic one. Second, your frozen gold set inherits this noise, which is why certification must compare the judge against pooled raters, never a single annotator's verdict.

JudgeBench (He et al.) is the correction for anyone reading the 85% headline as a universal license. On hard reasoning, code, and science items drawn from competition-style sources, the strongest judges' accuracy falls into the low-to-mid 70s percent, versus roughly 90% on easy chat. Reliability degrades exactly where model-selection decisions get made: the contested stratum, not the easy bulk, decides which model ships. An aggregate score earned on an easy-skewed suite does not transfer — certify per stratum, or the collapse stays invisible until a release gate trips over it.

Style is the third silent variable. AlpacaEval 2.0 introduced length-controlled win rates after demonstrating that raw LLM-judge win rates tracked response length; LMSYS added style control to Chatbot Arena for the same reason. Both adjustments reordered leaderboards — direct proof that judges score presentation alongside substance. Operationally, freeze the style-control settings inside the certified judge configuration: redeploy without them and the judge drifts toward verbose outputs, quietly inflating whichever candidate model rambles.

Governance councils should treat this table as the quarterly review artifact: refresh each row against the frozen gold set, and the moment a stratum slips below the certification floor, the flip rule above fires — no committee debate required.

Landis & Koch drew their agreement bands for radiologists; fifty years on, they are the load-bearing wall of eval governance. The switch table below exists to make one thing mechanical that most teams still leave to habit: which rater touches which tier. The two habitual failure modes are an all-human panel (too slow and too expensive for CI-scale volume) and an all-judge pipeline (uncertified, unaudited). By 2026 the unit economics are settled — the cost gap was quantified earlier in this guide — so the remaining design question is assignment, and assignment should be written down, not felt.

| Evidence | Figure | Licenses | Forbids |
| --- | --- | --- | --- |
| Zheng et al., MT-Bench (GPT-4 judge) | ~85% agreement vs ~81% human-human | Judge-first routing on routine chat evals | Treating chat parity as universal coverage |
| Zheng et al., position-swap test | A fifth to a quarter of single-order verdicts flip | Mandatory position ensembling | Single-pass verdicts in production |
| Zheng et al., self-enhancement probe | GPT-4 favored beyond human rate | Cross-family judge assignment | A vendor grading its own family |
| Ouyang et al., InstructGPT labeler study | ~72% labeler-author agreement (~0.6 kappa) | Pooled-rater gold sets | Any single annotator as truth |
| JudgeBench (He et al.), hard strata | Low-to-mid 70s% vs ~90% on easy chat | Per-stratum certification | Aggregate-only scorecards |
| AlpacaEval 2.0 / Arena style control | Both adjustments reordered leaderboards | Frozen style controls per certification | Raw win-rate comparisons |
| Rate cards: Scale AI/Outlier, Surge AI vs frontier APIs | Professional human rates vs low per-call API fees | Judge-first economics at volume | All-human panels on routine items |

The stakes taxonomy drives everything. **Tier 1** is routine regression evaluation: daily or weekly CI runs over thousands of items, where every decision is fully reversible because a bad call costs a rerun. **Tier 2** is the release gate: monthly, choosing among two or three finalist checkpoints, reversible until you ship and expensive afterward. **Tier 3** is safety, compliance, and externally reported claims — effectively irreversible once published. Each tier gets a different rater, and the table's job is to make that assignment mechanical rather than habitual.

![The Scoreboard — LLM Judges vs. Human Raters](https://static.mm-ais.com/article-images-pixabay/llm-judges-vs-human-raters-kappa-bands-a-b81af81d.jpg)

## The Switch Table

The kappa gate is explicit and anchored to Landis & Koch's conventional bands. Quarterly certification against the frozen gold set governs auto-adjudication: at kappa ≥ 0.75 ("substantial"), the judge may auto-adjudicate Tier 1 alone; at 0.60–0.75, a forced human confirmation sample kicks in; below 0.60, the tier flips back to all-human until recalibration passes. One edge case worth pre-empting: when kappa slips, do not stack additional judge layers on top. According to GovTech's AI practice assessment, each additional validation step with an imperfect validator compounds errors rather than reducing uncertainty — the remedy is human confirmation sampling, not more validators.

Platform leads should also compute the breakeven inequality instead of arguing unit price. Judge-first wins if and only if Cjudge + s·Chuman + perror·Eerror < Chuman, where s is the forced human share and perror·Eerror the expected downstream cost of judge mistakes. Solving for the critical share gives s* = 1 − (Cjudge + perror·Eerror)/Chuman. Because judge fees sit orders of magnitude below human adjudication rates, s* approaches 1 minus your normalized error cost — the margin is enormous at 2026 rates, which is why quality gates, not unit price, are the binding constraint. Plug in your own invoice and incident log; stop debating pennies.

Throughput settles Tier 1 on capability alone. A parallelized judge pool processes judgments orders of magnitude faster than a trained human panel, whose per-rater daily output is limited by attention and stamina. Same-day regression loops over multi-thousand-item suites are physically impossible with humans alone — and per GovTech's assessment, judges also apply criteria consistently, unlike humans whose energy and interpretation drift across days.

The winners, defended per row: Tier 1 goes to the certified LLM judge on volume, reversibility, and certified kappa. Tier 2 goes to a hybrid cascade — the judge ranks all candidates, and a 3-person human panel adjudicates the top and bottom decile before sign-off, because the tails are where checkpoint selection actually happens. Tier 3 stays all-human with dual independent rating and reconciliation, for two reasons: judges inherit the blind spots of the model family being scored, and a headline kappa earned mostly on easy, prevalence-skewed items licenses nothing about the contested slice — demand per-stratum, PABAK-adjusted readouts before anyone claims human-equivalence everywhere.

Dollar figures in the last column are deliberately ordinal: indicative cost varies with your vendor rate card, rubric length, and output-token settings, so pull your own last invoice rather than trusting any benchmark. Lift the table verbatim into your governance charter, name an owner for the quarterly certification readout, and compute s* from your own numbers before the next release gate convenes.

Panickssery, Bowman, and Feng demonstrated that LLM evaluators recognize and favor their own generations — a judge systematically grades outputs from its own model family more generously than equal-quality outputs from rivals. Hold that finding against every certification number in this guide, because it exposes what pooled agreement cannot: a kappa is a property of the suite it was earned on, not of the judge.

| Eval tier | Typical volume/mo | Decision reversibility | Required kappa vs gold | Assigned rater | Indicative cost |
| --- | --- | --- | --- | --- | --- |
| Tier 1 — routine regression (CI) | Tens of thousands (thousands per run × daily/weekly cadence) | Fully reversible — rerun the suite | κ ≥ 0.75 quarterly; 0.60–0.75 forces a human confirmation sample; < 0.60 returns tier to all-human | Certified LLM judge + fixed random human audit | Lowest — judge fees plus audit share |
| Tier 2 — release gate | Hundreds to low thousands (all candidates across 2–3 finalists) | Reversible until ship; costly after | κ ≥ 0.75 for ranking; deciles adjudicated by humans regardless | Hybrid cascade — judge ranks, 3-person panel adjudicates top/bottom decile | Middle — judge fees plus decile adjudication labor |
| Tier 3 — safety, compliance, external claims | Typically dozens to hundreds (smallest tier) | Effectively none once published | Not delegated — dual rating; report per-stratum agreement | All-human panel, dual independent rating + reconciliation | Highest — panel labor dominates |

The anchor evidence for judge-based evaluation comes overwhelmingly from public, English-language, general-chat preference data — mostly easy items, skewed toward clear-cut wins. Production tiers are none of those things. So when a judge certifies at or above the floor, retire the comfortable myth that it is thereby human-equivalent everywhere. The kappa was earned largely on the easy majority; on the contested slice — the low-teens share of items that actually separates two candidate models — judge accuracy sags hardest, occasionally by double-digit margins. And because Cohen's kappa is sensitive to prevalence, it can hide that collapse entirely unless you also report per-stratum agreement and a prevalence-adjusted statistic such as PABAK. What Kappa Hides covers the statistics; treat them here as a certification requirement, not a footnote.

![The Switch Table — LLM Judges vs. Human Raters](https://static.mm-ais.com/article-images-pixabay/llm-judges-vs-human-raters-kappa-bands-a-a327a823.jpg)

## What the Data Doesn't Tell You

Fidelity also refuses to sit still across cases. The same judge, rubric, and rater pool behave differently under six recurring conditions:

Four conditions bend the routing rule without overturning it. First, near-ties: when two frontier candidates finish inside the noise band, most items become contested, the confidence floor escalates a large share of traffic, and the cascade's economics invert — this, plus release gates and safety claims, is exactly where the all-human premium is justified. Second, mid-quarter shocks: certification is a snapshot, so a silent provider-side judge update or a traffic-mix shift invalidates it overnight; track the weekly judge-human disagreement rate on the audit stream as a leading indicator, and recertify immediately after any judge change. Third, adversarial and contaminated items: prompts carrying injection attempts, or resembling the judge's training corpus, manufacture meaningless agreement — screen the gold set and live items for both. Fourth, thin tiers: on low-volume tiers the fixed random audit yields too few human-labeled items per quarter to power the certification test; pool adjacent quarters, consolidate the tier, or keep it human-rated.

None of this argues the thesis backwards — it draws the boundary. Judge first, escalate floor-misses plus the fixed audit, and flip any tier back to all-human the moment certification slips below the floor. The habit that keeps the architecture honest is unglamorous: report per-stratum, distrust pooled averages, and treat every kappa as a dated measurement of one distribution.

| Condition | What happens | Verify before trusting |
| --- | --- | --- |
| Contested tail | Pooled kappa rides on easy items; the deciding stratum degrades most | Per-stratum kappa on items that changed the ranking |
| Same-family judging | Self-preference inflates scores for the judge's own model family | Cross-family assignment or disclosed overlap flags |
| Pairwise ordering | Position bias flips verdicts on near-ties, per Wang et al.'s fairness audit | Order-swapped double passes before tie-breaks |
| Vague rubric anchors | Unanchored scales widen the judge-human gap | Anchored exemplars for every score point |
| Non-English locales | Agreement typically sags outside English-heavy training mixes | Per-locale certification, not a global stamp |
| Verbose outputs | Judges tend to reward length over substance | Length-neutral rubric clauses; spot-check the longest decile |
| Rater-pool drift | The human baseline itself moves as guidelines or staff change | Rater guidelines versioned alongside the frozen gold set |

Ninety-four percent raw agreement, kappa 0.37 — same two raters, same hundred-item set. Construct it yourself: 92 of those items are clear passes, two competent reviewers flag different slices of the eight genuine failures, yet they agree on 94 verdicts. Chance expectation works out to 0.905, so Cohen's kappa lands near 0.37. Vendors quote the first number and bury the second. Of the three statistics the Inter-Rater Reliability guide catalogs — percentage agreement, Cohen's kappa, ICC — percentage agreement is the flattered one, and it headlines most judge decks.

There the durable myth dies: kappa at or above 0.8 versus human raters does not mean human-equivalent everywhere. Such scores are typically earned on suites where most items are easy and one class dominates, and Cohen's statistic is acutely sensitive to prevalence — at a 92/8 split, mediocrity grades as excellence. On the contested slice — the small share of items that actually decide model selection — judge accuracy can fall by double-digit points while the headline kappa sits still. Byrt, Bishop, and Carlin's correction (PABAK = 2·po − 1) strips prevalence out so skewed suites compare cleanly — but it also discards marginal-bias information, so publish raw prevalence beside every kappa and treat cross-suite comparisons lacking those columns as noise. Label Studio's engineering blog adds a subtler trap: Krippendorff's alpha shares kappa's formula skeleton, differing only in how observed and expected agreement are computed per data type, so "kappa-like" outputs from different tools are not automatically commensurable.

![What the Data Doesn&#039;t Tell You — LLM Judges vs. Human Raters](https://static.mm-ais.com/article-images-pixabay/llm-judges-vs-human-raters-kappa-bands-a-f2374b15.jpg)

## What Kappa Hides

Agreement is not correctness. Kappa measures match with human labels, not truth. One pointed Hacker News critique — user YeGoblynQueenne, dissecting a widely circulated judge-validation paper — reduced its design to "ChatGPT better approximates the labeling of D by human annotators than human annotators," calling the circularity absurd. GovTech's AI Practice names the right question — "How do we know if our LLM judges are actually good judges?" — but human agreement is the entry gate, not the finish line. On code and data tasks with executable tests, judge and humans can agree with each other while both miss the same failure; validate a judge subset against programmatic ground truth before trusting kappa alone.

Drift is the concealment nobody dashboards. Providers update and deprecate judge snapshots on their own schedules, and because rubric criteria live in natural-language prompts — one judge repurposed across tasks by editing text, as Wikipedia's LLM-as-a-judge entry notes — a silent endpoint swap changes the effective scorer even when your rubric is byte-identical. A January certification says nothing about June behavior. Pin snapshot versions, re-certify against the gold-set floor on any provider changelog event, and treat unexplained kappa movement as an incident. Allegro's engineering blog tracks Cohen's kappa over time as "continuous alignment," but no public benchmark publishes quarter-over-quarter judge stability, and GovTech's MetaEvaluator post records the default failure mode: every team validating judges independently, duplicatively, inconsistently.

The adversarial channel is worse. Policies optimized against a fixed judge learn to hack it — length inflation, keyword stuffing, sycophantic framing — and Gao, Schulman, and Hilton's 2023 overoptimization curves show proxy scores climbing while true quality falls. Humans catch judge-hacked outputs on inspection; anything trained against the judge cannot. Tier 3 never delegates.

A single headline also buries line-level variance. Judges track humans closely on verifiable lines — factual accuracy, instruction-following — and diverge on tonal and safety-```

## Frequently Asked Questions

**What agreement level do governance councils demand from LLM judges, and how does that compare to what human raters actually achieve?**

Governance councils demand judges clear a 0.75 agreement threshold, yet open-ended quality ratings from human panels converge near 0.6 inter-rater kappa.

**How much cheaper is a tuned LLM judge than a human panel on the same gold set?**

Against the same gold set, a tuned judge clears 0.8 agreement at roughly a hundredth of the unit price of a human panel.

**Why must pairwise judgments be run in both orders, and what does that cost?**

Judges favor whichever response appears first or last, so a verdict counts only when it survives the A/B order swap — a fix that doubles inference cost but requires no retraining, extra model, or new pipeline stage.

**How do you get per-item confidence when the judge API doesn't expose logprobs?**

Where the endpoint withholds verdict-token logprobs, sample k=5 verdicts at temperature 0.7 and take the majority, routing items below your confidence floor into the human queue.

**What happens under the cross-family rule when one candidate is Claude-family and the other is GPT-family?**

That pairing disqualifies both Sonnet-class and GPT-5-class judges, forcing you to reach for a third family or an open judge such as Prometheus 2.

**How did GPT-4-as-judge compare to human-human agreement on MT-Bench?**

In Zheng et al.'s NeurIPS 2023 paper, GPT-4 acting as judge agreed with human preferences on roughly 85% of MT-Bench comparisons — at or above the ~81% human-human agreement the authors measured on the same material.

## Quick answers

| What inter-rater kappa do human panels reach on open-ended quality ratings? | Human panels converge near 0.6 inter-rater kappa, which falls under the 0.75 threshold governance councils demand of judges. |
| --- | --- |
| How does a tuned LLM judge's agreement and cost compare to a human panel against the same gold set? | A tuned judge clears 0.8 agreement at roughly a hundredth of the unit price of a human panel. |
| What is the self-preference failure mode that survives calibration? | A judge shares distributional fingerprints like phrasing habits and formatting tics with outputs from its own model family and systematically over-scores them, as Panickssery et al. showed in their 2024 paper 'LLM Evaluators Recognize and Favor Their Own Generations.' |
| What is the standing mitigation for self-preference bias? | A cross-family rule: the judge may never share lineage with either candidate being scored. |
| What did Zheng et al.'s MT-Bench and Chatbot Arena paper (NeurIPS 2023) find about GPT-4 as a judge? | GPT-4 acting as judge agreed with human preferences on roughly 85% of MT-Bench comparisons, at or above the ~81% human-human agreement the authors measured on the same material. |

Also worth reading: **Datumbox 0 on 10K Rows: Why Accuracy Is the Worst Criterion**: [Datumbox 0 on 10K Rows:](https://enterpriseailabs.io/blog/datumbox-0-on-10k-rows-why-accuracy-is-the-worst-criterion.php) · **2026 LLM Routing: Latency Is Architectural, Not a Serving Artifact**: [2026 LLM Routing: Latency Is](https://enterpriseailabs.io/blog/2026-llm-routing-latency-is-architectural-not-a-serving-artifact.php) · **Miqu Breakthrough Achieving Top Scores on Open LLM Leaderboard**: [Miqu Breakthrough Achieving Top Scores](https://enterpriseailabs.io/blog/miqu_breakthrough_achieving_top_scores_on_open_llm_leaderboa.php)

### Related reading

- [Fortify Your AI Defenses Against Prompt Injection Using Structured Queries and Preference Optimization](https://enterpriseailabs.io/blog/fortify-your-ai-defenses-against-prompt-injection-using-structured-queries-and-preference-optimization.php)
- [Excel to slides reporting: 19 of 68 pilots passed Deloitte 2026 benchmark](https://enterpriseailabs.io/blog/excel-to-slides-reporting-19-of-68-pilots-passed-deloitte-2026-benchmark.php)
- [Enterprise Pilot Safety Checks: 0.5% Escape Block or Launch 2026](https://enterpriseailabs.io/blog/enterprise-pilot-safety-checks-05-escape-block-or-launch-2026.php)
- [Résumé Review Rules: 2 August 2026—Deployed OpenAI o3 Application Falls Under Annex III](https://enterpriseailabs.io/blog/rsum-review-rules-2-august-2026deployed-openai-o3-application-falls-under-annex-iii.php)
- [John Deere harvests data insights with new AI technology](https://enterpriseailabs.io/blog/john-deere-harvests-data-insights-with-new-ai-technology.php)
- [Nvidia's earnings show why CIOs need to think beyond the GPU](https://enterpriseailabs.io/blog/nvidias-earnings-show-why-cios-need-to-think-beyond-the-gpu.php)

### Latest

- [Excel to slides reporting: 19 of 68 pilots passed Deloitte 2026 benchmark](https://enterpriseailabs.io/blog/excel-to-slides-reporting-19-of-68-pilots-passed-deloitte-2026-benchmark.php)
- [Enterprise Pilot Safety Checks: 0.5% Escape Block or Launch 2026](https://enterpriseailabs.io/blog/enterprise-pilot-safety-checks-05-escape-block-or-launch-2026.php)
- [Résumé Review Rules: 2 August 2026—Deployed OpenAI o3 Application Falls Under...](https://enterpriseailabs.io/blog/rsum-review-rules-2-august-2026deployed-openai-o3-application-falls-under-annex-iii.php)
- [John Deere harvests data insights with new AI technology](https://enterpriseailabs.io/blog/john-deere-harvests-data-insights-with-new-ai-technology.php)

Canonical: https://enterpriseailabs.io/blog/llm-judges-vs-human-raters-kappa-bands-and-self-preference.php
Markdown: https://enterpriseailabs.io/blog/llm-judges-vs-human-raters-kappa-bands-and-self-preference.php/index.md
