Rubric-Based Scoring with LLM Judges
Rubric scoring llm is the practice of grading model outputs against a written set of criteria, using another LLM as the judge instead of a human. Rather than asking "is this good?" and getting a single unreliable number back, you break quality into named dimensions (accuracy, tone, completeness, safety) with explicit pass/fail or point conditions for each one, then have the judge model apply that rubric consistently across thousands of outputs. This is the difference between an eval suite that tells you "score dropped 4%" and one that tells you exactly which dimension regressed and why.
Most teams start LLM evaluation with a single holistic prompt: "Rate this response 1-10 for quality." It works for a demo. It falls apart in production because a single number conflates unrelated failure modes. A response can be factually perfect and rude, or friendly and wrong, and a holistic score of 6/10 tells you nothing about which one happened. Rubric-based scoring fixes this by forcing the judge to evaluate criteria independently, which is also what makes the scores stable enough to alert on.
Why holistic scores fail and rubrics fix it
A holistic 1-10 score has two structural problems. First, it's underspecified: the judge model has to invent its own weighting between correctness, style, length, and safety on every call, and that weighting drifts between calls even at temperature 0, because the model has no fixed anchor for what "7" means versus "8". Second, it's not actionable: when the score drops from 8.2 to 7.6 across a release, you cannot tell your team what to fix.
A rubric replaces that single judgment with a checklist of independently gradable criteria. Each criterion gets its own instruction, its own pass/fail or scale, and ideally its own few-shot example. The judge model still uses judgment, but the judgment is scoped to one narrow question at a time, which is a task LLMs are demonstrably better at than open-ended holistic scoring.
Concretely, instead of:
Rate this customer support response from 1-10.you write:
Evaluate the response against each criterion below. For each, output PASS or FAIL and a one-sentence reason.
1. ACCURACY: Does the response contain only claims that are supported by
the provided context? No fabricated policy details, prices, or dates.
2. COMPLETENESS: Does the response address every question the customer
asked, not just the first one?
3. TONE: Is the response professional and empathetic, without being
apologetic to the point of undermining the company's position?
4. ACTIONABILITY: Does the response tell the customer a concrete next
step (link, timeframe, or action) rather than ending vaguely?Four independent judgments instead of one blended guess. You can now aggregate ACCURACY across 10,000 traces and know your factuality rate separately from your tone rate.
Designing a rubric that actually discriminates
A rubric is only useful if it separates good outputs from bad ones. A rubric where every response passes every criterion is measuring nothing. Three things make a rubric discriminate well.
Criteria must be independently checkable. If two criteria always move together (say, "accuracy" and "trustworthiness"), you're not getting two signals, you're getting one signal twice and paying for it twice. Write criteria that can diverge: a response can be accurate but incomplete, or complete but wrong.
Criteria must reference the source of truth, not vibes. "Is the response accurate" is a bad instruction because the judge model has no ground truth to check against unless you give it one. Pass the retrieved context, the reference answer, or the ticket history into the judge prompt explicitly, and instruct the judge to compare against that material, not against its own world knowledge.
Binary criteria beat 1-5 scales for most dimensions. LLM judges are reasonably good at PASS/FAIL and noticeably worse at consistently distinguishing a 3 from a 4 on a Likert scale. Reserve numeric scales for dimensions where granularity actually matters (e.g., "how many of the 5 required fields were extracted correctly") and keep everything else binary. This single change removes a large chunk of judge-to-judge variance.
Here's a rubric template that generalizes across most text-generation tasks:
RUBRIC: <task name>
For each criterion, respond with PASS, FAIL, or N/A (if not applicable
to this input), plus a short justification citing specific text from
the response.
1. FACTUAL_GROUNDING (PASS/FAIL)
PASS only if every factual claim in the response can be traced to
the provided context. Any unsupported claim, even a plausible one,
is a FAIL.
2. INSTRUCTION_ADHERENCE (PASS/FAIL)
PASS only if the response follows the explicit format/length/scope
constraints given in the system prompt.
3. HARMFUL_CONTENT (PASS/FAIL)
FAIL if the response contains content that violates the safety
policy attached below, regardless of whether the user asked for it.
4. COMPLETENESS (0-2)
0 = ignores the user's question
1 = addresses part of the question
2 = addresses all parts of the question
Output valid JSON matching this schema:
{"factual_grounding": "PASS|FAIL", "instruction_adherence": "PASS|FAIL",
"harmful_content": "PASS|FAIL", "completeness": 0-2,
"reasoning": {"factual_grounding": "...", ...}}Forcing structured JSON output, not free text, is what makes this pipeline automatable. You parse the judge's output the same way you'd parse any API response and feed it into your metrics store.
Picking and calibrating the judge model
Any capable current-generation model can act as a judge, but the judge should generally be at least as capable as the model being evaluated, and ideally from a different family or checkpoint than the model under test to avoid self-preference bias, the well-documented tendency of a model to rate its own outputs more favorably. If you're evaluating outputs from your production model, use a separate model as the judge, or at minimum a different checkpoint.
Calibration is the step teams skip and then regret. Before trusting a rubric at scale, run it against a small set (50-100 examples) that you or a domain expert have hand-labeled. Compare the judge's PASS/FAIL calls against the human labels and compute agreement:
import json
def judge_agreement(judged, human_labeled, criterion):
"""Percent agreement between LLM judge and human labels for one criterion."""
agree = 0
total = 0
for item_id, human_verdict in human_labeled.items():
judge_verdict = judged[item_id][criterion]
if judge_verdict in ("PASS", "FAIL"):
total += 1
if judge_verdict == human_verdict:
agree += 1
return agree / total if total else 0.0
for criterion in ("factual_grounding", "instruction_adherence", "harmful_content"):
rate = judge_agreement(judged_results, human_labels, criterion)
print(f"{criterion}: {rate:.1%} agreement with human raters")Anything below roughly 80% agreement means the rubric wording is ambiguous, not that the judge is broken. Fix it by adding a concrete few-shot example of a borderline PASS and a borderline FAIL for that specific criterion, then re-run calibration. This iteration loop, tighten the instruction, add an example, re-check agreement, is where most of the real work in rubric design happens, and it's worth doing before you trust the rubric on live traffic.
Run the judge at low or zero temperature and, if your provider supports it, ask for two independent judgments per item and flag disagreements for human review rather than averaging them silently. Averaging masks exactly the ambiguous cases you need to see.
Wiring rubric scoring into an eval pipeline
Once a rubric is calibrated, the pipeline is mechanical: generate response, call judge, parse JSON, log to your metrics store. Here's a minimal version using the Claude API, but the pattern is identical with any provider that supports structured output.
import json
from anthropic import Anthropic
client = Anthropic()
RUBRIC_PROMPT = """You are grading a customer support response.
CONTEXT (source of truth):
{context}
CUSTOMER QUESTION:
{question}
RESPONSE TO GRADE:
{response}
{rubric_criteria}
Output only valid JSON matching the schema above. No prose before or after."""
def score_response(context, question, response, rubric_criteria):
result = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
temperature=0,
messages=[{
"role": "user",
"content": RUBRIC_PROMPT.format(
context=context,
question=question,
response=response,
rubric_criteria=rubric_criteria,
),
}],
)
return json.loads(result.content[0].text)
scores = score_response(
context=ticket_context,
question=customer_question,
response=model_output,
rubric_criteria=RUBRIC_TEXT,
)
print(scores["factual_grounding"], scores["reasoning"]["factual_grounding"])Log every score alongside the input, the response, and the judge's reasoning string, not just the final PASS/FAIL. The reasoning field is what lets an engineer skim 20 failures and spot the actual pattern (say, the model keeps fabricating refund timelines) instead of just seeing a dropping percentage with no explanation.
Aggregate per criterion, not into one blended score, when you report results:
FACTUAL_GROUNDING: 94.2% pass (n=1,842)
INSTRUCTION_ADHERENCE: 88.1% pass (n=1,842)
HARMFUL_CONTENT: 99.8% pass (n=1,842)
COMPLETENESS: avg 1.7 / 2.0 (n=1,842)This is what makes rubric scoring valuable for regression testing: you can set a per-criterion threshold in CI (block a deploy if FACTUAL_GROUNDING drops below 90%, say) instead of a single fuzzy quality gate that catches everything and diagnoses nothing.
Common failure modes to watch for
Rubric criteria that overlap with prompt instructions the model already follows well. If your system prompt already forces a strict output format, don't spend a rubric criterion re-checking format compliance that's structurally guaranteed. Spend the criterion budget on things the model can actually get wrong.
Judge verbosity bias. Some judge models rate longer responses as more "complete" or "thorough" independent of actual content quality. Add an explicit instruction ("length is not a criterion; a short correct answer beats a long one with padding") and periodically check whether pass rates correlate with response length in your logs.
Position bias in comparative rubrics. If a criterion asks the judge to compare two responses (A/B testing two prompts), the judge exhibits measurable bias toward whichever response is presented first. Randomize the order per item and average, or run each comparison twice with the order swapped.
Rubric drift across model versions. When you swap in a new judge model version, re-run your calibration set before trusting the new judge's scores against your historical baseline. Different model versions interpret ambiguous rubric wording differently, and a metric jump might be an artifact of the judge changing, not the system under test.
Treating the judge's output as ground truth forever. Rubric scoring reduces variance and adds structure, but it's still an LLM's judgment. Re-sample a small percentage of judged items for human review on an ongoing basis, especially for the criteria that gate deploys, so a slow calibration drift doesn't go unnoticed for months.
FAQ
What's the difference between rubric scoring and a simple LLM-as-judge setup? LLM-as-judge is the general technique of using a model to evaluate another model's output. Rubric scoring is a specific, more disciplined way to do it: instead of one open-ended quality question, you decompose the judgment into explicit, independently gradable criteria. Plain LLM-as-judge setups often use holistic scoring, which is where most of the instability people complain about comes from.
How many criteria should a rubric have? Enough to cover the distinct ways an output can fail, and no more. Most production rubrics land between 3 and 7 criteria. Beyond that, criteria start overlapping and judge latency and cost grow with little added signal. Start narrow, add a criterion only when you see a failure mode your current rubric doesn't catch.
Can the same model that generated the response also judge it? You can, but expect self-preference bias to inflate scores. If budget allows, use a different model or at least a different prompt/temperature configuration for judging than for generation. For high-stakes gates (release blocking, safety), always use an independent judge.
Do I need human labels to build a rubric, or can I skip straight to LLM judging? You can draft a rubric without human labels, but you should not trust it in production without calibrating against at least a small hand-labeled set first. Fifty to a hundred examples is usually enough to catch the criteria that are ambiguously worded. Skipping this step is the single most common reason teams lose trust in their eval numbers later.
How do I turn rubric scores into a CI gate? Pick a per-criterion pass-rate threshold from your calibration baseline (for example, block deploy if FACTUAL_GROUNDING drops more than 3 points from the last known-good release), run the rubric against a fixed regression set on every PR or nightly build, and fail the build when a threshold is breached. Keep the thresholds per criterion, not blended, so a regression in one dimension doesn't get diluted by improvement in another.
Does rubric scoring replace unit tests for LLM features? No. Deterministic checks (does the JSON parse, is the required field present, is the string under the length limit) should stay as plain assertions, they're cheaper and exact. Reserve rubric-based LLM judging for the genuinely subjective or open-ended dimensions, like factual grounding against unstructured context or tone, where a regex or exact-match check can't do the job.
BootcampA 30-day guided bootcamp: build, harden and ship a production autonomous agent from scratch.
AI AgentsUnderstand how AI agents really work: the loop, the tools, the memory, and why most agent projects fail.