← 提示词库 Anthropic/claude-code/skills/claude-api/shared/evals/eval-audit.md 原文 md
🌐 中英双语对照

Eval health checklist / 评估健康检查清单

This file is loaded whenever an eval is being built (build-eval.md) or climbed on (eval-hillclimb.md). It has two jobs. When you are writing the eval, every item below is a construction requirement - the runner, grader, and case set you produce should satisfy it by default, not after someone flags it. When the user brings an existing eval, it is the verification pass you run before building anything on top of it. Either way, run it once more before the first full paid pass and before round 1 of a hillclimb.

每当一个评估正在被构建(build-eval.md)或被爬坡调优(eval-hillclimb.md)时,都会加载本文件。它有两个职责。当你编写评估时,下面每一项都是构建要求——你产出的运行器(runner)、评分器和用例集应当默认满足它们,而不是等别人指出后才满足。当用户带来一个现有评估时,它是你在其上构建任何东西之前运行的验证环节。无论哪种情况,在第一次完整付费运行之前和 hillclimb 第 1 轮之前,都请再运行一遍本清单。

Before trusting an eval to tell you which model, prompt, or configuration is better, check that the eval itself is sound. A broken eval produces confident-looking numbers that point in the wrong direction, and a hillclimb over a broken eval just multiplies the misdirection: you will "improve" an artifact and ship nothing. In practice the most surprising eval results usually turn out to be bugs in the eval rather than facts about the model, so an hour of auditing up front routinely saves days of chasing phantom differences.

在信任一个评估来告诉你哪个模型、提示词或配置更好之前,先检查评估本身是否可靠。一个有缺陷的评估会产生看似可信却指向错误方向的数字,而在有缺陷的评估上做 hillclimb 只会把误导放大:你会"改进"一个假象而什么也没交付。实践中,最令人意外的评估结果往往是评估本身的 bug,而不是关于模型的事实,因此事先花一小时审计通常能省下数天追逐幻影差异的时间。

The checks are grouped into task design (are the cases right?), harness design (is the scaffolding right?), metrics hygiene (are cost and latency measured correctly?), grader design (is the scoring right?), and can it detect the change you're after (is there enough signal for the decision?). They are written as direct instructions: for each, look at the eval's actual code, config, and data, not its README. The final section, Reporting findings to the user, covers how to communicate what you find; the checks are declarative, but the report to the human is observations and suggestions, since the eval's author almost always has context that justifies choices an outsider would flag.

这些检查分为五组:任务设计(用例对不对?)、评测框架设计(脚手架对不对?)、指标卫生(成本和延迟测得对不对?)、评分器设计(打分对不对?)以及能否检测到你要找的变化(支撑决策的信号够不够?)。它们以直接指令的形式书写:对每一项,请查看评估的实际代码、配置和数据,而不是它的 README。最后一节向用户报告发现讲的是如何沟通你的发现;检查本身是指令式的,但呈给人的报告应是观察与建议,因为评估的作者几乎总有局外人看不到、却足以解释其选择的背景。

【评论】本清单把"读实际代码与数据、而非 README"作为审计原则;评估文档与实现脱节是常见的误差来源,这一点与一般代码审计的做法一致。

Before auditing further, run the eval once end-to-end on a handful of cases, or find a recent results file. A surprising number of eval-quality discussions turn out to be about code that does not currently run.

在进一步审计之前,先在少量用例上端到端运行一次评估,或找到一份近期的结果文件。相当多的评估质量讨论,最后发现讨论的其实是当前根本跑不起来的代码。

1. Task design / 1. 任务设计

These checks concern the cases themselves: what is being asked, what counts as correct, and whether the set as a whole can distinguish between the systems being compared.

这些检查关注用例本身:在问什么、什么算正确,以及整个用例集能否区分被比较的系统。

Auditing case sets at scale / 大规模审计用例集

The harness and grader are code you can read end to end; the case set may be hundreds of items you cannot. Do not try to read every case inline. Work in three tiers:

评测框架和评分器是你可以从头读到尾的代码;用例集却可能是数百条你读不过来的条目。不要试图逐条通读所有用例。分三个层级工作:

Tier 1: programmatic checks over the full set. Write a short script that loads every case and reports: exact- and near-duplicate rate; label or category balance; prompt-length and expected-answer-length distributions; schema validity and missing-field counts; obviously malformed rows. Cheap, exhaustive, and catches skew, duplicates, truncation, and broken rows regardless of set size.

**第一层级:对全集做程序化检查。**写一个简短脚本加载每个用例并报告:精确重复与近似重复率;标签或类别均衡度;提示词长度与期望答案长度分布;schema 有效性与缺失字段计数;明显畸形的行。成本低、覆盖全,且无论集合规模如何都能发现偏斜、重复、截断和坏行。

Tier 2: stratified sample for a close read. Draw twenty to fifty cases, stratified across tags[0] if it exists, otherwise uniformly at random, and apply the per-case checks below to those. Recommend the user read a handful themselves as well - a second pair of human eyes on raw cases catches things no checklist does. (This is what the build-eval inputs sign-off is for; the report's per-case table is the surface.)

**第二层级:分层抽样细读。**抽取二十到五十个用例,若存在 tags[0] 则按其分层,否则均匀随机抽取,并对它们应用下文的逐用例检查。同时建议用户自己也读几个——人眼对原始用例的第二遍审视能发现任何清单都发现不了的问题。(这正是 build-eval 输入签核的用途;报告的逐用例表格是其呈现面。)

Tier 3: per-case LLM auditor. For sets beyond a few hundred items, run one isolated model call per case with a tight audit prompt, collect a structured verdict, and aggregate. Ask before running it - the cost is roughly N cheap-model calls - and offer it explicitly: "I can run a per-case auditor over all N cases, ~$X. Want me to?"

**第三层级:逐用例 LLM 审计器。**对超过数百条的集合,对每个用例运行一次隔离的模型调用,配合精炼的审计提示词,收集结构化判定并汇总。运行前先询问——成本大约是 N 次廉价模型调用——并明确提议:"我可以对全部 N 个用例运行逐用例审计器,约 $X。需要吗?"

A per-case auditor prompt that works well (adapt field names to the eval's schema):

一个效果良好的逐用例审计器提示词(字段名请按评估的 schema 调整):

You are auditing a single case from an evaluation suite. Given the prompt, the reference answer, and a description of how the grader decides pass/fail, flag any of the following. Be conservative - only flag when reasonably confident.

PROMPT:
{prompt}

REFERENCE ANSWER:
{gold}

GRADER BEHAVIOUR:
{grader_description}

For each issue answer yes/no with a one-line reason if yes:
- ambiguous: could two careful experts reasonably disagree on the correct answer?
- gold_suspect: does the reference answer look wrong, incomplete, or arguable?
- answerable_from_memory: could a well-read model answer this without doing the intended work?
- grader_too_strict: are there clearly correct answers the grader as described would reject?
- grader_too_lenient: are there clearly wrong answers the grader as described would accept?
- trivially_cheatable: is there a shortcut that satisfies the grader without solving the task?
- other: anything else that would make this case's result misleading.

Return JSON: {"case_id": "...", "flags": {"ambiguous": {"flagged": bool, "reason": "..."}, ...}, "overall": "ok" | "review" | "broken"}

Cluster by flag type, surface the top issues with example case IDs, and feed them into the report (§6).

按标记类型聚类,列出主要问题及示例用例 ID,并纳入报告(§6)。

The per-case checks (apply to the tier-2 sample):

逐用例检查(应用于第二层级样本):

2. Harness design / 2. 评测框架设计

These checks concern the code around the model call. The central failure mode is conflation: any time a non-model artifact - an infra error, a truncated response, a broken tool, a retry delay - lands in the same column as a genuine model result, the eval attributes to the model something that belongs to the plumbing.

这些检查关注模型调用周边的代码。核心失败模式是混淆(conflation):任何时候,只要一个非模型产物——基础设施错误、被截断的响应、损坏的工具、重试延迟——落进与真实模型结果相同的列,评估就把本属于管道问题的东西归到了模型头上。

3. Metrics hygiene / 3. 指标卫生

Pass rate alone rarely answers the user's real question, which is some form of "what quality can I get for what cost and latency?" Check that each perf metric reflects the model under test rather than the rig around it.

仅有通过率很少能回答用户的真实问题,后者总是某种形式的"以怎样的成本和延迟能换来怎样的质量?"。要检查每项性能指标反映的是被测模型,而不是围绕它的装置。

4. Grader design / 4. 评分器设计

These checks concern the function that turns an output into a score - exact match, unit test, end-state check, or LLM judge.

这些检查关注把输出变成分数的函数——精确匹配、单元测试、终态检查,或 LLM 评审。

When the grader is an LLM judge / 当评分器是 LLM 评审时

5. Can it detect the change you're after? / 5. 它能否检测到你要找的变化?

An eval can be correct on every item above and still be useless for the decision at hand because it lacks the resolution to see the effect. Check this before the first full pass and again before round 1 of a hillclimb - discovering it after several paid rounds is the expensive way.

一个评估可以在上述每一项上都合格,却仍对手头的决策没用——因为它缺乏看清该效应所需的分辨率。在第一次完整运行之前检查这一点,并在 hillclimb 第 1 轮之前再查一次——在几个付费轮次之后才发现,是最昂贵的路径。

6. Reporting findings to the user / 6. 向用户报告发现

The checks above are directives to you; the report you hand the human should not read as one. The person who built the eval almost always has context you lack - a constraint, a deadline, a deliberate trade-off - and the purpose is to surface things worth a second look, not to grade their work.

上述检查是给你的指令;你交给人的报告不应读起来也像指令。评估的作者几乎总有你欠缺的背景——一个约束、一个期限、一个有意的权衡——报告的目的是浮出值得再看一眼的东西,而不是给他们的工作打分。

Two practices worth suggesting regardless of what the audit finds: treat the eval as a living suite (new production failure modes become cases, saturated items are hardened, the judge is re-calibrated when it drifts); and periodically have a strong model read the cases, rubric, and a few graded transcripts and ask where a reasonable person would disagree with the label - the tier-3 auditor is the scaled-up version of that.

无论审计结果如何,都值得建议两个做法:把评估当作活的套件(新的生产失败模式变成新用例,饱和条目被加固,评审漂移时重新校准);以及定期让一个强模型阅读用例、评分细则和若干已评分的轨迹,问它哪里讲道理的人会不同意该标签——第三层级审计器就是这一做法的放大版。