← 提示词库 Anthropic/claude-code/skills/claude-api/shared/evals/report/SCHEMA.md 原文 md
🌐 中英双语对照

Hillclimb state schema (v2) / Hillclimb 状态模式(v2)

state.json is the single handoff between an adapter (which reads
whatever your run directory looks like) and the renderer (which
produces report.html). Every field below is optional unless marked
required - the renderer shows what is present and hides what is
absent, so a minimal state with just metrics, variants and
examples renders fine, and a maximal one with reps, splits, judge
explanations, attachments and CIs renders all of those too.

state.json 是 adapter(适配器,负责读取你的运行目录的任何形态)与 renderer(渲染器,负责生成 report.html)之间唯一的交接物。除非标记为 required(必需),下文每个字段都是可选的——渲染器显示存在的内容、隐藏缺失的内容,因此只含 metrics、variants 和 examples 的最小状态可以正常渲染,而包含 reps、splits、评审解释、附件和置信区间(CI)的最大状态也能全部渲染。

Dialect: JSON. Arrays preserve order. Field names are snake_case.

方言:JSON。数组保持顺序。字段名为 snake_case。

Built-in adapter tolerance. adapter.load() is forgiving about
the on-disk input: in results.jsonl the case id may be spelled
prompt_id, id, or case_id; if _state.json omits metrics
they are inferred from the union of grade keys. The schema below
is what the adapter produces, not what it requires.

内置适配器的容错性。 adapter.load() 对磁盘上的输入颇为宽容:在 results.jsonl 中,用例 id 可以写成 prompt_id、id 或 case_id;如果 _state.json 省略了 metrics,则从各 grade 键的并集推断。下文的模式描述的是适配器产出的内容,而不是它要求的内容。

Top level / 顶层

{
  schema: "hillclimb/v2",

  source?: {                      // provenance - shown as a grey header bar
    path:         string,         // relative path of the data dir
    n_files:      number,
    content_sha:  string,         // sha256 over sorted (relpath, file-sha) pairs
    generated_at: string,         // ISO 8601
  },

  metrics: Metric[],              // REQUIRED - what each example is graded on
  perf_fields?: PerfField[],      // runtime fields to surface (default set below)

  variants: Variant[],            // REQUIRED - baseline first
  examples: Example[],            // REQUIRED - every row in the eval set
  metrics_md?: string,            // free-text rubric (markdown)

  // The next three are stderr/--check only: load() returns them in memory for
  // build-report.mjs to print, but they are NOT written to state.json or
  // report.html (they can carry absolute paths and fs error text).
  warnings?: string[],            // adapter diagnostics - stderr under --check
  errors?:   string[],            // only; build-report strips all three before
  trace_stats?: object[],         // writing state.json / report.html

  strtab?: { [key]: string },     // report.html embed only (never state.json):
                                  // strings >=1 KB that repeat across transcripts
                                  // are stored once here and referenced as
                                  // "\u0001S:<key>"; hc-adapt.js resolves them at
                                  // load. Tool payloads >24 KB are also clipped
                                  // in the embed with a pointer to the trace file.

  summary?: {
    narrative?:   string,         // markdown - model-authored running exec
                                  // summary; rewritten after every round,
                                  // finalized as the 4-part summary in Step 5
    best_variant?: string,        // variant id
    headline_metric?: string,     // metric id that test/val/train below
                                  // were computed over; titles the
                                  // score-by-split chart
    test?:  SplitScore,           // headline - shown with CI bars
    val?:   SplitScore,
    train?: SplitScore,
  },
}

Metric / 指标

{
  id:     string,                 // REQUIRED - key used in scores{}
  label?: string,                 // defaults to id; keep <=14 chars - the
                                  // legend has limited width and truncates
                                  // with an ellipsis
  kind:   "binary" | "float" | "judge",
                                  // binary -> % (n/N);  float -> mean±sd;
                                  // judge -> float score with per-rep `explanation`
  scale?: number,                 // upper bound of the raw score range;
                                  // default: 1 for binary, 10 for float/judge.
                                  // Set explicitly for anything else (e.g. 5, 100).
  better?: "higher" | "lower",    // default "higher"; drives delta colouring
}

PerfField / 性能字段

{ id: string, label?: string, unit?: string }

If perf_fields is absent the renderer uses the default set:
cost_usd, in_tokens, out_tokens, web_searches, tool_calls,
latency_s. The built-in adapter passes perf_fields (and
metrics) through from .claude/hillclimb/<flow>/_state.json when
present, so writing that file is how you override the columns
without writing a custom adapter.

如果缺少 perf_fields,渲染器使用默认集合:cost_usd、in_tokens、out_tokens、web_searches、tool_calls、latency_s。当 .claude/hillclimb/<flow>/_state.json 存在时,内置适配器会把其中的 perf_fields(和 metrics)透传过来,因此写这个文件就是在不编写自定义适配器的情况下覆盖这些列的方法。

Variant / 变体

{
  id:      string,                // REQUIRED - "baseline", "v1", ...
  label?:  string,
  description?: string,
  target?: "system_prompt" | "skill" | "tools" | "code",
  change_rationale?: string,      // markdown - rendered above the diffs
  diffs?: {
    incremental: [{ rel_path: string, unified_diff: string }],  // vN vs vN-1 (change.patch)
    cumulative:  [{ rel_path: string, unified_diff: string }],  // vN vs baseline (recomputed from snapshots)
  },
  model?: string | string[],      // distinct row.model values; "mixed" chip if >1
  suspicious?: { note: string },  // renderer shows a WARNING badge + tooltip
  errors?: { total: number, by_class: { [cls]: number }, truncated: number },
                                  // failed attempts from errors.jsonl + status:truncated rows;
                                  // shown as a "WARNING N not scored" badge, never in the means
  metrics?: { [metric_id]: number },
                                  // summary-only metrics ONLY - metrics that
                                  // appear in examples[].results are ignored
                                  // here (the UI derives those from the rows)
  paired?: { [split]: { [metric_id]: PairedDelta } },
                                  // paired per-case delta vs baseline, per
                                  // criterion. The renderer uses .significant
                                  // to gate cell heat-tinting (within-noise ->
                                  // neutral); the numbers stay here for audit.
}

The first variant is treated as the baseline. summary.best_variant
names the winner; if absent, the last variant is assumed.

第一个变体被视为基线。summary.best_variant 指出获胜者;若缺失,则假定最后一个变体为获胜者。

PairedDelta / 配对差值

{
  mean:  number,                  // mean of per-case (variant_mean - ref_mean)
  ci_lo: number, ci_hi: number,   // Wald CI over per-case deltas
  n:     number,                  // cases present in BOTH variants
  significant: boolean,           // CI excludes zero
}

A paired comparison: for each case present in both variants, take the
mean across that variant's reps minus the mean across the reference's
reps, then a CI over those per-case deltas. More powerful than
comparing two SplitScore CIs because between-case variance cancels -
two variants' unpaired CIs can overlap while the paired delta is
clearly non-zero.

配对比较:对同时出现在两个变体中的每个用例,取该变体各次复现的均值减去参照变体各次复现的均值,然后对这些逐用例差值计算置信区间。比比较两个 SplitScore 置信区间更有效,因为用例间方差被抵消——两个变体的非配对置信区间可以重叠,而配对差值却明显非零。

Example / 用例

{
  id:       string,               // REQUIRED
  prompt:   string,               // REQUIRED
  split?:   "train" | "val" | "test",
  tags?:    string[],             // ORDERED - tags[0] is the primary
                                  // grouping key the UI clusters rows by
                                  // (replaces v1's singular `category`);
                                  // further entries are secondary filters
  meta?:    { [k]: any },         // arbitrary sidecar data
  attachments?: Attachment[],     // input artifacts - render above the first
                                  // user turn in the transcript view
  results: { [variant_id]: RepResult[] },   // REQUIRED (may be empty per variant)
}

Attachment / 附件

{
  kind?: "image" | "svg" | "html" | "pdf" | "json" | "text" | "code"
       | "file" | "url",          // inferred from ref if omitted
  ref:  string,                   // path relative to the flow root, data: URI,
                                  // or URL. Paths under 2 MB are inlined as
                                  // data: at build time; larger -> download chip.
  alt?: string,
}

image/svg render inline; html in a sandboxed scrollable iframe; pdf
via the browser's native viewer in a scrollable embed; json/text/code
in a <pre>; file (docx/pptx/anything else) and url as a download/open
chip. Every kind has a Hide/Show toggle.

image/svg 内联渲染;html 在沙箱化的可滚动 iframe 中渲染;pdf 通过浏览器原生查看器在可滚动嵌入中渲染;json/text/code 在 <pre> 中渲染;file(docx/pptx/其他任何格式)和 url 以下载/打开小部件形式渲染。每种类型都有隐藏/显示切换开关。

RepResult / 单次复现结果

{
  rep?:        number,            // 0-based; default = array index
  status?:     string,            // present only when not 'ok' (e.g. 'truncated'); scores is {} then
  scores:      { [metric_id]: number },
  explanation?: { [metric_id]: string },    // judge rationale per metric
  model?:      string,            // model id that produced this rep (from the response)
  perf?:       { [perf_field_id]: number },
  attachment?: string,            // relative path to a per-rep output screenshot
  transcript?: Turn[],
}

Turn / 对话轮次

{
  role: "system" | "user" | "assistant" | "tool_call" | "tool_result",
  content:  string,               // markdown for user/assistant/system;
                                  // pretty-printed args/result for tool turns
  name?:    string,               // tool name (tool_call / tool_result)
  thinking?: string,              // assistant extended-thinking (collapsible)
  attachments?: Attachment[],     // artifacts produced/consumed at this turn -
                                  // render below the turn content. Use this for
                                  // files the model wrote, generated plots, etc.
}

On the input side, the built-in adapter reads traces/<id>.json
directly as a Turn[] list - each tool call / result is its own
{role: "tool_call", name, content} / {role: "tool_result", content}
entry. See build-eval.md §Step 3 for the trace-writing spec.

在输入侧,内置适配器把 traces/<id>.json 直接读取为 Turn[] 列表——每次工具调用/结果都是独立的 {role: "tool_call", name, content} / {role: "tool_result", content} 条目。轨迹写入规范见 build-eval.md §Step 3。

SplitScore / 分组得分

{
  score:  number,
  ci_lo?: number,
  ci_hi?: number,
  n?:     number,
  significant?: boolean,          // vs baseline - greys out & badges "within noise" when false
}

Rendering rules / 渲染规则

Writing your own adapter / 编写自己的适配器

adapter.load(path) -> dict is the only contract. If your data is not
laid out like .claude/hillclimb/<flow>/, write a function that reads
whatever you have and returns a dict matching this document, then call
render.render(state) directly (see build-report.mjs for the
one-liner). The renderer has no opinion about where the data came from.

adapter.load(path) -> dict 是唯一的契约。如果你的数据不是按 .claude/hillclimb/<flow>/ 的布局存放,就编写一个函数读取你手头的数据并返回符合本文档的 dict,然后直接调用 render.render(state)(一行写法见 build-report.mjs)。渲染器不关心数据从哪里来。

Pages beyond report.html / report.html 之外的页面

The builder's report.html stays the deliverable, and the build-eval
grading sign-off stays report.html too. Write a page yourself only
where a guide has you make one or the user asks for something
report.html does not show (with the lite report: a chart, the diff
on the page, a dashboard). If they already have a viewer they like,
use that instead. Build what they asked for and link to report.html
for the rest. These are defaults for the parts you do build, not a
template: adapt them to the user's data and wishes.

构建器的 report.html 仍是交付物,build-eval 的评分签署也仍是 report.html。只有当指南要求你创建页面、或用户要求 report.html 未展示的内容时(就 lite 报告而言:一张图表、页面上的 diff、一个仪表盘)才自己写页面。如果他们已有喜欢的查看器,就用那个。构建用户要求的部分,其余内容链接到 report.html。以下是你自行构建部分时的默认规范,而非模板:应根据用户的数据和意愿加以调整。

Any page:

任何页面:

A page of results follows these too. Run the builder first. Then
compute every number from the files - results.jsonl, _state.json
(split, best), errors.jsonl, vN/change.* - and never type one in.
Take per-case scores from trajectory/scores.tsv, which the builder
writes (with no node or bun to run it, compute them the same way
from results.jsonl), and take means the builder's way (per case
over status-ok reps, then over cases), so the page agrees with
report.html:

结果页面同样遵循这些规则。先运行构建器。然后从文件中计算每个数字——results.jsonl、_state.json(split、best)、errors.jsonl、vN/change.*——绝不要手工键入。逐用例得分取自构建器写出的 trajectory/scores.tsv(若没有可运行的 node 或 bun,就按同样方式从 results.jsonl 计算),均值按构建器的方式计算(每个用例先对状态为 ok 的复现取均值,再对用例取均值),使页面与 report.html 一致: