← 提示词库 Anthropic/claude-code/skills/claude-api/shared/preserved-thinking-migration.md 原文 md
🌐 中英双语对照

Preserved Thinking - Keeping Earlier Reasoning Valid Across a Conversation / 保留思考——让早前推理在整个对话中保持有效

If you arrived via /claude-api preserved-thinking-migration (or opened this file directly): this is the right file. Execute the steps below in order rather than summarizing the guide back to the user - presenting the break profile, the ranked causes, and the measured result of each fix IS part of the execution. Start with Step 0 (scope, quality bar, baseline) and finish with Step 4's two deliverables: the break profile and the changes.

**如果你是通过 /claude-api preserved-thinking-migration 到达此处(或直接打开了本文件):**文件没错。请按顺序执行下列步骤,而不是把本指南向用户复述一遍——呈现断点画像、按排序的成因清单以及每项修复的实测结果,本身就是执行的一部分。从第 0 步(范围、质量标准、基线)开始,以第 4 步的两项交付物收尾:断点画像与变更内容。

Preserved thinking is measured in units of conversations that keep their reasoning, not requests that pass. One edit to an earlier turn invalidates every thinking block after it, and the same stale block fails again on every later request that replays it - so a per-request count overstates the damage and a per-conversation count (did this conversation break, and at which turn) is the number that tells you whether a fix worked.

保留思考(Preserved thinking)的度量单位是保住推理的对话数,而不是通过的请求数。对早前某一轮的一处编辑会使其后所有 thinking 块失效,而同一个失效块在其后每个重放它的请求上都会再次失败——因此按请求计数会夸大损害,按对话计数(这条对话是否中断、中断在第几轮)才是能说明修复是否生效的数字。

What the check is, in one paragraph. On models with preserved thinking (Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5.5 today; the Preserved thinking page lists the models and the enforced accounts, and says later models will enforce the check for all users), each thinking block's signature records the conversation that produced it - the top-level system prompt, the set of tools, and every message before the block - and the model that produced it. When the transcript comes back, the API recomputes that record from what you sent and requires a match (a separate model check decides whether the current model can read the block at all; see "Switching models mid-conversation" in shared/preserved-thinking-migration/causes.md). Integrations that keep the history append-only never notice. Integrations that rewrite earlier turns between requests - truncation, client-side compaction, a re-rendered system prompt, a tool list that grows when a plugin connects, a per-turn reminder that is injected and then stripped, old tool results trimmed after the fact, media dropped by a size cap, a lossy round trip through the app's own message types - lose the reasoning after the edit point (drop_block) or fail the request (error). The check compares the conversation as you sent it, before any server-side edit, so Anthropic's own server-side compaction and context editing never count as edits. The published explanation lives in shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> Breaking change 3 (the three-step check and the append-only form of each edit; read it first if the check is new to you) - this workflow restates the rules only as a lookup, the cause table and keep list in shared/preserved-thinking-migration/causes.md; it finds which edits this harness makes, proves them with the API's own response, fixes them one at a time, and proves the fix the same way.

**这项检查是什么,一段话说清。**在支持保留思考的模型上(当前为 Claude Fable 5.1、Claude Opus 5.5 与 Claude Sonnet 5.5;Preserved thinking 页面列出了相关模型与被强制执行的账号,并说明后续模型将对所有用户强制执行该检查),每个 thinking 块的 signature 记录了产生它的对话——顶层 system 提示词、tools 集合以及该块之前的每条消息——以及产生它的模型。当对话记录被再次发回时,API 会根据你所发送的内容重新计算该记录并要求二者一致(另有独立的模型检查决定当前模型能否读取该块;见 shared/preserved-thinking-migration/causes.md 中的 "Switching models mid-conversation")。保持历史只追加(append-only)的集成永远不会碰到问题。会在请求之间改写早前各轮的集成——截断、客户端压缩、重新渲染的系统提示词、插件接入后变长的工具列表、先注入后剥离的每轮提醒、事后裁剪的旧工具结果、因大小上限被丢弃的媒体、经由应用自身消息类型的有损往返——会在编辑点之后丢失推理(drop_block)或导致请求失败(error)。该检查比较的是你发送时的对话、在任何服务端编辑之前的状态,因此 Anthropic 自己的服务端压缩与上下文编辑从不计为编辑。公开说明位于 shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> Breaking change 3(三步检查与每种编辑的只追加形式;如果你初次接触该检查,请先读它)——本工作流只是把规则当作查询来复述,完整的成因表与保留清单在 shared/preserved-thinking-migration/causes.md;它找出本 harness 造成哪些编辑,用 API 自己的响应加以证明,逐一修复,并以同样的方式证明修复有效。

【评论】该检查把对话身份绑定到历史前缀的字节级一致性上,保持只追加的集成天然免检;这一机制与提示词缓存的前缀命中逻辑同构,因此两类的修复往往是同一次改动。

Where this workflow sits: the migrate subcommand explains the check and lists the append-only form of each edit; this workflow is its executable form - scan, measure, fix, re-measure - for a harness that already exists. cost-optimize is a sibling, not a prerequisite: a history that invalidates its own thinking also restarts the prompt cache at the same point, so a fix here usually shows up in its cache-hit numbers too, and a prefix edit found there ("audit for mid-task cache-breakers") is the same finding as a break found here. Cache discipline and preserved-thinking discipline are very nearly the same discipline, so a harness that is already append-only for caching pays nothing extra here. Once the project has an eval, Step 3 runs as a hill-climb whose metric is the drop count - one change per round, measured, kept or reverted - and the hillclimb subcommand is the loop to use.

本工作流的定位:migrate 子命令解释该检查并列出每种编辑的只追加形式;本工作流是其可执行形态——扫描、度量、修复、重新度量——面向一个已经存在的 harness。cost-optimize 是同级工具,而非前置条件:使自身 thinking 失效的历史也会在同一点重启提示词缓存,因此这里的修复通常也会体现在其缓存命中数字上,在那里发现的前缀编辑("audit for mid-task cache-breakers")与在这里发现的断点是同一个发现。缓存纪律与保留思考纪律几乎是同一种纪律,因此为缓存已经做到只追加的 harness 在这里无需额外付出。一旦项目有了评测集,第 3 步就作为以丢块数为指标的山攀式迭代运行——每轮一个改动、度量、保留或回退——hillclimb 子命令就是该循环。

Two scripts ship with this guide, extracted beside it under shared/preserved-thinking-migration/, and are used by Steps 1 and 2. Both are dependency-free Python 3; neither needs credentials except the probe's live modes, and neither prints them. The commands below give their paths relative to this skill's base directory (the line at the top of the prompt); run them with that directory prefixed, from the user's project directory, so that relative capture paths resolve there. A reference file is extracted beside them, shared/preserved-thinking-migration/causes.md: the rules for a conversation that switches models, the "Cause -> detection -> fix" table, the keep list, and the failure modes to avoid. It is a separate file so that this guide fits in one Read; Read it when a step below sends you there (Step 1.4 at the latest), not before.

本指南附带两个脚本,随指南一并解包在 shared/preserved-thinking-migration/ 下,供第 1 步和第 2 步使用。两者都是无第三方依赖的 Python 3;除探针的实跑模式外都不需要凭证,也都不会打印凭证。下文命令给出的是相对于本 skill 基础目录(提示词顶部那一行)的路径;请以该目录为前缀、在用户的项目目录下运行,使相对的捕获路径在彼处解析。旁边还解包了一个参考文件 shared/preserved-thinking-migration/causes.md:切换模型的对话规则、"Cause -> detection -> fix" 表、保留清单以及要避免的失败模式。它单独成文件是为了让本指南一次 Read 就能读完;当下文某一步让你去读时(最迟第 1.4 步)再读,不要提前。


Severity tiers - one-off vs. recurring / 严重度分层——一次性问题与反复发生的问题

Tier every cause before planning the work. Tiers 0-2 are one-time fixes: apply them once and the harness stops breaking. Tiers 3-4 keep costing - they recur on every conversation that reaches them, which is why they are the ones to measure before deciding.

在规划工作之前,先给每个成因分层。第 0-2 层是一次性修复:执行一次,harness 就不再中断。第 3-4 层会持续产生损失——它们在每条触及它们的对话上反复出现,这正是要在决策前先对其度量的原因。

Tier Meaning Causes What to do
0 Fine - not an edit cache_control markers; reordering the tools array; server-side compaction and context editing; a retry; a regenerate, rewind, restored checkpoint or branch; the latest turn edited and resubmitted; a model switch (the model check is separate and is not a prefix edit) Nothing. None of these change the compared prefix - keep them out of the report
1 Accidental The system prompt re-rendered with per-request content (a date, a counter, live state); drift from an SDK or domain-model round trip Remove the edit. Nothing about the product needs it
2 Fixable The tool set changed mid-conversation; a same-name tool's description or schema rebuilt; a per-turn reminder injected then stripped; tail state re-rendered every turn Apply the append-only recipe (Step 3). Adding or withdrawing a tool has an append-only form (tool_addition / tool_removal, with the entry left in tools). A same-name description or schema change has one only under the inline-tools-2026-09-15 beta (Claude API), where a tool_addition carries the new definition (Step 3); without it, keep the first-sent bytes for the life of the conversation and accept the stale definition, or offer the changed text under a new name. Append tail state as new turns rather than rewriting it
3 Recurring Tail-kept compaction; background compaction (a second request writes the summary while the session continues, then it is swapped in); rolling truncation; a pinned document rewritten every turn Measure the drops and decide. Without the compact-2026-09-04 beta (on-demand compaction) no append-only client-side form exists for these: send drop_block from the swap onward, or strip the thinking from the kept turns; the recommended shape is simple compaction, done synchronously. With the beta, keep-tail and background compaction become append-only - see Step 3
4 Stop Prefix surgery - snipping, redacting, pruning old tool results, or removing content after a cache breakpoint to save cost; a missing predecessor; an unrecorded strip-and-retry No workaround exists. The reasoning after the edit point is lost; the fix is to stop doing it
层级 含义 成因 处理方式
0 正常——不是编辑 cache_control 标记;对 tools 数组重新排序;服务端压缩与上下文编辑;一次重试;重新生成、回退、恢复检查点或分支;对最新一轮的编辑与重新提交;模型切换(模型检查是独立的,不属于前缀编辑) 无需处理。以上均不改变被比较的前缀——不要把它们写进报告
1 意外 系统提示词被带每次请求内容(日期、计数器、实时状态)重新渲染;SDK 或领域模型往返造成的漂移 移除该编辑。产品并不需要它
2 可修复 工具集在对话中途变化;同名工具的描述或 schema 被重建;每轮提醒先注入后剥离;尾部状态每轮重新渲染 应用只追加配方(第 3 步)。新增或撤下工具有只追加形式(tool_addition / tool_removal,条目保留在 tools 中)。同名描述或 schema 变更只有在 inline-tools-2026-09-15 beta(Claude API)下才有只追加形式,此时 tool_addition 携带新定义(第 3 步);没有它时,就在对话存续期内保留首次发送的字节、接受过时定义,或者以新名称提供变更后的文本。尾部状态以追加新轮的方式提交,而不是重写
3 反复发生 保留尾部的压缩;后台压缩(会话继续的同时由第二个请求写入摘要,之后再换入);滚动截断;每轮重写的固定文档 度量丢块并决策。没有 compact-2026-09-04 beta(按需压缩)时,这些没有客户端只追加形式:从换入点起发送 drop_block,或剥离所保留各轮的 thinking;推荐形态是同步进行的简单压缩。有该 beta 时,保留尾部压缩与后台压缩可以变为只追加——见第 3 步
4 停止 前缀手术——为省成本而删节、遮蔽、修剪旧工具结果,或移除缓存断点之后的内容;缺失的前驱;未记录的"剥离后重试" 不存在变通办法。编辑点之后的推理已经丢失;修复方式就是停止这样做

A tier is a property of the cause, not of one conversation: rank by tier first, then by reasoning lost within a tier (Step 3).

层级是成因的属性,不是单条对话的属性:先按层级排序,再在层内按丢失的推理量排序(第 3 步)。

【评论】分层把"能否一次性修复"与"是否反复发生"区分开,让度量精力集中到持续产生损失的成因上,实质是一个按持续风险排定优先级的框架。

Step 0: Establish scope, quality bar, and baseline / 第 0 步:确定范围、质量标准与基线

Does your harness change earlier turns, the system prompt, or the tool list? If not, stop - there is nothing to migrate, and saying so plainly is the finding.

你的 harness 是否会改动早前各轮、系统提示词或工具列表?如果不会,就此停止——没有需要迁移的内容,直说这一点本身就是结论。

When a replayed thinking block no longer matches, the default is a 400 error. Dropping the thinking instead is opt-in, and it is not a fix: it trades a visible failure for the silent loss of that reasoning, and it is not free - dropped blocks aren't billed, but the session's token usage might still increase because Claude can sometimes think more to re-create the dropped thinking ("Failure modes to avoid" in causes.md).

当重放的 thinking 块不再匹配时,默认行为是 400 错误。改为丢弃 thinking 需要显式选择,而且那不是修复:它用一次可见的失败换来推理的无声丢失,并且并非没有代价——被丢弃的块不计费,但会话的 token 用量仍可能上升,因为 Claude 有时会为重建被丢弃的思考而多思考(causes.md 中的 "Failure modes to avoid")。

First, establish three things - from the request and the repository where they answer it, and from the user where they don't. This workflow is interactive by design: a capture of real request bodies, a test slice, and every live replay need the user's involvement or approval, and "which of these edits is deliberate" is a question only they can answer. State all three at the top of the report (the baseline may read "pending Step 2" at first).

**首先确定三件事——请求与代码仓库能回答的从那里取,不能回答的向用户问。**本工作流在设计上是交互式的:真实请求体的捕获、测试切片以及每一次实跑重放都需要用户的参与或批准,而"这些编辑中哪些是有意为之"只有用户能回答。在报告开头写明这三项(基线最初可以标注为"待第 2 步")。

  1. Scope. If an official Claude product or SDK (Claude Code, claude.ai, Claude Managed Agents, the Claude Agent SDK) manages the conversation history, there is nothing to migrate - say so and stop. Otherwise: if the request names files or directories, that is the scope. Otherwise it is every place the project builds the three parts of a request the check compares - the system prompt, the tools array, and the messages array - across every path that touches them between two requests of one conversation: the request builder, compaction or truncation, reminder or context injection, media handling, persistence and resume (anything that re-reads the conversation from a store and re-renders it), process restart and deploy, plugin or MCP connection, sub-agent transcripts, and model switching. Note distinct traffic classes (an interactive chat path and a background agent loop are different harnesses even on one key): the profile, the fixes, and every validation later run per class. A capture taken on an older model may carry request shapes Claude Fable 5.1 rejects before any check runs - thinking.type enabled or disabled, a forced tool_choice (any or tool), an assistant prefill as the last message, temperature, top_p or top_k with a thinking configuration - so list them now as things to convert before measuring (Step 2.1 lists what the probe converts; it leaves a prefill alone, and that 400 shows in its "not evaluated" line). Also establish which platform the code targets (Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud Vertex AI, or Microsoft Foundry) and which model it runs: the check applies only to models with preserved thinking. The Preserved thinking page states that beta names are the same on Amazon Bedrock and Google Cloud wherever the beta is available there; for any other platform, read that platform's own documentation and shared/platform-availability.md rather than assuming. Finally establish whether the organization is enforced today. The rule, as the Preserved thinking page reads on 2026-09-23: on Claude Fable 5.1 and Claude Opus 5.5 the check is enforced by default for accounts created on or after August 31, 2026, 00:00 UTC, with the same definition on the Claude API and on the cloud platforms; a request that sets prefix_mismatch_behavior opts in regardless of account age; and later models will enforce the check for all users. Confirm it from the responses production already gets: an enforced account sees input_transformations entries (with the header) or 400s whose text says "bound to a different conversation"; an account that is not enforced yet sees no dropped blocks and no 400s until a request sets the field; with the beta header alone, its responses list each block that would fail as a thinking_mismatch_allowed entry (Step 2.2). The page's own probe - send an edited history without the beta header: a 400 that names the header means the account is enforced by default, a 200 means it is not (the Message Batches API returns 200 either way). To confirm a 200, resend with the beta header and still no field: the response lists every thinking block after the edit as thinking_mismatch_allowed. Step 0.5 proves the enforced path - run it with a test key, never against production traffic.
    范围。如果由官方 Claude 产品或 SDK(Claude Code、claude.ai、Claude Managed Agents、Claude Agent SDK)管理对话历史,就没有需要迁移的内容——说明这一点并停止。否则:如果请求指定了文件或目录,那就是范围。否则范围是项目中构造该检查所比较的请求三要素——system 提示词、tools 数组和 messages 数组——的每一处,覆盖同一对话的两次请求之间触碰它们的每条路径:请求构造器、压缩或截断、提醒或上下文注入、媒体处理、持久化与恢复(任何从存储重读对话并重新渲染它的逻辑)、进程重启与部署、插件或 MCP 接入、子代理转录以及模型切换。区分不同的流量类别(交互式聊天路径与后台代理循环即使共用一个密钥也是不同的 harness):画像、修复以及此后的每项验证都按类别分别进行。在旧模型上采集的捕获可能带有 Claude Fable 5.1 在任何检查运行之前就会拒绝的请求形态——thinking.type 为 enabled 或 disabled、强制的 tool_choice(any 或 tool)、以助手预填充作为最后一条消息、thinking 配置下携带 temperature、top_p 或 top_k——因此现在就把它们列为度量前需要转换的事项(第 2.1 步列出探针会转换哪些;它不改动预填充,相应的 400 会出现在其 "not evaluated" 行中)。同时确定代码面向哪个平台(Claude API、Claude Platform on AWS、Amazon Bedrock、Google Cloud Vertex AI 或 Microsoft Foundry)以及运行哪个模型:该检查只适用于支持保留思考的模型。Preserved thinking 页面说明,在该 beta 可用的范围内,beta 名称在 Amazon Bedrock 与 Google Cloud 上相同;对其他任何平台,请阅读该平台自己的文档与 shared/platform-availability.md,不要臆测。最后确定组织当前是否被强制执行。按 Preserved thinking 页面在 2026-09-23 的表述:在 Claude Fable 5.1 与 Claude Opus 5.5 上,对 2026 年 8 月 31 日 00:00 UTC 及以后创建的账号,该检查默认强制执行,Claude API 与各云平台上的定义相同;设置了 prefix_mismatch_behavior 的请求无论账号新旧一律选择加入;后续模型将对所有用户强制执行该检查。可从生产环境已有的响应中确认:被强制执行的账号会看到 input_transformations 条目(带请求头时)或正文写明 "bound to a different conversation" 的 400;尚未被强制的账号在请求设置该字段之前不会看到丢块或 400;仅带 beta 请求头时,其响应会把每个本会失败的块列为 thinking_mismatch_allowed 条目(第 2.2 步)。页面自带的探测法——发送一份被编辑过的历史且不带 beta 请求头:返回指名该请求头的 400 表示账号默认被强制执行,返回 200 表示尚未(Message Batches API 无论如何都返回 200)。要确认 200,可带上 beta 请求头且仍不设该字段重发:响应会把编辑点之后的每个 thinking 块列为 thinking_mismatch_allowed。第 0.5 步验证被强制的路径——务必用测试密钥运行,绝不要对生产流量运行。
  2. Quality bar. Find the project's eval, test suite, or outcome checks for its model calls. The fixes in this workflow are behavior-preserving by construction (they change how the history is carried, not what the model is told), but two of them are not: replacing a compaction scheme, and choosing to drop thinking at a boundary. Those need the eval. If none exists, say so prominently in the report and do not stop: the drop count itself is a measurement, every fix that removes an edit is safe to propose, and the minimal eval recipe in Step 3 is the next step for the two that aren't.
    **质量标准。**找到项目针对其模型调用的评测集、测试套件或结果检查。本工作流中的修复在构造上是保行为的(它们改变的是历史如何被携带,而不是模型被告知的内容),但有两个例外:替换压缩方案,以及选择在某一边界丢弃 thinking。这两项需要评测集。如果没有,就在报告中显著说明,但不要停下:丢块计数本身就是一种度量,每个移除编辑的修复都可以放心提出,而对那两项例外,第 3 步的最小评测配方是下一步。
  3. Baseline. Two numbers, both from Step 2's probe on the same test slice: the share of conversations with at least one prefix break, and the turn at which each one first breaks. Record them before any fix. The three-arm protocol in Step 2 also asks for the eval score with the current harness and preserved thinking off (that is the eval's existing number) - write it down now if it exists.
    **基线。**两个数字,均来自第 2 步探针在同一测试切片上的结果:出现至少一次前缀断点的对话占比,以及每条对话首次断点所在的轮次。在任何修复之前先记录它们。第 2 步的三臂协议还要求提供当前 harness 在保留思考关闭时的评测分数(即评测集现有的数字)——如果存在,现在就记下来。

Step 0.5: Prove the mechanism is wired / 第 0.5 步:证明机制已经生效

Before trusting any measurement, run one deliberately broken conversation and confirm the API reports it. The probe's self-test sends three requests - one mint and two replays: it mints a thinking block with a one-turn conversation, replays it honestly as turn two (expect input_transformations: []), then replays it with the first user message edited (expect one {"type": "thinking_dropped", ...} entry with reason: "prefix_binding_mismatch", and - on the Claude API - the diagnosis header naming pattern=first_message_rewritten):

在信任任何度量之前,先运行一条故意弄坏的对话,并确认 API 会报告它。探针的自测发送三个请求——一次铸造、两次重放:先用一轮对话铸造一个 thinking 块,然后诚实地把它作为第二轮重放(预期得到 input_transformations: []),再在编辑第一条用户消息后重放(预期得到一条 {"type": "thinking_dropped", ...} 条目、reason 为 prefix_binding_mismatch,并且在 Claude API 上还有指名 pattern=first_message_rewritten 的诊断响应头):

python3 shared/preserved-thinking-migration/drop_block_probe.py --self-test --model <the target model id> --yes
python3 shared/preserved-thinking-migration/drop_block_probe.py --self-test --model <the target model id> --yes --mode error   # expect the 400 instead

The probe always sends the thinking-binding-controls-2026-08-01 beta value; add --beta <value> for any other beta your production requests carry (the bare command is enough on the Claude API). The probe is first-party only: it authenticates against /v1/messages with ANTHROPIC_API_KEY (sent as x-api-key) or ANTHROPIC_AUTH_TOKEN (sent as a bearer token; when both variables are set only the token is sent, because the API rejects a request carrying both headers) and cannot replay captures taken on Amazon Bedrock or Google Cloud Vertex AI - there, run the account's own client and read input_transformations from its responses.

探针总是发送 thinking-binding-controls-2026-08-01 这个 beta 值;如果生产请求还携带其他 beta,用 --beta <value> 补充(在 Claude API 上,不带附加参数的命令就足够)。探针仅支持第一方平台:它使用 ANTHROPIC_API_KEY(以 x-api-key 发送)或 ANTHROPIC_AUTH_TOKEN(以 bearer token 发送;两个变量都设置时只发送 token,因为 API 会拒绝同时携带两个请求头的请求)对 /v1/messages 进行认证,且无法重放在 Amazon Bedrock 或 Google Cloud Vertex AI 上采集的捕获——在那些平台上,请运行账号自己的客户端并从其响应中读取 input_transformations。

Three short requests (one mint, two replays), billed at the model's normal rates - state the cost and get the user's approval first, as for every run that exercises the model. If the edited replay comes back with no drop (the probe prints NOT WIRED and exits 1), stop: the check is not running for this request (the wrong model id, a platform without the controls, the header missing, the field misspelled, or a gateway or proxy between the harness and the API that drops the anthropic-beta header or the block_binding field - run the self-test through the same path production uses) and every later number would be meaningless. If the honest replay reports a drop, stop too: something in the probe's path is already editing the history - a proxy, an SDK middleware, a serializer - and that is finding number one. Record both responses in the report.

三个短请求(一次铸造、两次重放),按模型正常费率计费——像每次实际调用模型那样,先说明成本并取得用户批准。如果被编辑的重放没有返回丢块(探针打印 NOT WIRED 并以 1 退出),停止:该检查没有在这条请求上运行(模型 id 错误、平台不支持这些控制项、请求头缺失、字段拼错,或 harness 与 API 之间的网关或代理丢弃了 anthropic-beta 请求头或 block_binding 字段——请让自测走生产所用的同一条路径),此后的每个数字都毫无意义。如果诚实重放也报告了丢块,同样停止:探针路径上有东西正在编辑历史——某个代理、某个 SDK 中间件、某个序列化器——这就是第一条发现。把两个响应都记录进报告。

Step 1: Find the edits / 第 1 步:找出编辑点

Finding the edits has three sources, in order of evidence: the API's own diagnosis (Step 2), the diff of consecutive request bodies, and the code. Do the diff and the code read in this step; they tell you where to look before spending money on replays, and they are what localizes a break to a line of code after the API has named its shape.

寻找编辑点有三个来源,按证据强度排序:API 自己的诊断(第 2 步)、相邻请求体的差分,以及代码。差分与代码阅读在本步完成;它们让你在为重放花钱之前知道该看哪里,也是在 API 点明断点形态之后把断点定位到某行代码的手段。

1.1 Capture what the harness actually sends / 1.1 捕获 harness 实际发送的内容

Capture the exact request bodies of a few normal conversations - every request, in send order, JSON as it went over the wire, including the assistant turns with their thinking blocks and signature values exactly as the API returned them. Include one conversation that runs long enough to trigger compaction or truncation if the product has either, one that connects a plugin or tool mid-session if it can, and one that is resumed from storage or survives a process restart. Capture at the HTTP layer where possible (an SDK hook, a logging transport, a proxy) rather than from the application's own message objects - the second kind of capture hides exactly the re-serialization this check catches. If requests pass through a gateway, proxy, or model router on the way to the API, capture them as they leave that layer: a router can rewrite the system prompt, the tools, or the history after your code has built the request. Store one conversation per .jsonl file, one request body per line. Ten to thirty conversations are enough; they double as the test slice for Step 2.

捕获若干正常对话的确切请求体——每个请求、按发送顺序、线上传输的 JSON 原样,包括助手各轮及其 thinking 块与 signature 值(与 API 返回时完全一致)。如果产品具备相应机制,要包含一条长到足以触发压缩或截断的对话、一条能在会话中途接入插件或工具的对话,以及一条从存储恢复或经历了进程重启的对话。尽可能在 HTTP 层捕获(SDK 钩子、日志传输层、代理),而不要从应用自身的消息对象捕获——后一种捕获恰好会掩盖本检查所要抓的重新序列化。如果请求在到达 API 之前经过网关、代理或模型路由器,就在该层出口处捕获:路由器可能在你的代码构造完请求之后改写系统提示词、工具或历史。每条对话存一个 .jsonl 文件,每行一个请求体。十到三十条对话足够;它们可兼作第 2 步的测试切片。

Capture the request body and the anthropic-beta header only - never x-api-key or Authorization; an MCP server's authorization_token in the body is sent as the harness sent it (the connector's tools are part of what the check compares), so capture with a test-scoped token and rotate it afterwards. The probe reads nothing else and warns when a capture carries a credential.

只捕获请求体和 anthropic-beta 请求头——绝不要捕获 x-api-key 或 Authorization;正文中 MCP 服务器的 authorization_token 按 harness 发送的原样保留(连接器的工具是检查所比较内容的一部分),因此要用测试范围限定的令牌捕获,事后再轮换它。探针不读取其他任何内容,并在捕获带有凭证时告警。

Handling captures. A capture is the conversation as the end users had it, and it cannot be redacted without breaking the measurement (the check compares the bytes). Keep captures outside the repository (or ignored by version control), never commit them or an eval set derived from them, run the scripts from a machine that may hold that data, and delete the captures when the work is done. The probe's --json output is safe to share - it holds request ids, statuses, entries, headers, digests, token counts and file names, no message content; error text for the conversation check's own 400s is stored as a reconstruction of their fixed form (the block path, the fixed clause, and the first-changed-message diagnostic - never the server line itself), and every other error is reduced to its type and field path because API validation messages can echo request values (the full text still prints on the terminal); prefix_diff.py's output is not, because its attribution lines quote excerpts of the changed content.

**捕获数据的处理。**捕获是终端用户所见的那份对话,一旦脱敏就会破坏度量(该检查比较的是字节)。把捕获存放在仓库之外(或加入版本控制忽略),绝不提交它们或由其派生的评测集,在允许持有这些数据的机器上运行脚本,工作完成后删除捕获。探针的 --json 输出可以安全分享——其中只有请求 id、状态、条目、请求头、摘要、token 计数与文件名,没有消息内容;对话检查自身 400 的错误文本以固定形式的重构存储(块路径、固定从句与首条被改消息的诊断——绝不是服务端原文行),其他错误一律归约为类型与字段路径,因为 API 校验消息可能回显请求值(完整文本仍会打印到终端);prefix_diff.py 的输出则不可分享,因为其归因行会引用被改内容的片段。

If the application cannot capture bodies yet, adding that capture is itself the first diff of this workflow: it is the measurement channel for everything after it.

如果应用尚不能捕获请求体,那么补上这一捕获能力本身就是本工作流的第一个差分:它是其后一切度量的通道。

1.2 Diff consecutive pairs / 1.2 对相邻请求对做差分

python3 shared/preserved-thinking-migration/prefix_diff.py captures/conversation-0001.jsonl

For each pair of consecutive requests the script reports MATCH, MISMATCH with a verdict in the API's vocabulary - kind=system_changed; pattern=system_rerendered; sections=system; changed_validated=system.0 - and an attribution line that names the site and the first changed character:

对每一对相邻请求,脚本报告 MATCH、MISMATCH 以及用 API 词汇表述的判定——kind=system_changed; pattern=system_rerendered; sections=system; changed_validated=system.0——外加一条指明位置与首个被改字符的归因行:

system[0] changed at char 53: "... Be concise." -> "... Be concise. Current time: 2026-09-02T15:04:05Z."
tools: lookup_order description changed at char 23: "...order by id." -> "...order by id. Today is 2026-09-02."
messages[2] (user) content[1] (text) removed: {"text":"<reminder>Answer in one sentenc...
messages[1..2] removed (assistant, user)

Two lines matter as much as the verdict. replayed thinking blocks in the later request: N - when N is 0 the pair proves nothing about preserved thinking (there was no block to check), which is common for the first pair of every conversation and for harnesses that strip thinking; and the ! chain line (printed as CHAIN-BREAK, counted in the exit status), which is a break, not a warning - it fires when the already-sent turns come back with their thinking blocks changed in a way the API rejects: the kept blocks must be a contiguous window of the original sequence (dropping from the front, from the back, or both is fine), so a block removed from the middle, or a reorder, fails the block after the gap even though the rest of the prefix is untouched.

有两行与判定同样重要。replayed thinking blocks in the later request: N——N 为 0 时该请求对证明不了任何关于保留思考的事(没有块可查),这在对每条对话的第一对请求以及会剥离 thinking 的 harness 上很常见;还有 ! 链式行(打印为 CHAIN-BREAK,计入退出状态),它是断点而非警告——当已发送的各轮回放回来时其 thinking 块发生了 API 拒绝的变化,它就会触发:保留的块必须是原序列的一个连续窗口(从前端丢弃、从尾部丢弃或两者皆可),因此从中间移除一个块或重排块,会让缺口之后的那个块失败,即使前缀其余部分原封未动。

The comparison ignores what the API ignores: cache_control markers, string content versus a single text block, leading and trailing whitespace of a text block, whitespace-only text blocks, key order, the order of tools in the array (they are compared as a name-keyed set), a defer_loading tool that no tool result, tool-search result, or tool_addition has named yet, request parameters outside system / tools / messages, everything before the last server-side compaction block (the check restarts there; the diff says when it compared from one), and the thinking blocks themselves. Interior whitespace, tool_use.input bytes, tool-result text, image bytes, and everything else count. Treat the script's pattern as a guess in the API's words - the API's own header in Step 2 is the authority when the two differ.

比较会忽略 API 所忽略的:cache_control 标记、字符串内容与单个文本块的差异、文本块的首尾空白、仅含空白的文本块、键顺序、数组内工具的顺序(它们作为以名称为键的集合比较)、尚未被任何工具结果、工具搜索结果或 tool_addition 提及的 defer_loading 工具、system / tools / messages 之外的请求参数、最后一个服务端压缩块之前的所有内容(检查从那里重新开始;差分会说明何时是从压缩块开始比较的),以及 thinking 块本身。内部空白、tool_use.input 字节、工具结果文本、图像字节及其他一切都会计入。把脚本的 pattern 当作以 API 词汇作出的猜测——两者不一致时,以第 2 步中 API 自己的响应头为准。

1.3 Read the code / 1.3 阅读代码

python3 shared/preserved-thinking-migration/prefix_diff.py --scan path/to/repo

The scan prints file:line leads grouped by cause - timestamps and environment reads inside prompt builders, slicing of the messages array, tool lists mutated after session start, the opening message rebuilt from state, tool results trimmed after the fact, reminder tags stripped with a regex, thinking blocks filtered out, round trips through the app's own message model, media caps and URL re-signing. They are regex leads, not findings: read each one, and confirm it with the pair diff or the API's response before it goes in the report. The scan is optional; the checklist below is not. An application can always express an edit in words the patterns do not know, so a scan with no leads is not proof of compliance, and a scan with leads is a reading list - the pair diff is the instrument.

扫描按成因分组打印 file:line 线索——提示词构造器内的时间戳与环境读取、messages 数组的切片、会话开始后被改动的工具列表、由状态重建的开场消息、事后裁剪的工具结果、用正则剥离的提醒标签、被过滤掉的 thinking 块、经由应用自身消息模型的往返、媒体上限与 URL 重签名。它们只是正则线索,不是结论:逐条阅读,先用请求对差分或 API 响应加以确认,再写进报告。扫描是可选的;下面的清单不是。应用总可以用模式不认识的方式表达一次编辑,因此没有线索的扫描不是合规的证明,有线索的扫描则是一份阅读清单——请求对差分才是测量仪器。

Whatever the scan finds, read these by hand - this checklist is the mandatory part of Step 1.3; they are where the edits hide:

无论扫描结果如何,以下内容都要人工阅读——这份清单是第 1.3 步的强制部分;编辑点就藏在这里:

1.4 Name each edit and decide whether it is deliberate / 1.4 给每个编辑命名并判断其是否有意为之

For every pair-diff verdict and every confirmed lead, record: the cause in the API's words (the pattern), the site in the code, which traffic class it is on, and whether the edit is deliberate (a compaction the product relies on; a user-invoked reset that starts a new conversation) or accidental (a timestamp nobody needed in the system prompt; a plugin landing inline on request 2). Accidental edits are removed outright in Step 3. Deliberate ones are replaced by their append-only form, or - where none exists yet - measured and decided (Step 2's caveats, Step 3's last section). The table under "Cause -> detection -> fix" in shared/preserved-thinking-migration/causes.md is the lookup for both: Read that file now if you have not yet, and check every candidate against its "Keep list" before it goes in the report.

对每个请求对差分判定和每条已确认的线索,记录:以 API 词汇表述的成因(pattern)、代码中的位置、所处的流量类别,以及该编辑是有意的(产品依赖的压缩;用户发起的、开启新对话的重置)还是意外的(系统提示词里没人需要的时间戳;第 2 个请求上内联落地的插件)。意外编辑在第 3 步直接移除。有意编辑则替换为它们的只追加形式,或者——在尚不存在该形式时——加以度量并决策(见第 2 步的注意事项与第 3 步最后一节)。shared/preserved-thinking-migration/causes.md 中 "Cause -> detection -> fix" 下的表格是两者的查询表:如果还没读过该文件,现在就读,并让每个候选成因先对照其 "Keep list" 核查,再写进报告。

Step 2: Measure with drop_block on a test slice / 第 2 步:在测试切片上用 drop_block 度量

The API's response is the only ground truth. The client-side diff can miss what it cannot see (media bytes behind a URL, an edit in a part of the request the capture didn't include) and can flag what the API tolerates; the response cannot.

API 的响应是唯一的事实基准。客户端差分会漏掉它看不见的东西(URL 背后的媒体字节、捕获未包含的请求部分中的编辑),也可能标记 API 所容忍的东西;响应不会。

2.1 The request shape / 2.1 请求形态

Every request in the test slice carries the beta header and sets the behavior explicitly - this is what turns the check on for an organization that is not enforced by default, and it is what adds the report to the response:

测试切片中的每个请求都携带该 beta 请求头并显式设置行为——这既是为默认未强制执行的组织打开该检查的方法,也是把报告附加到响应的方法:

POST /v1/messages
anthropic-beta: thinking-binding-controls-2026-08-01

{"model": "<the target model>", "max_tokens": 4096,
 "thinking": {"type": "adaptive", "block_binding": {"prefix_mismatch_behavior": "drop_block"}},
 "system": ..., "tools": [...],
 "messages": [ ...the full history with thinking blocks replayed verbatim... ]}

Rules that save a debugging hour:

能省下一小时调试的规则:

2.2 What the response tells you / 2.2 响应告诉你什么

Detect drops from the request diff. prefix_diff.py over consecutive requests (Step 1.2) is the detector: it reads only what your harness sent, so it works whatever shape the response takes. The surfaces below confirm and explain what it finds.

**从请求差分中检测丢块。**对相邻请求运行的 prefix_diff.py(第 1.2 步)是检测器:它只读你的 harness 发送了什么,因此无论响应是什么形态它都能工作。下面的几个界面用于确认并解释它发现的东西。

Three surfaces, in order of reliability:

三个界面,按可靠度排序:

  1. input_transformations (response body, top-level, sibling of usage) - the contract. With the header it is present on every response from a thinking-capable model: [] when nothing was dropped and nothing failed, otherwise one entry per block, of two types. {"type": "thinking_dropped", "path": "messages.7.content.0", "reason": "prefix_binding_mismatch"}: the block was removed. thinking_mismatch_allowed (same path, reason always prefix_binding_mismatch): the block failed the prefix check on a request the API does not enforce (an older account with the field unset), so it reached the model unchanged and was billed; every block after the edit gets one, and a request that sets the field never does (the probe always sets the field, so it reads such an entry as the field stripped en route: turn not evaluated, run inconclusive). path indexes the messages array as you sent it. reason is prefix_binding_mismatch (your history changed - this workflow's subject) or model_binding_mismatch (the conversation switched to a model that cannot read the block - not a bug in your code; see "Switching models mid-conversation" in shared/preserved-thinking-migration/causes.md); the probe flags any reason it does not classify as WARN. Dropped blocks are not billed, whichever the reason. Ignore entries whose type or reason you don't recognize; later checks add values. When streaming, the array arrives on the message object in message_start (and again in the final message_delta only after a mid-stream server-side model fallback). Without the header the field is absent and drops are silent.
    input_transformations(响应体顶层,与 usage 同级)——正式契约。带请求头时,它出现在每个来自支持 thinking 的模型的响应上:[] 表示无丢块且无失败,否则每个块一条条目,共两种类型。{"type": "thinking_dropped", "path": "messages.7.content.0", "reason": "prefix_binding_mismatch"}:该块已被移除。thinking_mismatch_allowed(同样的 path,reason 恒为 prefix_binding_mismatch):该块在一条 API 不强制执行的请求上(字段未设置的旧账号)未通过前缀检查,因此它原样到达了模型并被计费;编辑点之后的每个块都会有一条,而设置了该字段的请求绝不会出现(探针总是设置该字段,因此它把这种条目解读为字段在途中被剥离:该轮未评估,运行无结论)。path 按你发送时的 messages 数组索引。reason 为 prefix_binding_mismatch(你的历史变了——这正是本工作流的主题)或 model_binding_mismatch(对话切换到了无法读取该块的模型——不是你代码的 bug;见 shared/preserved-thinking-migration/causes.md 中的 "Switching models mid-conversation");探针会把任何它未归类的 reason 标为 WARN。无论原因如何,被丢弃的块都不计费。不认识其 type 或 reason 的条目先忽略;后续的检查会新增取值。流式传输时,该数组随 message_start 中的 message 对象到达(只有在流中途发生服务端模型回退时,最终的 message_delta 中才会再次出现)。不带请求头时该字段不存在,丢块是无声的。
  2. A diagnosis header, if present. Some responses that report a drop (or a 400 in error mode) also carry a response header named anthropic-thinking-prefix-mismatch - which can be missed on streamed responses (the probe's --stream mode may then print no "why:" line), so read it as a second check and detect from the request diff. Anthropic has not published this header; it may change or stop without notice. Use it if present; never depend on it. If it is there, its pattern and changed_validated fields are hints - the shape of the edit, in the same words prefix_diff.py uses, and the first changed path in the request - and the probe prints them on its "why:" line. Treat its absence as "no diagnosis", not "no break": the detail (kind, pattern, changed path) is only given for blocks your own organization created - for other blocks the header carries only the bare fact - and a partner cloud's proxy is not guaranteed to forward it.
    **诊断响应头(如果存在)。**一些报告丢块的响应(或 error 模式下的 400)还携带名为 anthropic-thinking-prefix-mismatch 的响应头——在流式响应上可能取不到(此时探针的 --stream 模式可能不打印 "why:" 行),因此把它当作二次核对,检测仍以请求差分为准。**Anthropic 尚未公开该响应头;它可能不经通知即变更或停用。存在就用,绝不依赖。**如果它在,其 pattern 与 changed_validated 字段是提示——以 prefix_diff.py 同样的词汇描述编辑形态,以及请求中首个被改的路径——探针会在其 "why:" 行打印它们。把它的缺失读作"没有诊断",而不是"没有断点":细节(kind、pattern、被改路径)只对你的组织自己创建的块给出——对其他块,该响应头只携带裸事实——而且伙伴云的代理不保证转发它。
  3. The 400 text in error mode - the same diagnosis as one sentence, for code that never sees headers: messages.7.content.0: Invalid signatureinthinkingblock. The block is bound to a different conversation. Remove the block, or setthinking.block_binding.prefix_mismatch_behaviorto "drop_block". Content that preceded this block when it was created is missing from this request, starting atmessages.2. It usually ends with one sentence naming the first changed path, as in that example (the sentence varies with the kind of edit and is sometimes absent). The request is rejected before any output; retrying the same body fails the same way (what production code does instead: "Failure modes to avoid" in causes.md).
    error 模式下的 400 正文——同样的诊断浓缩成一句话,供永远看不到响应头的代码使用:messages.7.content.0: Invalid signatureinthinkingblock. The block is bound to a different conversation. Remove the block, or setthinking.block_binding.prefix_mismatch_behaviorto "drop_block". Content that preceded this block when it was created is missing from this request, starting atmessages.2. 它通常以一句话收尾,指名首个被改的路径,如该例所示(这句话随编辑类型而变,有时缺失)。请求在任何输出之前即被拒绝;用同一请求体重试仍以同样方式失败(生产代码应当改用什么:causes.md 中的 "Failure modes to avoid")。

The token-counting endpoint runs the same conversation check. /v1/messages/count_tokens applies it to the replayed blocks: in error mode it returns the same 400 (with the diagnosis header when that is present); in drop_block mode it returns 200 and leaves the dropped block out of the count. A harness that counts tokens before each request meets the 400 there first. The count endpoint costs nothing and samples nothing, so it is the cheapest first-break pass: replay the slice against it in error mode (drop_block_probe.py --count-tokens --mode error) before a paid /v1/messages replay. It returns no input_transformations, so it answers "is anything broken, and where", not "how many blocks".

token 计数端点运行同样的对话检查。/v1/messages/count_tokens 会把该检查应用于重放的块:error 模式下返回同样的 400(若诊断响应头存在则一并返回);drop_block 模式下返回 200,并把被丢弃的块排除在计数之外。在每次请求前统计 token 的 harness 会最先在这里碰到 400。计数端点不花钱也不采样,因此是最便宜的首次断点排查:在付费的 /v1/messages 重放之前,先用 error 模式对它重放切片(drop_block_probe.py --count-tokens --mode error)。它不返回 input_transformations,因此它回答"是否坏了、坏在哪里",而不是"丢了多少块"。

None of the three says which line of your code made the edit. That is what Step 1's diff and scan are for: the header's pattern and changed_validated tell you where in the request to look; the pair diff tells you what changed there; the scan tells you who wrote it.

三者没有一个会说出是你代码中的哪一行做的编辑。这正是第 1 步差分与扫描的用途:响应头的 pattern 与 changed_validated 告诉你在请求中的哪里去找;请求对差分告诉你那里变了什么;扫描告诉你那是谁写的。

2.3 Run the slice / 2.3 运行切片

python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --dry-run                       # validates the capture, prints the plan, sends nothing
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --count-tokens --mode error  # free first pass on the token-counting endpoint: 400s mark the breaks
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --json probe.json         # drop_block replay; one conversation per file
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --mode error              # the loud arm (prefix_mismatch_behavior "error"), if you want the 400 text

Nothing is sent without --yes: a bare invocation prints the plan and stops. For CI, a non-zero exit is the signal - the probe exits 0 with no break, 1 when a conversation was inconclusive, 2 when at least one conversation broke (not severity-ordered); the diff script exits 1 on any mismatch or chain break.

没有 --yes 就不会发送任何内容:不带参数调用只打印计划并停止。对 CI 而言,非零退出就是信号——探针在无断点时退出 0,有对话无法得出结论时退出 1,至少一条对话断掉时退出 2(不按严重度排序);差分脚本在任何不匹配或链式断点时退出 1。

Replaying a request does not run the application's own tools - the capture already holds their results - but server-side tools in the request (web search, code execution, MCP connectors) run again on the API side if the model calls them, and every request is billed; the probe caps max_tokens at 16 by default because the verdict is decided before the first output token, which keeps the reply short and the cost to the input side. The probe sets tool_choice to none on every request that declares tools (tool_choice is outside the compared prefix), so no tool - server-side or the harness's own - can be called during a replay; the 16-token cap is for cost, not safety. Do not remove tools from the capture to the same end: the tool set is part of what the check compares and removing one would turn every request into a tool_set_changed break. This spends real money: estimate it from the slice (requests × their input size at current rates - fetch the rates, don't quote remembered ones) and get the user's approval once, for the whole measurement budget of this workflow, before the first run. --max-requests and --max-conversations cap a first run. Use a key from the same organization as the capture - a dedicated key or workspace in that organization is ideal; a capture replayed with another organization's key is not diagnosed, so the run teaches nothing about the harness.

重放一个请求不会运行应用自己的工具——捕获里已经带着它们的结果——但请求中的服务端工具(网络搜索、代码执行、MCP 连接器)在模型调用它们时会在 API 侧再次运行,且每个请求都计费;探针默认把 max_tokens 封顶在 16,因为判定在第一个输出 token 之前就已确定,这样回复短、开销都落在输入侧。探针把每个声明了工具的请求的 tool_choice 设为 none(tool_choice 在被比较前缀之外),因此重放期间没有任何工具——无论服务端还是 harness 自己的——能被调用;16-token 上限是为了成本,不是为了安全。不要为同样的目的从捕获中删除工具:工具集是检查所比较内容的一部分,删一个就会把每个请求变成 tool_set_changed 断点。这会花真金白银:按切片估算(请求数 × 各自输入大小按现行费率——费率要去查,不要凭记忆引用),并在首次运行前为整个工作流的度量预算取得用户的一次性批准。--max-requests 与 --max-conversations 可以给首次运行设上限。使用与捕获同一组织下的密钥——该组织下的专用密钥或工作区最理想;用别的组织的密钥重放捕获不会得到诊断,这样的运行对 harness 毫无教益。

What to log per request (the probe records all of it; production telemetry should too): the conversation id and turn index; the number of thinking blocks replayed in the request; every input_transformations entry; the anthropic-thinking-prefix-mismatch header if present; the HTTP status and, on a 400, the error text; the request id; and a client-side digest of the three bound parts - a hash of the canonical system, of the name-keyed tools, and of messages up to each replayed thinking block - so that the request where the digest changed can be found without the header.

每个请求要记录什么(探针全部记录;生产遥测也应如此):对话 id 与轮次索引;请求中重放的 thinking 块数量;每条 input_transformations 条目;anthropic-thinking-prefix-mismatch 响应头(若存在);HTTP 状态及 400 时的错误文本;请求 id;以及三个受约束部分的客户端摘要——规范化 system 的哈希、以名称为键的 tools 的哈希、以及到每个被重放 thinking 块为止的 messages 的哈希——这样即使没有响应头也能找到摘要发生变化的那次请求。

The counting rule. Count new dropped blocks per conversation, not entries per request. A block that fails once fails again on every later request that replays it, so a ten-turn conversation with one edit at turn 3 shows entries on eight responses but has one break. Identify a dropped block by the thinking block itself - resolve each entry's path to the block in the body you sent and key it by its signature (a redacted_thinking block by its data) - not by the path: paths shift whenever the history is truncated or compacted, and the same block would be counted again at its new index. The probe does this and keeps the path in the record. Its per-conversation summary gives first_break_turn, distinct_dropped_blocks (the turns of reasoning the conversation lost) and the patterns seen; the share of conversations with a first break is the baseline from Step 0, and the first-break turn is what you compare against the code: the request where the digest changed is the request after the edit.

**计数规则。**按对话统计新丢块数,而不是按请求统计条目数。一个失败过一次的块会在其后的每个重放它的请求上再次失败,因此一条在第 3 轮有一处编辑的十轮对话会在八个响应上出现条目,但只有一处断点。识别丢块要靠 thinking 块本身——把每条条目的 path 解析到你发送的请求体中的那个块,并以它的 signature 为键(redacted_thinking 块以其 data 为键)——不要按路径:历史一旦截断或压缩,路径就会移动,同一个块会在新索引处再被数一次。探针就是这么做的,并在记录中保留路径。它的按对话摘要给出 first_break_turn、distinct_dropped_blocks(该对话丢掉的推理轮数)以及见到的 pattern;出现首次断点的对话占比就是第 0 步的基线,而首次断点轮次是你与代码对照的对象:摘要发生变化的那次请求就是编辑之后的第一个请求。

Suspect the test slice before the harness. A slice that never replays a thinking block produces a perfect score and proves nothing - the probe warns when a conversation's maximum replayed-thinking count is zero. Check that number first, on every run. The first request of every conversation replays nothing by definition; an adaptive-thinking model may answer a short turn with no thinking block at all (Claude Fable 5.1 at its default effort often does, so a slice of short exchanges can carry no thinking anywhere); a harness that strips thinking before sending has nothing to check. The slice must also be captured on the target model: a block minted by a model without preserved thinking carries no conversation record for the check to verify, so replaying such a capture can never show a prefix break (the probe prints the capture's model ids - check them). Before replaying, count the thinking blocks in the captured assistant turns (the probe's --dry-run prints the number it will replay). If the production traffic genuinely carries little thinking, say so in the report - that is a finding about exposure, not a pass - and, for a capture made specifically to test the harness, drive the conversations with tasks that need reasoning or capture them with output_config.effort raised ("max" on Claude Fable 5.1) so that the assistant turns hold thinking; keep everything else as production sends it. A zero on a slice whose conversations replay thinking on most turns is the result you want; a zero on any other slice is a broken measurement.

**先怀疑测试切片,再怀疑 harness。**从不重放 thinking 块的切片会得到满分且证明不了任何事——当某条对话重放 thinking 的最大计数为零时,探针会告警。每次运行都先核对那个数字。每条对话的第一个请求按定义不重放任何东西;自适应思考模型可能对简短的一轮完全不给 thinking 块(Claude Fable 5.1 在默认努力度下常常如此,因此由短交换组成的切片可能通篇没有 thinking);发送前剥离 thinking 的 harness 无从检查。切片还必须在目标模型上采集:由不支持保留思考的模型铸造的块不携带可供检查验证的对话记录,因此重放这种捕获永远显示不出前缀断点(探针会打印捕获中的模型 id——请核对)。重放前,先数一数捕获的助手各轮中的 thinking 块(探针的 --dry-run 会打印将要重放的数量)。如果生产流量确实几乎不带 thinking,就在报告中如实说明——这是关于暴露面的发现,而不是通过——并且,对于专门为测试 harness 而采集的捕获,用需要推理的任务驱动对话,或在采集时调高 output_config.effort(Claude Fable 5.1 上为 "max"),使助手各轮带有 thinking;其余一切保持与生产发送一致。在一个多数轮次都重放 thinking 的切片上得到零,是你想要的结果;在任何其他切片上得到零,是坏掉的度量。

【评论】把"零断点"本身当作待核实的度量结果而非直接通过,是评测方法学中的负控制思路:先确认测量通道真的能测出问题,数字才有意义。

2.4 The three-arm protocol (validation against the eval) / 2.4 三臂协议(用评测集验证)

When the project has an eval, treat the migration as an A/B/C experiment on one frozen set of inputs. Arm 1 is the harness as it is today, check not enforced: the eval's existing score, no new run. Arm 2 is the same harness with drop_block set in the eval runner's configuration (never in production code), with input_transformations recorded per response and joined to each conversation's score and token usage; the join answers whether the conversations that lost reasoning scored worse or used more tokens, and by how much. Arm 3 is the harness after Step 3's fixes, again with drop_block set; the target is no prefix_binding_mismatch entries on the slice (model-check entries are counted separately: "Switching models mid-conversation" in causes.md) and a score within noise of Arm 1. Report the three scores, the token usage and the two drop counts side by side, per traffic class.

当项目有评测集时,把迁移当作在一组冻结输入上的 A/B/C 实验。臂 1 是现状 harness、检查未强制执行:用评测集现有分数,不新跑。臂 2 是同一 harness、在评测运行器的配置中设置 drop_block(绝不写进生产代码),逐响应记录 input_transformations 并与每条对话的分数和 token 用量关联;这一关联回答丢掉推理的对话是否分数更差或 token 更多、差多少。臂 3 是第 3 步修复后的 harness,同样设置 drop_block;目标是切片上不再有 prefix_binding_mismatch 条目(模型检查条目单独计数:causes.md 中的 "Switching models mid-conversation"),且分数与臂 1 的差异在噪声范围内。按流量类别并排报告三个分数、token 用量与两个丢块数。

2.5 Caveats that change what the measurement means / 2.5 会改变度量含义的注意事项

Step 3: Fix one cause per diff, re-measure, keep or revert / 第 3 步:每个差分修一个成因,重新度量,保留或回退

Work the causes in order of turns of reasoning lost: for each cause, sum distinct_dropped_blocks over the conversations whose first-break diagnosis carried that pattern (the probe's per-conversation summary gives both; a conversation with an early break loses more blocks than one that breaks late) - the cause that breaks every conversation at turn 2 comes before the one that breaks a tenth of them at turn 30.

按丢失推理的轮数排序处理各成因:对每个成因,把首次断点诊断为该 pattern 的各对话的 distinct_dropped_blocks 求和(探针的按对话摘要两者都给;断点早的对话比断点晚的丢块更多)——在第 2 轮断掉每条对话的成因,排在第 30 轮只断十分之一对话的成因之前。

When several causes hit the same request - the usual case in a harness that grew over time - the unit of work is the attribution line, not the pattern. Run prefix_diff.py on the first-break pair of each conversation; every line it prints (system[0] changed at char 78, messages[0] (user) content[0] (text) changed, messages[2] (user) content[1] (text) removed, ...) is one edit with one site in the code, and the API reports only the earliest of them (the header names the first failing block's cause; the kind says multiple). Rank the lines, fix each as its own diff, and measure each fix with the pair diff: the fix is right when that line disappears from the first-break pair. Use the probe for the end-to-end re-measure only after the whole set of lines on that pair is gone - the drop count cannot move while any edit on the first-break pair remains, so a probe run after a single correct fix will show the same breaks. Do not read that as "the fix did nothing" and revert it; read the pair diff.

当多个成因命中同一个请求时——随时间生长的 harness 中的常态——工作单元是归因行,而不是 pattern。对每条对话的首次断点请求对运行 prefix_diff.py;它打印的每一行(system[0] changed at char 78、messages[0] (user) content[0] (text) changed、messages[2] (user) content[1] (text) removed 等)都是一次编辑、对应代码中的一处,而 API 只报告其中最早的(响应头指名第一个失败块的成因;kind 为 multiple)。给这些行排序,每行作为独立的 diff 修复,并用请求对差分度量每次修复:当那一行从首次断点请求对中消失时,修复就是对的。只有该请求对上的全部行都消失之后,才用探针做端到端重新度量——只要首次断点请求对上还有任何编辑,丢块数就不可能动,因此单个正确修复之后的探针运行仍会显示同样的断点。不要把它读成"修复没起作用"然后回退;去读请求对差分。

Each cause that earns a place becomes its own diff (one cause per diff, so a revert is clean and the effect attributes). Diffs are proposed by default - presented to the user with the attribution line they clear and the measurement that will prove it - and applied only when the user asks; then measured: the pair diff first, then - once the first-break pair is clean - re-run the probe on the same slice, compare the share of conversations with a break and the first-break turn against the previous kept state, and, when the eval exists and the change is one of the two behavior-affecting kinds, re-run Arm 3. A diff whose attribution line goes to MATCH is kept; a diff that changes nothing in the pair diff is either a miss (the slice didn't exercise that path - extend the slice, not the claim) or a wrong diagnosis; a diff that clears its line but moves the eval is reverted and recorded. Never keep or revert on one conversation's swing; the slice is the unit.

每个入选的成因都成为它自己的 diff(一个成因一个 diff,回退才干净、效果才可归因)。diff 默认只提出——连同它将清除的归因行与将证明它的度量一起呈现给用户——仅在用户要求时应用;应用后度量:先看请求对差分,然后在首次断点请求对干净之后,在同一切片上重跑探针,把断点对话占比与首次断点轮次同上一个保留状态比较,并在评测集存在且改动属于两种影响行为的类型之一时重跑臂 3。归因行变为 MATCH 的 diff 保留;在请求对差分中毫无变化的 diff,要么打偏了(切片没经过那条路径——扩大切片,而不是扩大结论),要么诊断错了;清掉了自己的行却使评测变动的 diff 回退并记录。绝不因单条对话的波动而保留或回退;切片才是单元。

Accidental edits are removed. The recipes below are the append-only forms for the deliberate ones, the same shapes Anthropic's own agent products use, since they face every one of these problems (the "Cause -> detection -> fix" table in shared/preserved-thinking-migration/causes.md maps each pattern to its recipe here, and its model-switch section covers a harness that routes between models). They share one principle: the transcript is the source of truth, and everything the model needs to know later is added at the tail, never written into the head.

意外编辑直接移除。下面的配方是有意编辑的只追加形式,与 Anthropic 自己的代理产品所用的形态相同,因为它们面对其中每一个问题(shared/preserved-thinking-migration/causes.md 中的 "Cause -> detection -> fix" 表把每个 pattern 映射到此处的配方,其模型切换一节覆盖在模型间路由的 harness)。它们共享一条原则:转录文本是事实来源,模型此后需要知道的一切都追加在尾部,绝不写入头部。

Freeze the rendered prompt; deliver changes as appended messages. Render the system prompt once, at conversation start, and store the rendered bytes with the conversation record; every later request of that conversation sends the stored bytes - across process restarts, deploys, template updates, and client versions - not whatever this turn would render. Move every per-session or per-request fact out of system and out of the opening message: date and time, the user or account line, working directory, environment, instruction or memory files, feature flags, model and client version. Announce them once in the first user turn (or in a role: "system" message appended after it), and afterwards send only deltas, as a new appended message that says what changed ("Primary working directory: /repo/worktrees/x (was /repo)"; "Instruction files were re-read; these differ from their earlier copies: ..."). Mid-conversation role: "system" messages carry system-prompt authority and become part of the conversation record later blocks are checked against, so a change delivered this way is as strong as a re-rendered prompt and invalidates nothing; a plain one needs no beta header on Claude Fable 5.1 (only clear_at and the tool-change blocks below do). The only times the prompt is rendered again are deliberate boundaries - a new conversation, a user-invoked reset, the request after a full compaction - and a deliberate boundary is declared in the logs so a diff at that point is not mistaken for a bug.

**冻结渲染后的提示词;把变更作为追加消息下发。**在对话开始时渲染一次系统提示词,把渲染后的字节随对话记录存储;该对话此后的每个请求都发送存储的字节——跨越进程重启、部署、模板更新与客户端版本——而不是本轮会渲染出的任何内容。把每个按会话或按请求的事实移出 system 和开场消息:日期与时间、用户或账号信息行、工作目录、环境、指令或记忆文件、特性开关、模型与客户端版本。在第一条用户轮次中宣告一次(或在它之后追加的 role: "system" 消息中宣告),此后只发送增量,作为说明变化内容的新追加消息("Primary working directory: /repo/worktrees/x (was /repo)";"Instruction files were re-read; these differ from their earlier copies: ...")。对话中途的 role: "system" 消息携带系统提示词级别的权威,并成为其后各块受检所依据的对话记录的一部分,因此以这种方式下发的变更与重新渲染提示词同等有力,且不使任何东西失效;在 Claude Fable 5.1 上,普通的一条不需要 beta 请求头(只有 clear_at 和下文的工具变更块需要)。重新渲染提示词的时机只剩有意边界——新对话、用户发起的重置、完整压缩之后的第一个请求——并且要在日志中声明这一有意边界,使该点上的差分不被误认为 bug。

Declare the initial tool set at the start; never edit an entry; surface late tools and withdrawals by reference. For a small fixed set, build the full tools array before the first request, with defer_loading: true on tools that may not be ready; store the array as sent and replay it. For a large catalogue the model will mostly never use, leave not-yet-enabled tools out of tools and, on the turn one becomes available - or a tool connects later (an MCP server, a plugin, a permission granted mid-session) - append it to tools with defer_loading: true (a deferred tool nothing has referenced yet is outside the compared prefix, so appending it is safe; the Preserved thinking page documents this form) and announce it with a tool_addition block in an appended role: "system" message (beta mid-conversation-tool-changes-2026-07-01; it must follow a user message, such as the tool_result turn; right after a paused assistant turn that ends in a server-tool result a text-only system message is accepted but a tool change is a 400, so resume that turn first); never append a regular tool. A tool whose provider goes away is never removed from the tools array: to withdraw it, announce a tool_removal block in an appended role: "system" message (same beta) and leave the definition in place, returning an ordinary "not available" error if the model still calls it. Freeze each tool's description and schema text for the life of the conversation - a refreshed token, a date, a live listing, or a version string inside a description is a re-render.

**在开始时声明初始工具集;绝不编辑条目;后到的工具与撤下都按引用呈现。**对小型固定集合,在第一个请求之前构建完整 tools 数组,对可能尚未就绪的工具加 defer_loading: true;按发送原样存储该数组并重放。对模型多半永远用不到的大型目录,把尚未启用的工具留在 tools 之外,在某个工具可用的那一轮——或某个工具中途接入时(一台 MCP 服务器、一个插件、会话中途授予的权限)——用 defer_loading: true 把它追加进 tools(尚未被任何东西引用的延迟工具在被比较前缀之外,因此追加是安全的;Preserved thinking 页面记录了这一形式),并用追加的 role: "system" 消息中的 tool_addition 块宣告它(beta mid-conversation-tool-changes-2026-07-01;它必须跟在用户消息之后,例如 tool_result 轮;在以服务端工具结果结尾的暂停助手轮之后,纯文本系统消息可被接受,但工具变更会是 400,因此先恢复那一轮);绝不追加常规工具。提供方消失的工具绝不从 tools 数组移除:要撤下它,在追加的 role: "system" 消息中宣告一个 tool_removal 块(同一 beta),定义留在原地,模型若仍调用就返回普通的 "not available" 错误。把每个工具的描述与 schema 文本在对话存续期内冻结——描述里刷新的令牌、日期、实时列表或版本字符串都是一次重新渲染。

With the tool search tool in the request, keep a tool out of tools until it is available. Search can find and call a deferred tool (the Tool search page), so append the tool with defer_loading: true and a tool_addition block on the turn it becomes available.

**请求中带工具搜索工具时,工具在可用之前不放进 tools。**搜索能找到并调用延迟工具(Tool search 页面),因此在工具可用的那一轮,用 defer_loading: true 加一个 tool_addition 块追加它。

Keep per-turn reminders in the history. A reminder that should apply to one turn goes out as a turn-scoped system message - {"role": "system", "clear_at": "next_user_message", "content": "..."} appended after the tool_result message it applies to (beta mid-conversation-system-clear-at-2026-08-21; a role: "system" message must follow a user message, or an assistant message that ends in a server-tool result - anywhere else it is a 400, not a binding failure) - and every earlier copy stays where it is, byte for byte: a cleared message renders nothing, costs no input tokens, and is still part of the conversation record the thinking is checked against. Three details from the platform page: a turn-scoped message carries text content only; it takes no cache_control marker, so put the cache breakpoint on the user turn before it; and a user message that holds only tool_result blocks counts as the "next user message" that clears it. Without that beta, append the reminder as a text block after the tool_result blocks in the same user message and leave earlier copies in place; the model acts on the newest one. Rewording, rebuilding from current state, or deleting a copy already sent is an edit like any other. The same rule covers any mid-conversation role: "system" message: persist it with the transcript and replay it, including in sub-agent transcripts.

把每轮提醒留在历史里。只应作用于某一轮的提醒,以轮次作用域系统消息发送——{"role": "system", "clear_at": "next_user_message", "content": "..."} 追加在它所作用的 tool_result 消息之后(beta mid-conversation-system-clear-at-2026-08-21;role: "system" 消息必须跟在用户消息或以服务端工具结果结尾的助手消息之后——放在别处是 400,不是绑定失败)——并且更早的每一份副本原地不动,逐字节保留:被清除的消息不渲染任何内容、不耗输入 token,且仍是 thinking 受检所依据的对话记录的一部分。平台页面的三个细节:轮次作用域消息只携带 text 内容;它不接受 cache_control 标记,因此把缓存断点放在它之前的用户轮次上;只含 tool_result 块的用户消息也算作清除它的"下一条用户消息"。没有该 beta 时,把提醒作为 text 块追加在同一用户消息的 tool_result 块之后,更早的副本原地保留;模型按最新的一份行动。改写措辞、按当前状态重建、或删除已发送的副本,都是与其他无异的编辑。同一条规则覆盖任何对话中途的 role: "system" 消息:随转录持久化并重放,包括在子代理转录中。

Size old messages before they are sent, never after. Tool results, documents, and images are sized at ingestion - truncate the output, downscale the image, count the tokens - before the first request that carries them, and never touched again. Later trimming goes through server-side context editing (tool-result clearing, thinking clearing) or server-side compaction, which do not count as edits. For images specifically, the options in order of preference: (a) downscale at ingestion so that keeping every image in the history is affordable - the only option that loses nothing; (b) return generated or fetched images inside the tool_result of the tool that produced them, because server-side context editing can clear old tool results without a client edit, whereas there is no server-side way to prune an image that sits in a plain user turn; (c) if a client-side cap over user-turn images is unavoidable, make what it strips a deterministic function of the append-only history, stripping down to the cap minus a headroom so that a crossing happens once every N images rather than on every request - and say plainly that each crossing is still an edit that costs the reasoning after it; (d) the Files API (file_id) is for content whose bytes would otherwise drift between turns (a re-fetched URL, a re-encoded upload) - it does not reduce the tokens an image costs, so it is not a cap.

**在发送之前、绝不是之后给旧消息定尺寸。**工具结果、文档与图像在摄取时定型——截断输出、缩小图像、统计 token——在携带它们的第一个请求之前完成,此后绝不再碰。其后的裁剪经由服务端上下文编辑(工具结果清理、thinking 清理)或服务端压缩进行,它们不计为编辑。具体到图像,按优先级排序的选项:(a) 在摄取时缩小,让历史中保留每张图像都负担得起——唯一不损失任何东西的选项;(b) 把生成或获取的图像放在产生它们的工具的 tool_result 里返回,因为服务端上下文编辑无需客户端编辑即可清理旧工具结果,而躺在普通用户轮次中的图像没有服务端修剪手段;(c) 如果用户轮次图像上的客户端上限不可避免,让它剥离的内容成为只追加历史的确定性函数,剥到上限减去余量为止,使越线每 N 张图像发生一次而不是每个请求都发生——并明确说明每次越线仍是一次编辑、都会损失其后的推理;(d) Files API(file_id)用于那些字节否则会在轮次间漂移的内容(重新抓取的 URL、重新编码的上传)——它不减少图像消耗的 token,因此不是上限。

Store and echo wire bytes; never rebuild history from a domain model. Persist the messages array exactly as sent and the assistant content exactly as received - in particular tool_use.input as the API produced it (keep a normalized copy for your own execution if you need one, but echo the original) and text blocks untrimmed. Replay those bytes. A round trip through ORM objects, dataclasses, or a "normalize" pass is where interior whitespace, number formatting, key coercion, and string-versus-block shapes drift; the check tolerates leading and trailing whitespace and the string-versus-single-text-block shape, and nothing else.

**存储并回显线上字节;绝不从领域模型重建历史。**把 messages 数组按发送原样持久化、助手内容按接收原样持久化——尤其是 API 产出的 tool_use.input(自己执行时如需规范化副本可以留一份,但回显原件)以及未裁剪的文本块。重放那些字节。经过 ORM 对象、dataclass 或某个"规范化"步骤的往返,正是内部空白、数字格式、键类型强转、字符串与块形态发生漂移的地方;该检查容忍首尾空白与字符串对单个文本块的形态,其余一概不容。

Compact in a shape the check honours. Prefer server-side compaction (its instructions parameter takes your own summarization prompt; on Claude 5.1 and later models threshold compaction with custom instructions summarizes without the earlier thinking, while on-demand compaction's summarizer always reads it) or context editing - the checked prefix restarts at the compaction block. Client-side, the recommended shape is simple compaction: when the conversation grows too long, summarize it into one message and start the next request with that summary and the new user turn, replaying nothing older - no earlier turns, no earlier thinking. The summary is a plain user message, so there is no thinking left to fail the check. Threshold compaction writes its summary with the model named in the request; an on-demand compaction request can name another supported model, but kept turns' thinking stays valid only if every compaction request since it was produced ran on a model with preserved thinking. Never compact in the middle of a tool round (an assistant turn whose tool_use is still waiting on its tool_result goes back with its thinking intact). "Summarize the last N turns and drop the rest" is simple compaction too, as long as nothing older than the summary is replayed verbatim. If the product keeps a verbatim tail, strip the thinking and redacted_thinking blocks from the retained turns - text and tool calls stay - and make the strip a deterministic function of the stored compacted transcript (every assistant turn older than the compaction boundary goes out without thinking), so the same bytes go out on every later request and after a restart; a harness that re-derives the transcript from a store that still holds the thinking needs a recorded marker to get the same result. A rolling keep-last-N scheme pays this at every compaction (Step 2.5 compares the strip with drop_block). Name compaction in the logs as the one sanctioned boundary where the prefix legitimately changes.

**以检查认可的方式压缩。**优先服务端压缩(其 instructions 参数接受你自己的摘要提示词;在 Claude 5.1 及之后的模型上,带自定义 instructions 的阈值压缩在摘要时不带早前的 thinking,而按需压缩的摘要器始终读取它)或上下文编辑——受检前缀从压缩块重新开始。客户端侧,推荐形态是简单压缩:对话过长时,把它摘成一个消息,下一个请求以该摘要和新用户轮次开始,不重放任何更早内容——没有更早的轮次,没有更早的 thinking。摘要是普通的用户消息,因此没有留下会检查失败的 thinking。阈值压缩用它请求中所指名的模型写摘要;按需压缩请求可以指名另一个受支持的模型,但被保留轮次的 thinking 只有在它产生以来的每个压缩请求都运行在支持保留思考的模型上时才保持有效。绝不在工具轮中途压缩(tool_use 还在等 tool_result 的助手轮,其 thinking 完好地随行返回)。"摘要最近 N 轮、丢弃其余"也是简单压缩,只要不逐字重放比摘要更早的任何内容。如果产品保留逐字尾部,就从被保留轮次中剥离 thinking 与 redacted_thinking 块——文本与工具调用保留——并让这次剥离成为对已存储压缩转录的确定性函数(早于压缩边界的每个助手轮次不带 thinking 发出),这样每个后续请求以及重启之后发出的都是同样的字节;从仍存有 thinking 的存储中重新推导转录的 harness,需要一个有记录的标记才能得到同样结果。滚动的 keep-last-N 方案在每次压缩时都要付这笔代价(第 2.5 步比较了剥离与 drop_block)。在日志中把压缩称为前缀合法变化的唯一受认可边界。

Where the product can use one of two newer betas, its append-only form is one more scheme to measure: compact-2026-09-04 (on-demand compaction; not on Amazon Bedrock - its page's Compatibility list names the models and platforms) for background and keep-tail compaction, inline-tools-2026-09-15 (Claude API) for tools learned mid-session, same-name tool changes and connectors that re-list. Recipes and rules: "Append-only forms under newer betas" in shared/preserved-thinking-migration/causes.md; check the platform pages for availability first.

产品能用两个较新 beta 之一时,其只追加形式就是又一种待度量的方案:compact-2026-09-04(按需压缩;Amazon Bedrock 上没有——其页面的 Compatibility 列表列出了模型与平台)用于后台压缩与保留尾部压缩,inline-tools-2026-09-15(Claude API)用于会话中途学到的工具、同名工具变更与重新列出的连接器。配方与规则:shared/preserved-thinking-migration/causes.md 中的 "Append-only forms under newer betas";先查平台页面确认可用性。

Remove thinking only as a contiguous run, and record every strip. The check accepts any contiguous window of the original thinking blocks - a run dropped from the front (the oldest first; after a compaction block, the oldest after it), a run dropped from the back, or both - and nothing else: a block removed from the middle, or a reorder, fails the block after the gap and every one after that. Two shapes follow from it. When a 400 forces a strip-and-retry, strip from the rejected block onward (a trailing run) and persist that the strip happened, so later requests send the same stripped history, not the refused blocks. And never thin the middle. Once a block is removed, leave it out: putting it back invalidates the thinking produced while it was gone.

**只以连续区段移除 thinking,并记录每一次剥离。**该检查接受原 thinking 块序列的任何连续窗口——从头部丢弃的区段(最旧的在前;压缩块之后,则是压缩块之后最旧的)、从尾部丢弃的区段,或两者兼有——此外什么都不接受:从中间移除一个块或重排块,会让缺口之后的那个块及其后每一个失败。由此得到两种形态。当 400 迫使剥离重试时,从被拒块起向后剥离(一个尾部区段),并持久记录这次剥离,使后续请求发送同样的已剥离历史,而不是那些被拒的块。并且绝不抽稀中段。一个块一旦被移除,就让它保持移除:把它放回去,会使它缺席期间产生的 thinking 全部失效。

Keep conversations apart. A block from another conversation fails as kind=unrelated / pattern=foreign_prefix. That is a session-keying bug, not a prefix edit; fix the key. Reasoning cannot be carried into a new conversation: a branch that replays the history unchanged up to the fork keeps it; anything else starts from a summary.

**把对话彼此隔开。**来自另一条对话的块以 kind=unrelated / pattern=foreign_prefix 失败。那是会话键的 bug,不是前缀编辑;修键。推理无法带入新对话:在分叉点之前原样重放历史的分支可以保住它;其余一切都从摘要开始。

When no natural append-only form exists for a shape - keep-tail or background compaction without the compact-2026-09-04 beta, a same-name tool definition change or a connector that re-lists without the inline-tools-2026-09-15 beta (the rename-and-withdraw form above exists but is rarely worth it) - the diff is the decision, not a code change: measure the cost with Arm 2, choose "error" (a mismatch can only mean a bug, fail loudly) or "drop_block" (degrade, keep serving) for production, set it explicitly under the header, and log input_transformations or the 400s either way. Record the cause as measured and decided in the report, with the setting chosen and what it costs, so the next person doesn't re-litigate it.

当某种形态没有天然的只追加形式时——没有 compact-2026-09-04 beta 的保留尾部或后台压缩,没有 inline-tools-2026-09-15 beta 的同名工具定义变更或重新列出的连接器(上面的"改名并撤下"形式存在但很少值得)——这个 diff 就是决策本身,而不是代码改动:用臂 2 度量代价,为生产选择 "error"(不匹配只可能是 bug,要响亮地失败)或 "drop_block"(降级、继续服务),在请求头下显式设置,无论哪种都记录 input_transformations 或 400。把该成因以已度量并已决策记入报告,连同所选设置及其代价,免得下一个人重新争论一遍。

Minimal eval recipe, for the two behavior-affecting fixes when no eval exists: a frozen set of 20-30 real conversations from the capture; a per-conversation judgment that is cheapest for the workload (golden outputs to diff against, a short rubric, or an automated checker); a runner that replays one configuration and reports pass rate beside the probe's break share. Three arms, approved as one budget.

最小评测配方,供不存在评测集时用于两种影响行为的修复:从捕获中取 20-30 条真实对话的冻结集合;对该工作负载最便宜的逐对话判分(供差分的标准输出、一份简短评分细则、或一个自动化检查器);一个重放单一配置、把通过率与探针断点占比并排报告的运行器。三臂,作为一笔预算一起批准。

Finally, the production setting is its own diff, and the last one. Under the header, choose "error" or "drop_block" and set it explicitly; do not leave the field unset, because the defaults differ by surface and by account age, and an unset field on an account that is not yet enforced means the check is recorded, not applied. Apply that diff only with the user's explicit approval, only after the slice shows zero new drops for every cause that was fixed, and never "error" in production while a measured-but-unfixed cause remains - "error" belongs in CI, where one multi-turn capture per traffic class replays with it so that a new prefix edit fails the build. In production, alert on the first prefix_binding_mismatch entry per conversation, not on the count.

最后,生产设置本身是一个 diff,而且是最后一个。在请求头下选择 "error" 或 "drop_block" 并显式设置;不要把字段留空,因为默认值随入口和账号年龄而不同,且在尚未被强制的账号上留空意味着检查只被记录、不被执行。该 diff 只在用户明确批准后应用,且只在切片对每个已修复成因都显示零新增丢块之后,并且只要还有已度量但未修复的成因,生产中就绝不用 "error"——"error" 属于 CI:每个流量类别一条多轮捕获带着它重放,新的前缀编辑就会让构建失败。在生产中,对每条对话的首个 prefix_binding_mismatch 条目告警,而不是对计数告警。

Before the report, check the run against "Failure modes to avoid" in shared/preserved-thinking-migration/causes.md.

写报告之前,用 shared/preserved-thinking-migration/causes.md 中的 "Failure modes to avoid" 核对整个运行过程。

Step 4: Deliverables / 第 4 步:交付物

  1. The break profile and plan: the Step 0 assumptions (scope, traffic classes, platform and model, enforcement status, quality bar), the Step 2 baseline (share of conversations with a break, first-break turn distribution, replayed-thinking coverage of the slice), and the causes found - each with its pattern, its site in the code, whether it is deliberate, and the turns of reasoning it costs - ranked by reasoning lost. Causes with no append-only form are listed as measured and decided, with the setting chosen.
    断点画像与计划:第 0 步的假设(范围、流量类别、平台与模型、强制执行状态、质量标准),第 2 步的基线(断点对话占比、首次断点轮次分布、切片的 thinking 重放覆盖),以及找到的成因——每个都带其 pattern、代码中的位置、是否有意、以及它损失的推理轮数——按丢失推理量排序。没有只追加形式的成因列为已度量并已决策,连同所选设置。
  2. The changes: one diff per cause, in the order proposed (and applied, when the user asked for that), each tagged applied and measured (break share and first-break turn before and after; Arm 3 score where the eval exists), proposed (with the expected effect), or needs an eval (the two behavior-affecting kinds without one). Plus the production setting chosen ("error" or "drop_block") as its own, last, explicitly approved diff - applied only once the slice is clean for the fixed causes, and never "error" while an unfixed cause remains - where it is set, the CI replay, and the alert. "No changes recommended" - the slice replayed thinking on most turns and nothing was dropped - is a successful outcome; say it plainly.
    变更内容:每个成因一个 diff,按提出顺序排列(用户要求时也已应用),每个都标注 applied and measured(前后断点占比与首次断点轮次;评测集存在时还有臂 3 分数)、proposed(附预期效果)或 needs an eval(没有评测集的两种影响行为的类型)。外加所选的生产设置("error" 或 "drop_block"),作为独立的、最后一个、经明确批准的 diff——只在切片对已修复成因干净之后应用,且只要还有未修复成因就绝不用 "error"——设置在哪里、CI 重放与告警。"不建议变更"——切片多数轮次都重放了 thinking 而没有丢弃任何东西——也是成功的结果;直说即可。

Report skeleton (section order and required columns - keep the rest flexible):

报告骨架(小节顺序与必需的列——其余保持灵活):

Sources and live references / 资料来源与在线参考

Facts about the check, the request fields, and the response surfaces above are snapshots; the pages win where they differ. Fetch them when the user needs the full write-ups or current availability:

上文关于该检查、请求字段与响应界面的事实都是快照;两者不一致时以页面为准。当用户需要完整说明或最新可用性时,抓取这些页面: