← 提示词库 Anthropic/claude-code/skills/claude-api/shared/preserved-thinking-migration/causes.md 原文 md
🌐 中英双语对照

Preserved Thinking - Causes, Fixes, and the Keep List / Preserved Thinking——成因、修复与保留清单

Reference for shared/preserved-thinking-migration.md (the preserved-thinking-migration workflow). This file is not a workflow: it holds four lookups that the guide uses - the rules for a conversation that switches models, the "Cause -> detection -> fix" table, the keep list, and the failure modes to avoid - kept apart from the guide so that the guide fits in one Read. Read this file when the guide sends you here (Step 1.4, Step 2, or Step 3), and use the step numbers below as references into that guide.

shared/preserved-thinking-migration.md 的参考文件(即 preserved-thinking-migration 工作流)。本文件不是一个工作流:它保存着指南会用到的四项查询内容——对话切换模型时的规则、"Cause -> detection -> fix" 表、保留清单,以及需要避免的失败模式——将其与指南分开存放,以便指南能一次 Read 读完。当指南把你指引到这里时(Step 1.4、Step 2 或 Step 3)阅读本文件,并以文中的步骤编号作为回指该指南的参照。

Switching models mid-conversation / 在对话中途切换模型

A harness that routes one conversation to more than one model - a cheaper model for easy turns, a fallback when the primary is unavailable, an upgrade from the model the conversation started on - meets a second check that shares the response with the prefix check. The rules below are what the API does today, as observed on the Claude API with claude-opus-5, claude-opus-5-5, claude-fable-5, and claude-fable-5-1 in both prefix_mismatch_behavior modes; the Preserved thinking page (under "Sources and live references" in shared/preserved-thinking-migration.md) states the same rules, and the Extended thinking page (section "Only for the model that produced it, or a newer one") lists which models read which blocks.

把同一段对话路由到多个模型的 harness——简单轮次交给更便宜的模型、主模型不可用时回退、或从对话最初使用的模型升级——会遇到一个与 prefix 检查并列的第二项检查。以下规则是该 API 当前的行为,观测自 Claude API 上的 claude-opus-5、claude-opus-5-5、claude-fable-5 与 claude-fable-5-1,覆盖两种 prefix_mismatch_behavior 模式;Preserved thinking 页面(位于 shared/preserved-thinking-migration.md 的 "Sources and live references" 之下)陈述了相同的规则,Extended thinking 页面("Only for the model that produced it, or a newer one" 一节)列出了哪些模型可以读取哪些 block。

Is the model part of what the signature records? Yes. A thinking block's signature records the model that produced it alongside the conversation record and the previous thinking block. When the block is replayed, the API first asks whether the model now reading it reads blocks from the model that produced it (the page's rule: the model that produced it, or a newer one) - the model check - and only then whether the conversation before the block is unchanged - the prefix check. The model field of the request is not part of the conversation record, so changing it is not an edit: a model switch by itself never fails the prefix check.

模型是否属于 signature 所记录的内容? 是。thinking block 的 signature 在记录对话记录与上一个 thinking block 的同时,也记录了产生它的模型。当该 block 被重放时,API 首先询问当前读取它的模型能否读取由产生它的模型所产的 block(即该页面的规则:产生它的模型或更新的模型)——这是模型检查——然后才询问该 block 之前的对话是否未被改动——这是 prefix 检查。请求的 model 字段不属于对话记录,因此更改它不构成编辑:仅切换模型本身永远不会导致 prefix 检查失败。

Does a switch break the check? As of 2026-09-03 the API behaves as follows: a model switch does not fail a request - upgrade, downgrade, a round trip, and a downgrade combined with an edit all return 200, in "error" mode and in "drop_block" mode. "error" makes edits loud; it does not make a model switch loud, because the model check has no error setting: a block the current model cannot read is dropped, not rejected. Read the Preserved thinking page for the current rule. What differs by direction is whether the earlier reasoning is used:

切换会破坏检查吗? 截至 2026-09-03,API 的行为如下:模型切换不会使请求失败——升级、降级、往返,以及降级叠加编辑,在 "error" 模式与 "drop_block" 模式下均返回 200。"error" 让编辑变得响亮;它不会让模型切换变得响亮,因为模型检查没有报错设置:当前模型无法读取的 block 会被丢弃,而不是拒绝。当前规则请阅读 Preserved thinking 页面。随方向不同而不同的是:更早的推理是否被使用。

Does changing a tool description break the check? Yes. Tools are compared as their full definitions - name, description, and input_schema - so changing even one tool's description invalidates every thinking block minted before the change, reported as prefix_binding_mismatch with pattern=tool_schema_changed (the API reports a description edit and a schema edit with the same word). Adding or removing a plain tool has the same effect, reported as tool_set_changed. Reordering tools is fine, and a defer_loading: true tool is outside the comparison until something references it. The fix is the tool_schema_changed and tool_set_changed rows: freeze each tool's text for the life of the conversation, and add or withdraw tools by reference - or, under inline-tools-2026-09-15 (Claude API), append a tool_addition carrying the new definition ("Append-only forms under newer betas", below).

修改工具描述会破坏检查吗? 会。工具以其完整定义参与比较——名称、描述与 input_schema——因此哪怕只改一个工具的描述,也会使改动之前新铸的每一个 thinking block 失效,报告为 prefix_binding_mismatch 且 pattern=tool_schema_changed(API 用同一个词报告描述修改与 schema 修改)。新增或移除普通工具的效果相同,报告为 tool_set_changed。对工具重新排序没有影响,defer_loading: true 的工具在被引用之前不参与比较。修复方式见 tool_schema_changed 与 tool_set_changed 两行:在对话的整个生命周期内冻结每个工具的文本,通过引用来增撤工具——或者在 inline-tools-2026-09-15(Claude API)之下,追加一个携带新定义的 tool_addition(见下文"Append-only forms under newer betas")。

Cause -> detection -> fix. A model switch produces findings in two shapes, and only the second is a harness bug:

Cause -> detection -> fix(成因 -> 检测 -> 修复)。 模型切换会产生两种形态的发现,只有第二种是 harness 缺陷:

  1. The conversation is routed to a model that cannot read its thinking. Detection: model_binding_mismatch entries on the older model's turns; the probe prints them as model_drops=N per turn and model_drop_turns per conversation, apart from the prefix-break count, and prefix_diff.py prints model_switch=A->B on the pair where the model field changed (never as a MISMATCH). Fix: a routing decision, not a history edit. Send the same messages - thinking blocks included - to every model, and leave the beta header and block_binding field in place across the switch (the object is accepted on every model that accepts thinking). If the product needs the reasoning on those turns, keep the conversation on the model that produced it and switch models at conversation boundaries; if it needs the older model on those turns, accept that they run without the newer reasoning. Report it as reasoning lost by routing - conversations and turns affected - beside the prefix breaks, so the owner can decide.
    1. 对话被路由到无法读取其 thinking 的模型。 检测:较旧模型轮次上的 model_binding_mismatch 条目;探针按轮打印 model_drops=N、按对话打印 model_drop_turns,与 prefix 断裂计数分开;在 model 字段变化的那一对上,prefix_diff.py 打印 model_switch=A->B(绝不作为 MISMATCH)。修复:这是一个路由决策,不是历史编辑。把同一份 messages——包括 thinking block——发给每个模型,并在切换期间保持 beta header 与 block_binding 字段原位(每个接受 thinking 的模型都接受该对象)。如果产品需要这些轮次上的推理,就让对话留在产生它的模型上,只在对话边界切换模型;如果产品需要这些轮次由较旧模型承担,就接受它们在没有新推理的情况下运行。将其报告为因路由而丢失的推理——受影响的对话数与轮次数——与 prefix 断裂并列呈现,交由负责人决断。
  2. The switch triggers an edit. Three kinds, each an ordinary prefix break with the switch as its trigger: (a) the harness strips thinking blocks on a switch, or rebuilds the history from what each model was shown - the diff shows the blocks removed (predecessor_missing if from the middle), and the reasoning is gone for good when the conversation returns; fix: stop stripping, the API already leaves out what the current model cannot read. (b) A re-rendered system prompt or tool list that persists past the switch - a fallback banner that stays in system from then on, a prompt that depends on which model answered last - reaches Claude Fable 5.1 together with blocks minted under the old text: system_rerendered, tool_set_changed, or tool_schema_changed on the first Claude Fable 5.1 turn after it (on a downgrade the older model's turn in between reports only the model drop). A prompt or tool set that is a pure function of the model being called is different: every Claude Fable 5.1 request then carries the same text the blocks were minted under, so the check finds nothing (verified: the original prompt restored on the return to Claude Fable 5.1 gave 200 and [] with the block fed, in both modes), even though the pair diff flags the switch pairs and the older model's turns still lose the reasoning through the model check. Fix for the persisting case: the rows for those patterns - keep each model's prompt and tool text stable across that model's own turns, and deliver anything that must change mid-conversation as an appended role: "system" message. (c) A request shape the target model does not accept - a thinking configuration it rejects (a 400 from request validation whose text names the field, not a binding failure), or mid-conversation role: "system" messages, and with them tool_addition, tool_removal and clear_at, on a model the Mid-conversation system messages page does not list (it names Claude Sonnet 5 as unsupported: "Use the top-level system field there instead"); fix: one request body that every model in the route accepts.
    1. 切换触发了编辑。 三种情形,每一种都是一次普通的 prefix 断裂,切换只是其触发器:(a) harness 在切换时剥离 thinking block,或按各模型"所见"重建历史——diff 显示 block 被移除(若从中间移除则为 predecessor_missing),且当对话切回时推理已永久丢失;修复:停止剥离,API 本来就会略去当前模型读不了的内容。(b) 一次在切换之后持续存在的系统提示词或工具列表重渲染——一条从此驻留在 system 中的回退横幅、一条取决于上一个由哪个模型作答的提示词——与在旧文本下新铸的 block 一同抵达 Claude Fable 5.1:在其后的第一个 Claude Fable 5.1 轮次上出现 system_rerendered、tool_set_changed 或 tool_schema_changed(在降级场景下,中间那个较旧模型的轮次只报告模型丢弃)。作为"被调用模型的纯函数"的提示词或工具集则不同:此时每个 Claude Fable 5.1 请求携带的都是 block 新铸时所用的同一文本,检查一无所获(已验证:切回 Claude Fable 5.1 时恢复原始提示词,两种模式下均返回 200 和 [] 且 block 被正常喂入),尽管成对 diff 仍会标记切换对,且较旧模型的轮次仍经由模型检查丢失推理。针对持续存在情形的修复:见这些模式对应的行——让每个模型的提示词与工具文本在该模型自己的轮次间保持稳定,把任何必须在对话中途变化的内容作为追加的 role: "system" 消息送达。(c) 目标模型不接受某种请求形态——它拒绝的某个 thinking 配置(来自请求校验的 400,其文本点名字段名,而非绑定失败)、或对话中途的 role: "system" 消息,以及随之的 tool_addition、tool_removal 与 clear_at,被用在了 Mid-conversation system messages 页面未列出的模型上(该页面点名 Claude Sonnet 5 不支持:"Use the top-level system field there instead");修复:让路由中的每个模型都接受同一份请求体。

【评论】该文件把"模型路由导致的推理丢失"与"harness 缺陷"明确区分为两类发现,是典型的把对未公开 API 行为的观测结论固化为工程操作规程的内部文档写法。

Cause -> detection -> fix / 成因 -> 检测 -> 修复

The pattern column is the word the API's diagnosis header and the diff script both use. "Diff shows" is the attribution line from prefix_diff.py; "scan lead" is the heuristic --scan reports; the fix is the append-only form from Step 3 of shared/preserved-thinking-migration.md. One caution on reading the pattern word: for a summary-plus-tail compaction the header's word depends on how much was removed - the same code path reports tail_kept, compaction_summary, or unknown with kind=blocks_replaced on different conversations - so identify that cause by the kind (blocks_removed or blocks_replaced) together with the diff's attribution lines, not by one pattern word.

pattern 列是 API 的诊断 header 与 diff 脚本共同使用的词。"Diff shows" 是 prefix_diff.py 给出的归因行;"scan lead" 是 --scan 报告的启发式线索;修复方式即 shared/preserved-thinking-migration.md Step 3 的 append-only 形式。阅读 pattern 词时须注意一点:对于"摘要+尾部"式压缩,header 给出的词取决于移除了多少——同一代码路径在不同对话上会分别报告 tail_kept、compaction_summary 或 unknown 且 kind=blocks_replaced——因此要通过 kind(blocks_removed 或 blocks_replaced)结合 diff 的归因行来判定该成因,而不是靠单个 pattern 词。

Tier pattern (and kind) What the harness did Diff shows Scan lead Fix
1 system_rerendered (system_changed) The system prompt was rebuilt with per-request content: time, cwd, account line, memory or instruction files, flags, version strings system[i] changed at char N time and environment reads near prompt builders; templates rendered per request Render once, store the bytes with the conversation, replay; per-session facts go into the first turn; changes go out as appended role: "system" messages
2 tool_set_changed (tools_changed) A tool was added or removed after the first request: a plugin or MCP server connected late, a provider disconnected, a permission changed tools: X added / removed tools mutated after session start; a tool listing fetched per request Declare the full set at start; append a late tool with defer_loading: true (safe while unreferenced) and announce it with tool_addition in an appended system message - never append a regular tool; never remove one from the array - withdraw it with a tool_removal block and leave the definition in place, returning an ordinary "not available" error if the model still calls it
2 tool_schema_changed (tools_changed) Same tool names, different description or schema text: a date, a refreshed token, a live listing, a version inside a description tools: X description changed at char N descriptions or schemas built from templates or state Freeze each tool's text for the conversation; store and replay the definitions as sent. Without the inline-tools-2026-09-15 beta no append-only form expresses a same-name change - a new name is the only way to offer changed text; under it, append a tool_addition carrying the new definition instead (the betas section below)
1-2 system_and_tools_changed (multiple) Both re-rendered, messages untouched: a connector landing on request 2, or a restart or resume re-deriving both both of the above startup, resume, reconnect paths Replay the stored prompt and tool text across restarts, and across model switches (a switch is not a boundary: a re-render that persists into later requests on the model that minted the blocks is this break, with the switch as its trigger; a prompt that is a pure function of the model called is stable on each model's own turns and is not); the only declared boundaries are a new conversation, a user-invoked reset, and the request after a full compaction
1 first_message_rewritten (blocks_modified / blocks_removed, often with system in sections) The opening user message carried context rebuilt from live state: environment, instructions, a session date messages[0] (user) content[j] changed at char N messages[0] assigned after creation; a context header rendered per request Announce context once and freeze it; send later changes as an appended message describing the delta
3 rolling_truncation (blocks_removed) The oldest turns were dropped whole - a sliding window messages[0..k] removed messages[-N:], keep-last, window size Without the compact-2026-09-04 beta no client-side form keeps the thinking (with it, on-demand compaction does, but the window must summarize rather than only drop - the betas section below). The choices: server-side compaction or context editing; simple compaction (summary plus new turn, nothing older); or keep the window and strip the retained turns' thinking as a deterministic, recorded strip, or send drop_block (equivalent in effect) - measured
3 tail_kept (blocks_removed) A run of older turns removed (or replaced by a summary the record can't see) with the first message kept and the newest turns verbatim - keep-tail compaction or keep-first truncation messages[i..j] removed, messages[0] intact summarize(messages[:-k]) plus messages[-k:] Same as above; the retained turns' thinking cannot verify without the compact-2026-09-04 beta (it does behind an on-demand compaction block, the betas section below) - send drop_block from the compaction onward or strip that thinking as a recorded decision, never mid tool-round; measure, decide, and record the decision
3 compaction_summary (blocks_replaced) Older turns replaced in place by a shorter summary, the tail intact messages[i..j] replaced by 1 message(s) same Same; or move the summary to simple compaction (replay nothing older than the summary); or, under compact-2026-09-04, on-demand compaction (the betas section below)
4 tool_results_rewritten (blocks_modified) Old tool_result content trimmed or cleared after it was sent messages[i] (user) content[j] (tool_result) changed at char N tool-result truncation applied to earlier turns Bound outputs before the first send; later clearing through server-side context editing (clear_tool_uses_20250919, beta context-management-2025-06-27); a client-side prune only at a declared boundary, as a pure function of the growing history
1 tool_use_rewritten (blocks_modified) Old tool_use.input re-encoded or normalized on replay content[j] (tool_use) changed input normalizers, to_dict on tool calls Echo tool_use.input exactly as received; normalize a copy for execution only
1 reserialized (blocks_modified) Many blocks differ slightly across types: a lossy round trip through the app's own message model (interior whitespace, number formatting, coerced keys, trimmed text) many changed at char N lines across messages from_dict/to_dict, JSON re-encoding of history, .strip() on content Persist and replay the wire JSON; never rebuild messages from domain objects
2 reminder_stripped / history_block_stripped / block_inserted (blocks_removed / blocks_inserted) A per-turn text block injected into a user turn and removed on the next request (or added after the fact) messages[i] (user) content[j] (text) removed / inserted regex strips of reminder tags; inject-then-strip helpers Turn-scoped system message (clear_at: "next_user_message") appended after the tool results, every earlier copy left in place; without the beta, a text block after the tool_result blocks, left in place
2 system_block_rerendered / system_blocks_stripped A mid-conversation role: "system" message re-rendered in place, or several dropped (a sub-agent transcript replayed without them) messages[i] (system) changed / removed transcript stores that don't keep system messages Persist them with the transcript and replay verbatim
3 media_stripped (blocks_removed / blocks_modified) Images or documents in earlier turns dropped, downsized, or replaced by a placeholder - a client media cap content[j] (image) removed image caps, resizing of stored turns Downscale at ingestion; return images inside the producing tool's tool_result so server-side context editing can clear them; if a user-turn cap is unavoidable, strip deterministically to cap-minus-headroom and accept that each crossing is an edit; file_id only for bytes that would drift
1 image_url_resigned (pattern) / media_content_changed (kind) A URL-sourced image or document whose block changed, or whose bytes differ from the first fetch content[j] (image) changed URL re-signing, re-uploads The check compares the bytes, not the URL string: a rotated URL to the same bytes is fine; for content referenced across turns use a file_id or base64
4 predecessor_missing / predecessor_reordered (kinds of the chain check; the header reads kind=predecessor_missing; pattern=not_applicable) A thinking block removed from the middle, or re-ordered, with the rest of the prefix intact ! messages[i] re-sent with a different set of thinking blocks filters on type == "thinking"; serializers that drop empty fields or unknown block types (a thinking block with empty text is still a block); a hand-rolled stream parser that loses the signature_delta; strip-and-retry without a record Keep the replayed thinking blocks a contiguous window of the original (drop from the front or the back, never the middle); make any forced strip deterministic and recorded so it replays identically
3 unknown with kind blocks_removed or blocks_replaced A shortening of the history that the API does not name more specifically, or several edits at once messages[i..j] removed / replaced by 1 message(s) the same leads as the truncation and compaction rows Read the attribution lines; the fix is the truncation or compaction one above
- foreign_prefix (unrelated) A block from another conversation replayed (on a long conversation; a short one reports an ordinary multiple / system_and_tools_changed) no pair diff (it is a different conversation) session keys, multiplexed stores Fix the session keying
- (no pattern) A drop or 400 on a pair where nothing you sent differs - the diff shows no change and the digests match no diff - Not a harness bug; report the request id to Anthropic
4 thinking_modified (a separate 400, "cannot be modified") A replayed thinking block's text differs from what the API returned - truncated, summarized, re-wrapped ! the thinking text of the block in messages[i] differs stores that trim or reformat thinking text Store and replay thinking blocks byte for byte
0 model_binding_mismatch (model check; no pattern, no header) The conversation was routed to a model that cannot read its earlier thinking - a downgrade, a cheaper-model route, a fallback model_switch=A->B on the pair, verdict unchanged; the probe's model_drops model ids chosen per turn; fallback or router code Not a history edit: send the same messages to every model, keep the field and header in place, and decide the routing - pin the conversation to the producing model, or accept that the older model's turns run without the newer reasoning; report it as reasoning lost by routing
1-2 system_rerendered / tool_set_changed / tool_schema_changed / predecessor_missing, triggered by a switch A re-rendered prompt or tool list that persists past the switch (a fallback banner, a prompt keyed to the last responder), or thinking stripped on the switch; on a downgrade the edit is reported only on the next Claude Fable 5.1 turn. A prompt that is a pure function of the model called is not this break the row for that pattern, on the pair after the switch (the diff also flags a per-model prompt that the API accepts - confirm with the probe) prompts or tool lists that change with the route; thinking filtered on a switch The row for that pattern: each model's prompt and tool text stable across its own turns, changes as an appended role: "system" message, and never strip thinking on a switch
层级 pattern(及 kind) harness 做了什么 Diff 显示 扫描线索 修复
1 system_rerendered(system_changed) 系统提示词被按请求内容重建:时间、cwd、账户信息行、记忆或指令文件、标志位、版本字符串 system[i] changed at char N 提示词构建器附近读取时间与环境;模板按请求渲染 只渲染一次,把字节随对话存储并重放;会话级事实放进第一轮;变更以追加的 role: "system" 消息发出
2 tool_set_changed(tools_changed) 首个请求之后新增或移除了工具:插件或 MCP 服务器晚接入、提供方断开、权限变化 tools: X added / removed 会话开始后 tools 被改动;每次请求都拉取工具列表 开局声明完整集合;晚到的工具以 defer_loading: true 追加(未被引用时安全),并用追加系统消息中的 tool_addition 予以宣告——绝不追加常规工具;绝不从数组中移除工具——用 tool_removal block 将其撤下并保留定义,若模型仍调用则返回普通的 "not available" 错误
2 tool_schema_changed(tools_changed) 工具名相同,描述或 schema 文本不同:日期、刷新的 token、实时清单、描述中的版本号 tools: X description changed at char N 由模板或状态构建的描述或 schema 在整个对话中冻结每个工具的文本;按发送原样存储并重放定义。没有 inline-tools-2026-09-15 beta 时,没有任何 append-only 形式能表达同名变更——改用新名称是提供变更文本的唯一途径;在该 beta 下,改为追加携带新定义的 tool_addition(见下文 betas 一节)
1-2 system_and_tools_changed(multiple) 两者都被重渲染,消息未动:连接器落在第 2 个请求上,或重启/恢复时两者都被重新推导 以上两者 启动、恢复、重连路径 跨重启、跨模型切换重放存储的提示词与工具文本(切换不是边界:若重渲染持续存在于产生 block 的模型的后续请求中,即是此种断裂,切换只是其触发器;作为被调用模型纯函数的提示词在该模型自己的轮次上是稳定的,不属于此类);唯一声明的边界是新对话、用户主动重置,以及完整压缩之后的请求
1 first_message_rewritten(blocks_modified / blocks_removed,常伴有 sections 中的 system) 开场用户消息携带了从实时状态重建的上下文:环境、指令、会话日期 messages[0] (user) content[j] changed at char N messages[0] 在创建之后才赋值;上下文头按请求渲染 上下文只宣告一次并冻结;后续变化作为描述差异(delta)的追加消息发送
3 rolling_truncation(blocks_removed) 最旧的轮次被整轮丢弃——滑动窗口 messages[0..k] removed messages[-N:]、保留末尾、窗口大小 没有 compact-2026-09-04 beta 时,没有任何客户端侧形式能保住 thinking(有了它,按需压缩可以,但窗口必须做摘要而不只是丢弃——见下文 betas 一节)。可选方案:服务端压缩或上下文编辑;简单压缩(摘要加新一轮,不留更早内容);或保留窗口,把被保留轮次的 thinking 作为确定性的、留有记录的剥离移除,或发送 drop_block(效果等同)——均须度量
3 tail_kept(blocks_removed) 移除了一段较早的轮次(或替换为记录不可见的摘要),首条消息保留、最新轮次逐字保留——保留尾部的压缩或保留首部的截断 messages[i..j] removed,messages[0] 完好 summarize(messages[:-k]) 加 messages[-k:] 同上;没有 compact-2026-09-04 beta 时,被保留轮次的 thinking 无法通过校验(在按需压缩 block 之下则可以,见下文 betas 一节)——从压缩那一点起发送 drop_block,或将该 thinking 作为留有记录的决策剥离,绝不在工具轮中途进行;度量、决策并记录该决策
3 compaction_summary(blocks_replaced) 较早的轮次被原地替换为更短的摘要,尾部完好 messages[i..j] replaced by 1 message(s) 同上 同上;或把摘要改为简单压缩(摘要之前的内容一律不重放);或在 compact-2026-09-04 下使用按需压缩(见下文 betas 一节)
4 tool_results_rewritten(blocks_modified) 已发送的旧 tool_result 内容事后被裁剪或清空 messages[i] (user) content[j] (tool_result) changed at char N 对更早轮次施加的工具结果截断 首次发送前就限制输出规模;事后清理通过服务端上下文编辑完成(clear_tool_uses_20250919,beta context-management-2025-06-27);客户端侧修剪只在声明的边界上、作为随历史增长的纯函数进行
1 tool_use_rewritten(blocks_modified) 旧的 tool_use.input 在重放时被重新编码或归一化 content[j] (tool_use) changed 输入归一化器、对工具调用使用 to_dict tool_use.input 按收到时的原样回显;仅为执行而归一化一个副本
1 reserialized(blocks_modified) 大量 block 在各类型上略有差异:经由应用自身消息模型的有损往返(内部空白、数字格式、键被强转、文本被裁剪) 大量跨消息的 changed at char N 行 from_dict/to_dict、对历史做 JSON 重编码、对内容调用 .strip() 持久化并重放线路上的 JSON;绝不从领域对象重建消息
2 reminder_stripped / history_block_stripped / block_inserted(blocks_removed / blocks_inserted) 注入到用户轮次的按轮文本 block 在下一个请求中被移除(或事后被加入) messages[i] (user) content[j] (text) removed / inserted 用正则剥离提醒标签;先注入后剥离的辅助函数 轮次作用域的系统消息(clear_at: "next_user_message")追加在工具结果之后,此前每一份都原位保留;没有该 beta 时,在 tool_result block 之后放一个文本 block 并原位保留
2 system_block_rerendered / system_blocks_stripped 对话中途的 role: "system" 消息被原地重渲染,或几条被丢弃(子代理转写被重放时未带它们) messages[i] (system) changed / removed 不保存系统消息的转写存储 随转写持久化并逐字重放
3 media_stripped(blocks_removed / blocks_modified) 较早轮次中的图像或文档被丢弃、缩小或替换为占位符——客户端媒体上限 content[j] (image) removed 图像上限、对已存储轮次做缩放 在摄取时就缩小尺寸;让图像随产生它的工具的 tool_result 返回,以便服务端上下文编辑能清除它们;如果用户轮上限不可避免,则确定性地裁剪到"上限减去余量"并接受每次越线都是一次编辑;file_id 只用于本就会漂移的字节
1 image_url_resigned(pattern)/ media_content_changed(kind) URL 来源的图像或文档其 block 发生变化,或其字节与首次抓取不同 content[j] (image) changed URL 重新签名、重新上传 检查比较的是字节而非 URL 字符串:指向相同字节的轮换 URL 没有问题;跨轮引用的内容使用 file_id 或 base64
4 predecessor_missing / predecessor_reordered(链检查的若干 kind;header 显示 kind=predecessor_missing; pattern=not_applicable) 一个 thinking block 从中间被移除或被重新排序,其余 prefix 完好 ! messages[i] re-sent with a different set of thinking blocks 针对 type == "thinking" 的过滤器;丢弃空字段或未知 block 类型的序列化器(文本为空的 thinking block 仍是 block);丢失 signature_delta 的手写流解析器;不留记录的"剥离后重试" 让重放的 thinking block 保持为原序列的连续窗口(从前部或后部丢弃,绝不从中间丢弃);任何被迫的剥离都要确定性且有记录,以便重放完全一致
3 unknown,kind 为 blocks_removed 或 blocks_replaced API 未更具体命名的对话缩短,或同时发生多处编辑 messages[i..j] removed / replaced by 1 message(s) 与截断、压缩各行相同的线索 阅读归因行;修复方式即上文截断或压缩对应行
- foreign_prefix(unrelated) 重放了来自另一段对话的 block(长对话上;短对话报告普通的 multiple / system_and_tools_changed) 无成对 diff(这是另一段对话) 会话键、多路复用的存储 修好会话键
- (无 pattern) 在你发送的内容毫无差异的一对请求上出现丢弃或 400——diff 无变化、摘要一致 无 diff - 不是 harness 缺陷;将请求 id 报告给 Anthropic
4 thinking_modified(单独的 400,"cannot be modified") 重放的 thinking block 文本与 API 返回的不同——被截断、摘要、重新折行 ! the thinking text of the block in messages[i] differs 会修剪或重排 thinking 文本的存储 逐字节存储并重放 thinking block
0 model_binding_mismatch(模型检查;无 pattern、无 header) 对话被路由到无法读取其更早 thinking 的模型——降级、更便宜模型的路由、回退 该对上 model_switch=A->B,判定不变;探针的 model_drops 按轮选择模型 id;回退或路由代码 不是历史编辑:向每个模型发送相同的 messages,保持字段与 header 原位,并对路由做出决策——把对话钉在产生推理的模型上,或接受较旧模型的轮次在没有新推理的情况下运行;将其报告为因路由而丢失的推理
1-2 system_rerendered / tool_set_changed / tool_schema_changed / predecessor_missing,由切换触发 重渲染后持续存在过切换的提示词或工具列表(回退横幅、按最后作答者键控的提示词),或切换时剥离了 thinking;在降级场景下,该编辑只在下一个 Claude Fable 5.1 轮次上被报告。作为被调用模型纯函数的提示词不属于此种断裂 该模式对应的行,出现在切换后的那一对上(diff 也会标记 API 所接受的按模型提示词——用探针确认) 随路由变化的提示词或工具列表;切换时对 thinking 的过滤 该模式对应行:每个模型的提示词与工具文本在其自己的轮次间保持稳定,变更以追加的 role: "system" 消息送达,且绝不在切换时剥离 thinking

Append-only forms under newer betas / 较新 beta 下的 append-only 形式

Two newer betas add an append-only form for shapes the table above marks as having none without them. On-demand compaction (compact-2026-09-04) is on the Claude API, Claude Platform on AWS, Google Cloud and Microsoft Foundry, not on Amazon Bedrock; the Compatibility list on its page names the models and platforms, and the Models API reports each model's capabilities.compaction with the beta header. Defining a tool inside a message (inline-tools-2026-09-15) is Claude API only; by-reference tool changes under the older mid-conversation-tool-changes-2026-07-01 header also work on Amazon Bedrock and Google Cloud. Where a beta is not available, treat those shapes as the table says: measure and decide, freeze and replay.

两个较新的 beta 为上表标注"没有它们就没有对应形式"的形态补充了 append-only 形式。按需压缩(compact-2026-09-04)可用于 Claude API、Claude Platform on AWS、Google Cloud 与 Microsoft Foundry,不用于 Amazon Bedrock;其页面上的 Compatibility 列表给出了模型与平台,Models API 会在 beta header 下报告每个模型的 capabilities.compaction。在消息内定义工具(inline-tools-2026-09-15)仅限 Claude API;在较旧的 mid-conversation-tool-changes-2026-07-01 header 下按引用变更工具的方式在 Amazon Bedrock 与 Google Cloud 上同样可用。在 beta 不可用之处,按表格所述对待这些形态:度量并决策,冻结并重放。

Background and keep-tail compaction, the append-only way: on-demand compaction (beta compact-2026-09-04, "compaction": {"type": "summarize"}; the on-demand compaction page, https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand). Send a compaction request that carries exactly the messages of a request you already sent, on the conversation's model and under its system, tools and thinking settings, with the compaction field and a max_tokens large enough for a summary; send the beta header on it and on every request that carries the block. Exactly those messages, because the kept turns must directly follow the summarized messages and the first kept message must not be one the API would merge into the last summarized one (the same role, or a role: "system" message). A compaction request whose last assistant turn is still waiting on a tool result is rejected, so send the results first. Leave out output_config.format, stop_sequences and a tool_choice of type any or tool (the API rejects a compaction request that carries them), and never send output_config.task_budget.remaining on the compaction request or on any request that carries the block (a 400). The response holds one signed compaction block and nothing else (stop_reason: "compaction"). When no summary could be written there is no block: the response is still a 200 with empty content, and stop_reason is the summarization call's own - max_tokens (cut off), model_context_window_exceeded (no room for the summarization prompt), refusal, tool_use, or end_turn (no text) - so give max_tokens a few thousand tokens at least, resend with more room or fewer messages as the reason suggests, or continue without one. Custom instructions replace the server's summarization prompt whole (a blank value counts as absent), so ask for text only and no tool call yourself; the summarizer reads the whole conversation either way, earlier thinking included (unlike threshold compaction with custom instructions on Claude 5.1 and later models, which leaves earlier thinking out). Keep taking turns against the full history while it runs, and do not edit anything already sent.

后台压缩与保留尾部压缩的 append-only 做法:按需压缩(beta compact-2026-09-04,"compaction": {"type": "summarize"};按需压缩页面,https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand)。 发送一个压缩请求,其 messages 与你已发送过的某个请求完全一致,使用对话所在的模型并沿用其 system、tools 与 thinking 设置,附上 compaction 字段和一个足以容纳摘要的 max_tokens;在该请求以及其后每个携带该 block 的请求上都要发送 beta header。必须恰好是那些消息,因为被保留的轮次必须紧跟在被摘要的消息之后,且第一条被保留的消息不能是 API 会并入最后一条被摘要消息的那条(角色相同,或是 role: "system" 消息)。最后一个 assistant 轮次仍在等待工具结果的压缩请求会被拒绝,所以要先送出结果。不要带 output_config.format、stop_sequences 以及 any 或 tool 类型的 tool_choice(API 拒绝携带它们的压缩请求),也不要在压缩请求或任何携带该 block 的请求上发送 output_config.task_budget.remaining(会返回 400)。响应只包含一个已签名的 compaction block,别无其他(stop_reason: "compaction")。当写不出摘要时则没有 block:响应仍是 200 但 content 为空,stop_reason 为摘要调用自身的取值——max_tokens(被截断)、model_context_window_exceeded(放不下摘要提示词)、refusal、tool_use 或 end_turn(无文本)——因此至少给 max_tokens 留出几千 token,按原因提示用更大空间或更少消息重发,或干脆放弃摘要继续。自定义 instructions 会整体替换服务端的摘要提示词(空值视同缺省),因此请自行要求只输出文本、不调用工具;无论哪种方式,摘要器都会读取整段对话,包括更早的 thinking(不同于 Claude 5.1 及以后模型上带自定义指令的阈值压缩,后者会把更早的 thinking 排除在外)。压缩运行期间可继续在完整历史上进行轮次,且不要改动任何已发送的内容。

On the first request after the block arrives, drop exactly the messages you sent to the compaction request from the front of the history and put the block first, as an assistant message of its own (the request that does this adopts the block; the API also accepts it as the first content block of the first kept message, and prefix_diff.py checks the kept turns only in the message-of-its-own form); everything appended since stays, thinking included, and keeps verifying. Keep nothing else from the dropped messages: summarized messages re-sent after the block are not rejected - the model sees them twice, summary then verbatim - so drop them yourself. Keep the block first on every later request; to compact again, send compaction on a request that starts with the current block, and from then on send only the newest block (a request that carries more than one compaction block is a 400). Do not send compaction and context_management in the same request. Text instructions in role: "system" messages and tool_addition / tool_removal blocks that sat inside the summarized messages stop applying at adoption (tool changes excepted when the block carries tool_changes - next sentence): re-declare them in a role: "system" message on the first request that adopts the block, directly after that request's new user turn (which comes after the kept turns - a system message between the block and the kept turns breaks their thinking), and leave it there afterward. Where the compaction request also carried inline-tools-2026-09-15, the block records the summarized messages' net tool changes in a tool_changes field: send the block back unmodified and they carry over by themselves, so re-declare no tool change; a block without that field carries none, so re-declare as above. The block is accepted on any model that supports the beta, with any later system or tools, but the kept turns' thinking verifies only on a model that can read it, only if every compaction request since that thinking was produced ran on a model with preserved thinking, and only while system and the tools other than defer_loading: true ones stay what the compaction request had: changing either can invalidate the kept turns' thinking and has no other effect, so to change them without losing any, compact the whole conversation first (keeping no turns), then change them on the next request. Nothing before the block is sent, but the kept turns are still checked against the summarized messages as they stood when you sent the compaction request, so do not touch them in between; prefix_diff.py compares them against the earlier request by aligning on a kept thinking block both carry, and when the compaction request itself is in the capture it checks that request's system and tools against the conversation's, compares the adopting request's system and tools against that request, and checks that the dropped range is the messages it carried.

在该 block 到达后的第一个请求上,把你发给压缩请求的那些消息从历史前部精确移除,并把该 block 放在首位,作为一条独立的 assistant 消息(执行此动作的请求即采纳该 block;API 也接受将其作为第一条被保留消息的第一个 content block,而 prefix_diff.py 只在"独立消息"形式下检查被保留的轮次);此后追加的一切原样保留,thinking 也包括在内,并继续通过校验。被移除消息的其余内容一件都不要保留:block 之后再重发被摘要的消息不会被拒绝——模型会看到它们两次,先摘要后原文——所以要由你自己丢弃它们。此后的每个请求都让该 block 保持首位;要再次压缩,就在一个以当前 block 开头的请求上发送 compaction,从那之后只发送最新的 block(携带多于一个 compaction block 的请求是 400)。不要在同一请求中同时发送 compaction 与 context_management。位于被摘要消息内部的 role: "system" 文本消息与 tool_addition / tool_removal block,在采纳时停止生效(block 携带 tool_changes 时工具变更除外——见下一句):在采纳该 block 的第一个请求上、紧跟该请求新的 user 轮之后(它位于被保留轮次之后——block 与被保留轮次之间插入系统消息会破坏其 thinking),用一条 role: "system" 消息重新声明它们,此后保持原位。若压缩请求同时带有 inline-tools-2026-09-15,该 block 会把被摘要消息的净工具变化记录在 tool_changes 字段中:把该 block 原样发回,它们会自行延续,因此不要重新声明任何工具变更;不带该字段的 block 即无工具变化,按上文重新声明即可。该 block 在任何支持该 beta 的模型上都被接受,可配任意后续的 system 或 tools,但被保留轮次的 thinking 只有在能读取它的模型上、且在该 thinking 产生之后的每个压缩请求都运行于支持 preserved thinking 的模型上、并且 system 与除 defer_loading: true 之外的工具保持与压缩请求相同时才通过校验:改动任一项都可能使被保留轮次的 thinking 失效,且没有其他作用,因此要在不丢失任何内容的前提下变更它们,先把整段对话压缩一遍(不保留任何轮次),再在下一个请求上变更。block 之前的内容不会被发送,但被保留的轮次仍会对照你发送压缩请求那一刻被摘要消息的形态进行检查,所以在此期间不要动它们;prefix_diff.py 通过双方共有的某个被保留 thinking block 对齐来将它们与更早的请求比较,而当压缩请求本身也在抓取中时,它会将该请求的 system 和 tools 与对话的进行比较、将采纳请求的 system 和 tools 与该请求的进行比较,并检查被移除的范围正是它所携带的那些消息。

Tool changes by value (beta inline-tools-2026-09-15, Claude API; the Mid-conversation system messages and tool changes page, "Define tools in a message"). Keep tools exactly as the first request sent it, on every request, and make every later change by appending one role: "system" message: a tool_addition whose tool is {"type": "tool_definition", "definition": {...}} with the full tools entry (name, description, input_schema) inside definition, for a new tool or for a same-name tool whose description or schema changed - a different definition replaces the tool from that message on, an identical one changes nothing (safe to resend on a retry) - and a tool_removal by reference to withdraw one. Nothing already sent moves, so earlier thinking keeps verifying; rewording or deleting a definition message already sent is an edit like any other. The header also covers changes by reference, so it replaces mid-conversation-tool-changes-2026-07-01; the placement rules are the same, including no tool change directly after a paused assistant turn. A tool added this way may itself be defer_loading: true (inside definition); a tool already known at the first request belongs in tools with defer_loading: true, shown later by reference. Keep at least one non-deferred tool in tools (a tool search tool counts): otherwise the first tool defined by value costs one full cache miss. cache_control goes on the block or in the definition, not both, and never on a deferred definition. Reusing a name for a different type of tool is a 400, and during the beta some tool types (computer use among them) cannot be defined in a message - declare those in tools and add them by reference. A definition stays in the history after the tool is replaced or withdrawn, so a beta header that a dated tool type needs goes on every later request of the conversation. For an MCP connector server, add mcp-client-2026-09-15 (in place of mcp-client-2025-11-20, which it includes): the definition can then be an mcp_toolset (connection details stay in mcp_servers), and a response for which the API fetched a server's tool list starts with one mcp_tool_listing block per server fetched (code that reads content[0] must skip them) - send the assistant message back as it came, those blocks included, keep the header on every request that carries one, and later requests reuse that list instead of asking the server again.

按值变更工具(beta inline-tools-2026-09-15,Claude API;Mid-conversation system messages and tool changes 页面,"Define tools in a message")。 在每个请求上都让 tools 与第一个请求所发送的完全一致,此后的每一次变更都通过追加一条 role: "system" 消息完成:tool_addition,其 tool 为 {"type": "tool_definition", "definition": {...}},把完整的 tools 条目(名称、描述、input_schema)放进 definition,用于新工具,或用于描述或 schema 发生变化的同名工具——不同的定义从该消息起替换该工具,相同的定义不改变任何东西(重试时重发是安全的)——以及按引用撤下工具的 tool_removal。已发送的内容一概不动,因此更早的 thinking 继续通过校验;改写或删除已发送的定义消息与其他任何编辑无异。该 header 同时覆盖按引用的变更,因此它取代了 mid-conversation-tool-changes-2026-07-01;放置规则相同,包括不能紧跟暂停的 assistant 轮次做工具变更。以此方式添加的工具自身可以是 defer_loading: true(在 definition 内部);首个请求时已知的工具应放入 tools 并带 defer_loading: true,之后再按引用展示。tools 中至少保留一个非延迟工具(工具搜索工具也算):否则第一个按值定义的工具要付出整整一次缓存未命中。cache_control 要么放在 block 上、要么放在定义里,不能两者都放,也绝不要放在延迟定义上。把同一个名称复用于不同类型的工具是 400,并且 beta 期间某些工具类型(计算机使用在其中)不能在消息中定义——在 tools 中声明它们,并按引用添加。定义在工具被替换或撤下之后仍留在历史中,因此带日期的工具 type 所需的 beta header 要挂在该对话此后每个请求上。对于 MCP 连接器服务器,添加 mcp-client-2026-09-15(取代 mcp-client-2025-11-20,前者包含后者):此时 definition 可以是 mcp_toolset(连接细节留在 mcp_servers 中),而 API 为其抓取了某服务器工具列表的响应会以每个被抓取服务器一个 mcp_tool_listing block 开头(读取 content[0] 的代码必须跳过它们)——把 assistant 消息按原样发回,包括这些 block,在每个携带它们的请求上保持该 header,后续请求会复用该列表而不再询问服务器。

Keep list - what never to flag / 保留清单——哪些情形绝不该被标记

The scan and the diff will tempt you to report things the check does not care about. These stay out of the report (or go in a "checked, fine" line):

扫描与 diff 会诱使你报告一些检查并不在意的东西。以下内容不得进入报告(或只放入一行 "checked, fine"):

【评论】"Keep list" 以白名单方式列出不构成 prefix 编辑的差异(参数变更、URL 重签、键序等),其目的是把自动化扫描的误报压到最低,与前文的诊断表互为补充。

Failure modes to avoid / 需要避免的失败模式