fix: v0.6.2 修复工具调用不稳定与会话停止 — 纯 tool_calls 轮丢失 assistant 消息导致 API 400
CI / 类型检查 + Lint + 单元测试 (push) Failing after 5m43s
CI / 全量测试 (Electron ABI) (push) Failing after 5m20s
CI / 产物编译验证 (push) Successful in 10m5s

【根因(main.log 实证)】
19:04 / 19:05 / 19:06 三次会话终止均为同一报错:
  DeepSeek 400 "Messages with role 'tool' must be a response to a preceding
  message with 'tool_calls'"

缺陷链:engine 主循环仅在 step.thought 存在(该轮有文本或思考内容)时才
将 assistant 消息加入请求历史。当模型发起纯工具调用(零文本零思考 —
DeepSeek 高频行为)时:
  - assistant(tool_calls) 消息不进 messages
  - 但 tool 结果消息照常 push
  → 下一轮请求出现孤立 tool 消息 → 协议 400(不可重试)→ 会话 ERROR 终止
"不稳定" = 模型每轮是否附带文本是概率性行为:带文本正常,纯调用必崩。
DB 持久化侧同源缺陷(if (!step.thought) continue)导致这些步骤的
assistant 与 tool 结果全部不落库 — 重启后工具上下文丢失,模型重复调用。

【修复】
- engine.ts: 有 toolCalls 的轮次必 push assistant(content=null,C-6 规范)
- agent.ts: 持久化条件同步修复(无 thought 但有 toolCalls 的步骤落库)
- 回归测试: 纯 tool_calls 轮后第二次请求中 tool 消息前必须是带
  tool_calls 的 assistant(请求契约断言,engine-toolchain.test.ts)

【纵深防御 — 孤立 tool 消息过滤】
- openai-format.ts(DeepSeek/Agnes/MiMo/OpenAI 四家共享): 构建请求时
  按 tool_call_id 配对过滤孤立 tool 消息(任何来源的历史污染不再 400 死锁)
- anthropic.adapter.ts: tool_use/tool_result 同策略配对过滤
- 单测 ×6: 正常配对保留 / 孤立丢弃 / id 不匹配丢弃 / 多轮配对 /
  includeImages 原位转换 / 非 vision 静默丢弃

【多模态索引对齐收敛】
4 家 adapter 的 images 处理循环原按未过滤的 nonSystemMsgs[i-1] 对齐索引,
孤立 tool 过滤引入后会错位 — 统一收进 buildOpenAICompatibleMessages
(includeImages 参数,基于 sanitized 序列原位转换),4 家 adapter 删除
各自的索引对齐循环(DeepSeek vision 判断 / OpenAI 推理模型拒绝保留在 adapter)。

【终止原因可见化】
MAX_ITERATIONS / TIMEOUT 终止此前无任何提示(用户感知"会话直接停止")—
前端 DONE 事件非 completed 终止原因显示为 system 消息。

【v0.6.1 回归缓解】
web_fetch timeoutMs 120s → 240s:浏览器回退串行化后并发 3 个排队最坏
~127.5s,旧值让排队末位抓取被工具超时杀掉(表现为抓取不稳定)。

【验证】
lint 0/0;typecheck 双工程 0 错误;test:electron 259/259(+7);
electron-vite build 成功
This commit is contained in:
2026-08-22 19:34:16 +08:00
parent 80cf5b482c
commit a7090214b1
14 changed files with 524 additions and 214 deletions
@@ -10,7 +10,8 @@
* @see apis/mimo-api-docs-20260715.html
*/
import type { MetonaRequest, MetonaToolDef } from '../../types';
import log from 'electron-log';
import type { MetonaMessage, MetonaRequest, MetonaToolDef } from '../../types';
/**
* 构建 OpenAI 兼容的 messages 数组
@@ -19,13 +20,15 @@ import type { MetonaRequest, MetonaToolDef } from '../../types';
* - System Prompt 拼接(静态区 + 动态区 + 安全准则)
* - 工具调用历史保留(reasoning_content + tool_calls
* - 工具结果注入(tool_call_id + content
* - 孤立 tool 消息过滤(纵深防御,见函数内注释)
* - 多模态图片(includeImages=true 时转换为 image_url content parts
*
* 注意:图片(多模态)处理不属于此共享函数。
* 各 Provider 对多模态的支持不同(DeepSeek 不支持,Agnes/Ollama 支持但格式各异),
* 应在各自 Adapter 的 toNativeRequest 中处理。
* @param includeImages true 时将消息的 images 转为 OpenAI image_url content parts
* Agnes/MiMo/OpenAI 全系、DeepSeek 仅 vision 模型传 true
*/
export function buildOpenAICompatibleMessages(
request: MetonaRequest,
includeImages = false,
): Array<Record<string, unknown>> {
const systemContent = [
request.systemPrompt.roleDefinition,
@@ -36,60 +39,97 @@ export function buildOpenAICompatibleMessages(
.filter(Boolean)
.join('\n\n');
const nonSystemMessages = request.messages
.filter((m) => m.role !== 'system')
.map((m) => {
// v0.3.0 修复: assistant 消息有 tool_calls 但 content 为空时,content 设为 null
// DeepSeek/OpenAI API 要求有 tool_calls 的 assistant 消息 content 必须为 null 而非空字符串
const msg: Record<string, unknown> = {
role: m.role,
content: m.content,
};
// 注意:图片(多模态)处理不在此共享函数中。
// DeepSeek 非 vision 模型不支持多模态,images 被静默丢弃是正确行为
// vision 模型在 DeepSeekAdapter.toNativeRequest 中独立处理)。
// Agnes/MiMo/OpenAI 各自的 toNativeRequest 中有独立的 images 处理。
// 审查修复: #27 曾在此添加 images 处理,但 DeepSeek 非 vision 模型会导致 API 400,已撤销。
// === Assistant 消息 ===
if (m.role === 'assistant') {
// 工具调用历史
if (m.toolCalls?.length) {
msg.tool_calls = m.toolCalls.map((tc) => ({
id: tc.id,
type: 'function',
function: { name: tc.name, arguments: JSON.stringify(tc.args) },
}));
// 有 tool_calls 时 content 必须为 nullAPI 规范)
if (!m.content) msg.content = null;
}
// 推理内容(无论是否有工具调用,都保留 reasoning_content
if (m.reasoningContent) {
msg.reasoning_content = m.reasoningContent;
}
// 崩溃修复(纵深防御): 过滤孤立 tool 消息 — 其 tool_call_id 不属于任何前置
// assistant(tool_calls) 消息。OpenAI/DeepSeek 协议要求 tool 消息必须紧跟带
// tool_calls 的 assistant,违反直接 400 且不可重试(会话死锁)。正常链路由
// engine 保证配对;此处兜底任何来源的历史污染(旧版本数据/导入/边界场景)。
const nonSystem = request.messages.filter((m) => m.role !== 'system');
const sanitized: MetonaMessage[] = [];
/** 已出现且尚未被 tool 结果回应的 tool_call id 集合 */
const pendingToolCallIds = new Set<string>();
let droppedOrphans = 0;
for (const m of nonSystem) {
if (m.role === 'assistant' && m.toolCalls?.length) {
for (const tc of m.toolCalls) pendingToolCallIds.add(tc.id);
sanitized.push(m);
continue;
}
if (m.role === 'tool' && m.toolResult) {
if (pendingToolCallIds.has(m.toolResult.toolCallId)) {
pendingToolCallIds.delete(m.toolResult.toolCallId);
sanitized.push(m);
} else {
droppedOrphans++;
}
continue;
}
sanitized.push(m);
}
if (droppedOrphans > 0) {
log.warn(
`[OpenAIFormat] Dropped ${droppedOrphans} orphan tool message(s) without matching assistant tool_calls`,
);
}
// === 工具执行结果 ===
if (m.role === 'tool' && m.toolResult) {
msg.tool_call_id = m.toolResult.toolCallId;
// CE-2 修复: 工具失败时 result 为 null,优先用 error 字段作为 content
// 否则 LLM 看到 "null" 不知道失败原因,可能重复调用导致死循环
msg.content = m.toolResult.error
? m.toolResult.error
: typeof m.toolResult.result === 'string'
? m.toolResult.result
: JSON.stringify(m.toolResult.result);
// #26 修复: 确保 tool 消息 content 不为 undefined
// JSON.stringify(undefined) 返回 undefined(非字符串),会导致 content 字段在序列化后消失
// OpenAI/DeepSeek/Agnes API 严格要求 tool 消息必须有 content 字段,缺失会返回 400
if (msg.content === undefined) msg.content = '';
const messages = sanitized.map((m) => {
// v0.3.0 修复: assistant 消息有 tool_calls 但 content 为空时,content 设为 null
// DeepSeek/OpenAI API 要求有 tool_calls 的 assistant 消息 content 必须为 null 而非空字符串
const msg: Record<string, unknown> = {
role: m.role,
content: m.content,
};
// === 多模态图片(includeImages=true 时) ===
// v0.6.2 收敛: 原先 4 家 adapter 各自按 nonSystemMsgs[i-1] 索引对齐处理 images
// 孤立 tool 过滤引入后索引错位。统一收进共享函数,基于 sanitized 原位转换。
// DeepSeek 非 vision 模型传 falseimages 静默丢弃是正确行为,见 #27 审查撤销记录)。
if (includeImages && m.images?.length) {
const contentParts: Array<Record<string, unknown>> = [];
if (m.content) contentParts.push({ type: 'text', text: m.content });
for (const img of m.images) {
contentParts.push({ type: 'image_url', image_url: { url: img.url } });
}
msg.content = contentParts;
}
return msg;
});
// === Assistant 消息 ===
if (m.role === 'assistant') {
// 工具调用历史
if (m.toolCalls?.length) {
msg.tool_calls = m.toolCalls.map((tc) => ({
id: tc.id,
type: 'function',
function: { name: tc.name, arguments: JSON.stringify(tc.args) },
}));
// 有 tool_calls 时 content 必须为 nullAPI 规范)
if (!m.content) msg.content = null;
}
// 推理内容(无论是否有工具调用,都保留 reasoning_content
if (m.reasoningContent) {
msg.reasoning_content = m.reasoningContent;
}
}
return [{ role: 'system', content: systemContent }, ...nonSystemMessages];
// === 工具执行结果 ===
if (m.role === 'tool' && m.toolResult) {
msg.tool_call_id = m.toolResult.toolCallId;
// CE-2 修复: 工具失败时 result 为 null,优先用 error 字段作为 content
// 否则 LLM 看到 "null" 不知道失败原因,可能重复调用导致死循环
msg.content = m.toolResult.error
? m.toolResult.error
: typeof m.toolResult.result === 'string'
? m.toolResult.result
: JSON.stringify(m.toolResult.result);
// #26 修复: 确保 tool 消息 content 不为 undefined
// JSON.stringify(undefined) 返回 undefined(非字符串),会导致 content 字段在序列化后消失
// OpenAI/DeepSeek/Agnes API 严格要求 tool 消息必须有 content 字段,缺失会返回 400
if (msg.content === undefined) msg.content = '';
}
return msg;
});
return [{ role: 'system', content: systemContent }, ...messages];
}
/**