P0 会话可靠性收口(根治"模型思考着会话就停止"): - P0-1 finish_reason 全链路贯通:DONE 事件与 IterationStep 新增 finishReason,OpenAI 共享 SSE / Anthropic message_delta.stop_reason / Ollama done_reason 三路采集,TRACE 层弃用硬编码 'stop' 记录真值 - P0-2 空响应守卫 + 降级重试:零产出流→可重试错误走退避;思考耗尽输出预算(reasoning-only + length)→自动关闭思考降级重试一次;仍失败→OUTPUT_LENGTH_EXCEEDED 结构化错误 + 故障转移;附带根治 abort 恰逢零工具调用轮被 COMPLETED 抢占的真实缺陷 - P0-3 思考×能力×预算三对齐:DeepSeek/MiMo/Agnes/Ollama 四家 supportsThinking=false 强制不发思考参数;小输出预算告警;设置页联动提示 - P0-4 渲染层可见性:截断/空完成/友好错误三类提示,i18n 全部出层 - P0-5 回归四件套:reasoning-only 终止判定、集成级空闲超时、504 引擎重试归类、思考中 abort→USER_INTERRUPT、P4-2 强制收尾路径 FEAT-1:LLM 设置新增「最大输出上限」——Provider 支持矩阵显隐 + 模型上限钳制提示 + 超限保存警告 + llm.maxTokens 热生效 P1 修复面收口: - 渲染层三缺陷根治:后台会话回放缓冲(2000 条/4MB 有界 + agent:getReplayState + 事件总线)+ abort 双层自愈 + sendMessage 收尾兜底 + 中断卡片清扫 - 工具 abort 信号全覆盖:web_search/web_fetch/http_request/code_search/git 系列/delegate_task 全部接入引擎中断;web_search 时间预算收敛(720s→≤240s);移除伪造 ToolExecutionContext 与死代码 - 安全:本地 Pinned CONNECT 代理根治浏览器通道 DNS rebinding(校验期 IP pinning,可注入 resolver 表测);配置 URL 域名解析深校验(DeepCheckSoftFailure 软失败);SSE 空 error 帧防御修复;Ollama generate/embed AbortSignal.any 合并 - 缺陷清单:UTF-16 BOM 读取、tmp 同毫秒碰撞(nanoid 后缀)、code_search JS 回退参数对称(case_sensitive/前后文独立)、list_directory include_node_modules、崩溃自愈退避(60s 窗 ≥3 次停 reload)、MemoryViewer/Sidebar i18n 收口 P2 能力演进: - 会话回收站:SCHEMA_VERSION 3 + 迁移 10(deleted_at,存在性守卫),软删除/恢复/彻底删除/30 天自动清理(启动+24h),searchMessages 聚合剔除,Sidebar 回收站面板 - 会话回放播放器:sessions:listRecordings/readRecording(白名单+目录边界+20MB 上限),SessionReplayPlayer 时间轴/步进/变速,Trace 面板入口 - electron-updater 自动更新:双轨(手动 feed 比对保留),生产环境启动静默检查 + update:status 广播 + app:updateInstall + LogsSettings UpdatePanel + builder publish 配置 - @ 文件提及:workspace.listFiles/readFileClip(边界/512KB/NUL 拒绝/MEMORY.md 保护),ChatInput Fuse 联想+键盘导航+附件管线注入 - MCP Resources/Prompts 发现:可选能力 try/catch 降级,mcp:listServerContents,MCPSettings 展开视图 - 文档对齐:内部 API 标准 HTML(Adapter 清单补 MiMo/已实现注记/STREAM_RESET/DONE.finishReason/ repetition_truncation 映射);README v0.8.0 亮点表 P3 测试基建: - 新增 4 个测试文件:engine-stream-contract(6)、engine-stream-reliability(4:集成空闲超时/504 重试/思考中 abort/P4-2 强制收尾)、thinking-capability-gate(7)、pinned-proxy(9,含深校验 5)、session-trash(5,DB 域)、use-agent-stream hook 级(5)、agent.test 回放缓冲(2) - 契约更新:orchestrator 被中断 SubAgent success=false(abort 优先级修复语义)、SSE 空 error 帧、UTF-16 正常读取、DeepSeek 未配置思考显式 disabled、迁移矩阵 v2→3 - 弱断言根治:registry WEBP 单向断言、hooks-contracts 自比恒真、memory 空 token 补强 全量验证:typecheck 0 错误 / lint 0 问题 / 系统 Node 2144 通过(301 DB 用例按 ABI 跳过)/ Electron ABI 2445/2445 全量通过 0 跳过
827 lines
32 KiB
TypeScript
827 lines
32 KiB
TypeScript
/**
|
||
* Ollama Provider Adapter
|
||
*
|
||
* 本地推理引擎,支持 Tool Calling、Thinking 模式、NDJSON 流式。
|
||
* 无需 API Key,连接本地 http://localhost:11434。
|
||
*
|
||
* 完整实现 Ollama API 文档中所有端点和参数:
|
||
* - POST /api/chat (对话)
|
||
* - POST /api/generate (补全)
|
||
* - POST /api/embed (嵌入)
|
||
* - GET /api/tags (列出模型)
|
||
* - POST /api/show (模型详情)
|
||
* - POST /api/pull (下载模型)
|
||
* - GET /api/ps (运行中模型)
|
||
* - GET /api/version (版本)
|
||
* - think 参数(Thinking 模式)
|
||
* - options 参数(temperature/top_k/top_p/stop/num_ctx/num_predict)
|
||
* - format 参数(structured output)
|
||
* - images 参数(多模态)
|
||
* - tool_calls 流式处理
|
||
*
|
||
* @see apis/ollama-api-docs-20260518.html
|
||
*/
|
||
|
||
import { BaseAdapter } from './base-adapter';
|
||
import { truncatedArgumentsPayload, readStreamChunkWithIdleTimeout } from './shared/sse-stream';
|
||
import log from 'electron-log';
|
||
import { nanoid } from 'nanoid';
|
||
import type { MetonaRequest, MetonaResponse, MetonaStreamEvent } from '../types';
|
||
import { MetonaFinishReason, MetonaStreamEventType } from '../types';
|
||
import type { MetonaModelInfo } from '../types/metona-adapter';
|
||
|
||
export class OllamaAdapter extends BaseAdapter {
|
||
// H-2 修复: provider → providerId(规范要求)
|
||
override readonly providerId: string = 'ollama';
|
||
readonly supportedModels = ['qwen3:latest', 'gemma3:latest', 'deepseek-r1:latest'];
|
||
readonly supportsToolCalling = true;
|
||
readonly supportsThinking = true;
|
||
|
||
// H-2 修复: Ollama 本地模型默认上下文窗口(可由 options.num_ctx 覆盖)
|
||
private static readonly DEFAULT_CONTEXT_WINDOW = 4096;
|
||
|
||
private baseURL: string;
|
||
|
||
constructor(config: ConstructorParameters<typeof BaseAdapter>[0]) {
|
||
super(config);
|
||
this.baseURL = config.baseURL || 'http://localhost:11434';
|
||
// v0.6.4 P4-1: 每个适配器实例(= 每会话独立引擎)启动时做一次 /api/show 探测,
|
||
// 把 num_ctx 实测值填充进 getContextWindow 缓存。fire-and-forget:失败静默,
|
||
// 不阻塞/不影响首个请求;此后压缩预算基于实测窗口而非保守默认 4096。
|
||
this.refreshContextWindow();
|
||
}
|
||
|
||
// ===== POST /api/chat =====
|
||
|
||
// H-2 修复: chat → send(规范要求)
|
||
async send(request: MetonaRequest): Promise<MetonaResponse> {
|
||
// #2 修复: toNativeRequest 改为 async(需下载 URL 图片转 base64)
|
||
const nativeRequest = await this.toNativeRequest(request);
|
||
|
||
// #24 修复: 使用 fetchWithTimeout 替代 getFetchSignal + fetch,确保 timer 清理
|
||
const response = await this.fetchWithTimeout(
|
||
`${this.baseURL}/api/chat`,
|
||
{
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify({ ...nativeRequest, stream: false }),
|
||
},
|
||
this.config.timeoutMs ?? 300_000,
|
||
);
|
||
|
||
if (!response.ok) {
|
||
await this.throwHttpError(response, 'Ollama API error');
|
||
}
|
||
|
||
const data = (await response.json()) as Record<string, unknown>;
|
||
return this.toMetonaResponse(data, request.meta.requestId, request.meta.iteration);
|
||
}
|
||
|
||
// H-2 修复: chatStream → sendStream(规范要求)
|
||
async *sendStream(request: MetonaRequest): AsyncIterable<MetonaStreamEvent> {
|
||
// #2 修复: toNativeRequest 改为 async(需下载 URL 图片转 base64)
|
||
const nativeRequest = await this.toNativeRequest(request);
|
||
|
||
// #24 修复: 使用 fetchWithTimeout 替代 getFetchSignal + fetch,确保 timer 清理
|
||
const response = await this.fetchWithTimeout(
|
||
`${this.baseURL}/api/chat`,
|
||
{
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify({ ...nativeRequest, stream: true }),
|
||
},
|
||
this.config.timeoutMs ?? 300_000,
|
||
);
|
||
|
||
if (!response.ok || !response.body) {
|
||
await this.throwHttpError(response, 'Ollama stream error');
|
||
}
|
||
|
||
// 非空断言:上方 if 已确保 response.body 不为 null
|
||
const reader = response.body!.getReader();
|
||
const decoder = new TextDecoder();
|
||
let seq = 0;
|
||
let buffer = '';
|
||
let streamEndedNormally = false;
|
||
|
||
while (true) {
|
||
// v0.7.4 P1-2: 空闲超时 — 本地模型加载/推理期间服务器可能长时间不推数据,
|
||
// 共享辅助在连续 60s 无数据时抛 SseUpstreamError(504) 进重试通道
|
||
const { done, value } = await readStreamChunkWithIdleTimeout(reader);
|
||
if (done) break;
|
||
|
||
buffer += decoder.decode(value, { stream: true });
|
||
const lines = buffer.split('\n');
|
||
buffer = lines.pop() ?? '';
|
||
|
||
for (const line of lines) {
|
||
const trimmed = line.trim();
|
||
if (!trimmed) continue;
|
||
|
||
try {
|
||
const chunk = JSON.parse(trimmed);
|
||
|
||
// 思考内容
|
||
if (chunk.message?.thinking) {
|
||
yield {
|
||
type: MetonaStreamEventType.REASONING_DELTA,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
delta: chunk.message.thinking,
|
||
};
|
||
}
|
||
|
||
// 文本内容
|
||
if (chunk.message?.content) {
|
||
yield {
|
||
type: MetonaStreamEventType.TEXT_DELTA,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
delta: chunk.message.content,
|
||
};
|
||
}
|
||
|
||
// 工具调用(Ollama 在最后一个 chunk 中整块返回)
|
||
if (chunk.message?.tool_calls) {
|
||
for (const tc of chunk.message.tool_calls) {
|
||
const args = tc.function?.arguments;
|
||
// v0.6.4 缺口修复: NDJSON 路径的截断自愈 —— 原实现 JSON.parse 抛错会
|
||
// 落入外层 catch:该 tool call 整体静默丢弃,且同一行剩余处理
|
||
// (含 done/USAGE 检查)一并被跳过,与 v0.6.3 已根治的 OpenAI 共享层
|
||
// 旧行为完全相同。现独立捕获并转为 _truncatedArguments 自愈载荷,
|
||
// 同时保证本 chunk 的后续分支照常执行。
|
||
let parsedArgs: Record<string, unknown>;
|
||
if (typeof args === 'string') {
|
||
try {
|
||
parsedArgs = JSON.parse(args);
|
||
} catch (parseErr) {
|
||
const sample = args.slice(-120);
|
||
log.warn(
|
||
`[Ollama] Tool call args truncated (unparseable JSON, ${(parseErr as Error).message}). Tail: ...${sample}`,
|
||
);
|
||
parsedArgs = truncatedArgumentsPayload((parseErr as Error).message, sample);
|
||
}
|
||
} else {
|
||
parsedArgs = (args as Record<string, unknown>) ?? {};
|
||
}
|
||
yield {
|
||
type: MetonaStreamEventType.TOOL_CALL_COMPLETE,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
toolCall: {
|
||
// L-9 修复: 统一使用 nanoid 生成工具调用 ID(与 sse-stream.ts 一致)
|
||
id: `tc_${nanoid(8)}`,
|
||
name: tc.function?.name ?? '',
|
||
args: parsedArgs,
|
||
iteration: request.meta.iteration,
|
||
timestamp: Date.now(),
|
||
},
|
||
};
|
||
}
|
||
}
|
||
|
||
// 流结束
|
||
if (chunk.done) {
|
||
streamEndedNormally = true;
|
||
// v0.8.0 P0-1: 采集 done_reason —— Ollama 的停止原因在最终 done chunk
|
||
// 携带;旧实现流式路径完全忽略(length 截断不可见),现归一化后随
|
||
// DONE 事件上交引擎
|
||
const doneReason = mapOllamaDoneReason(
|
||
chunk.done_reason as string | undefined,
|
||
Array.isArray(chunk.message?.tool_calls) && chunk.message.tool_calls.length > 0,
|
||
);
|
||
const finishReason = doneReason === MetonaFinishReason.LENGTH ? 'length' : undefined;
|
||
// 发送 usage 信息
|
||
yield {
|
||
type: MetonaStreamEventType.USAGE,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
usage: {
|
||
inputTokens: chunk.prompt_eval_count ?? 0,
|
||
outputTokens: chunk.eval_count ?? 0,
|
||
totalTokens: (chunk.prompt_eval_count ?? 0) + (chunk.eval_count ?? 0),
|
||
},
|
||
};
|
||
|
||
yield {
|
||
type: MetonaStreamEventType.DONE,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
// v0.8.0 P0-1: 仅 length 需要显式上交(stop/tool_calls 语义由
|
||
// 事件流本身表达;load/unload 属本地引擎状态非响应语义)
|
||
...(finishReason ? { finishReason } : {}),
|
||
};
|
||
return;
|
||
}
|
||
} catch (parseErr) {
|
||
// P2-8 修复: 与 sse-stream.ts 一致,记录解析失败行便于诊断
|
||
log.warn(
|
||
`[Ollama] Failed to parse NDJSON line: ${(parseErr as Error).message}`,
|
||
trimmed.slice(0, 200),
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
// 流未正常结束(连接断开等),补发 DONE 事件防止 Agent Loop 挂起
|
||
if (!streamEndedNormally) {
|
||
yield {
|
||
type: MetonaStreamEventType.DONE,
|
||
requestId: request.meta.requestId,
|
||
sessionId: request.meta.sessionId,
|
||
iteration: request.meta.iteration,
|
||
seq: seq++,
|
||
timestamp: Date.now(),
|
||
};
|
||
}
|
||
}
|
||
|
||
// ===== POST /api/generate =====
|
||
|
||
async generate(params: {
|
||
model: string;
|
||
prompt: string;
|
||
suffix?: string;
|
||
system?: string;
|
||
stream?: boolean;
|
||
think?: boolean | string;
|
||
format?: string | object;
|
||
images?: string[];
|
||
options?: Record<string, unknown>;
|
||
/** v0.8.0 P1-3.2: 外部取消信号(与固定 300s 超时合并,任一触发即中止) */
|
||
signal?: AbortSignal;
|
||
}): Promise<{
|
||
response: string;
|
||
thinking?: string;
|
||
done: boolean;
|
||
totalDuration: number;
|
||
evalCount: number;
|
||
}> {
|
||
const response = await fetch(`${this.baseURL}/api/generate`, {
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify({ ...params, stream: false, signal: undefined }),
|
||
// v0.8.0 P1-3.2 根治: 旧实现固定 AbortSignal.timeout(300_000) —— 外部中断
|
||
// 无法取消该请求(最长 5 分钟资源悬挂)。现用 AbortSignal.any 合并外部
|
||
// signal 与超时信号,任一触发即中止(Node ≥20.3 / Electron 35 满足)。
|
||
signal: params.signal
|
||
? AbortSignal.any([params.signal, AbortSignal.timeout(300_000)])
|
||
: AbortSignal.timeout(300_000),
|
||
});
|
||
|
||
if (!response.ok) throw new Error(`Ollama generate error: ${response.status}`);
|
||
const data = (await response.json()) as {
|
||
response?: string;
|
||
thinking?: string;
|
||
done?: boolean;
|
||
total_duration?: number;
|
||
eval_count?: number;
|
||
};
|
||
|
||
return {
|
||
response: data.response ?? '',
|
||
thinking: data.thinking,
|
||
done: data.done ?? true,
|
||
totalDuration: data.total_duration ?? 0,
|
||
evalCount: data.eval_count ?? 0,
|
||
};
|
||
}
|
||
|
||
// ===== POST /api/embed =====
|
||
|
||
async embed(params: {
|
||
model: string;
|
||
input: string | string[];
|
||
dimensions?: number;
|
||
/** v0.8.0 P1-3.2: 外部取消信号(与固定 60s 超时合并) */
|
||
signal?: AbortSignal;
|
||
}): Promise<{ embeddings: number[][]; totalDuration: number }> {
|
||
const response = await fetch(`${this.baseURL}/api/embed`, {
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify(params),
|
||
// v0.8.0 P1-3.2: 与 generate 同口径 —— 外部 signal 与超时合并
|
||
signal: params.signal
|
||
? AbortSignal.any([params.signal, AbortSignal.timeout(60_000)])
|
||
: AbortSignal.timeout(60_000),
|
||
});
|
||
|
||
if (!response.ok) throw new Error(`Ollama embed error: ${response.status}`);
|
||
const data = (await response.json()) as { embeddings?: number[][]; total_duration?: number };
|
||
|
||
return {
|
||
embeddings: data.embeddings ?? [],
|
||
totalDuration: data.total_duration ?? 0,
|
||
};
|
||
}
|
||
|
||
// ===== GET /api/tags =====
|
||
|
||
/**
|
||
* H-2 修复: 返回 MetonaModelInfo[](规范要求)
|
||
*
|
||
* Ollama /api/tags 返回模型列表含详细信息(name, size, details),
|
||
* 转换为 MetonaModelInfo 并补充默认元数据。
|
||
*/
|
||
async listModels(): Promise<MetonaModelInfo[]> {
|
||
try {
|
||
const response = await fetch(`${this.baseURL}/api/tags`, {
|
||
signal: AbortSignal.timeout(10_000),
|
||
});
|
||
if (response.ok) {
|
||
const data = (await response.json()) as {
|
||
models?: Array<{
|
||
name: string;
|
||
size?: number;
|
||
details?: { parameter_size?: string; quantization_level?: string; family?: string };
|
||
}>;
|
||
};
|
||
if (data.models?.length) {
|
||
// v0.6.4 P4-1: 能力标志改为逐模型 /api/show 实测探测;单个探测失败
|
||
// 该模型回退保守 true(不可用时行为与旧实现一致,fail-open 保可用性)
|
||
// v0.7.3 P1-4: supportsVision 随探测结果透出(undefined = 未知 → 前端保守放行),
|
||
// 供上传入口拒绝不支持图片的本地语言模型
|
||
const enriched = await Promise.all(
|
||
data.models.map(async (m) => {
|
||
const caps = await this.probeCapabilities(m.name);
|
||
return {
|
||
id: m.name,
|
||
name: m.name,
|
||
// Ollama 模型上下文窗口由 options.num_ctx 决定,此处给保守值
|
||
contextWindow: OllamaAdapter.DEFAULT_CONTEXT_WINDOW,
|
||
supportsToolCalling: caps ? caps.supportsTools : true,
|
||
supportsThinking: caps ? caps.supportsThinking : true,
|
||
supportsVision: caps ? caps.supportsVision : undefined,
|
||
description: m.details
|
||
? `${m.details.family ?? 'unknown'} / ${m.details.parameter_size ?? '?'} / ${m.details.quantization_level ?? '?'}`
|
||
: undefined,
|
||
};
|
||
}),
|
||
);
|
||
return enriched;
|
||
}
|
||
}
|
||
} catch {
|
||
// API 不可用时降级
|
||
}
|
||
// 回退到 supportedModels
|
||
return this.supportedModels.map((id) => ({ id }));
|
||
}
|
||
|
||
/**
|
||
* H-2 修复: 获取上下文窗口大小(规范要求)
|
||
*
|
||
* Ollama 上下文窗口由 options.num_ctx 决定(默认 4096),
|
||
* Engine 应通过 MetonaRequest.params.contextLength 显式设置。
|
||
* 此处返回默认值,供 Engine 在未指定时参考。
|
||
*/
|
||
override getContextWindow(): number {
|
||
return this.cachedContextWindow ?? OllamaAdapter.DEFAULT_CONTEXT_WINDOW;
|
||
}
|
||
|
||
/**
|
||
* v0.6.4 P4-1: 从 /api/show 的 parameters 区解析 num_ctx 真值。
|
||
*
|
||
* 契约约束:IMetonaProviderAdapter.getContextWindow 是同步接口(引擎压缩判定
|
||
* 依赖同步取值),无法在内部 await。因此采用"机会主义缓存"策略:
|
||
* send/sendStream 启动时 fire-and-forget 刷新缓存;首次请求前返回默认 4096,
|
||
* 之后永远返回实测值。压缩预算的准确性随使用逐渐收敛到真值。
|
||
*/
|
||
private cachedContextWindow: number | null = null;
|
||
private refreshingContextWindow = false;
|
||
/** v0.8.0 P0-3: /api/show 探测只发一次(构造函数发起),失败不重试(fail-open) */
|
||
private probeAttempted = false;
|
||
/**
|
||
* v0.8.0 P0-3: /api/show capabilities 探测缓存 —— 模型是否支持思考。
|
||
* null = 未探测/探测失败(fail-open 放行,与 listModels 能力回退策略一致);
|
||
* false = 服务端明确不支持 → toNativeRequest 不发 think 参数。
|
||
*/
|
||
private cachedThinkingSupport: boolean | null = null;
|
||
|
||
private refreshContextWindow(): void {
|
||
if (this.refreshingContextWindow || this.probeAttempted) return;
|
||
this.probeAttempted = true;
|
||
this.refreshingContextWindow = true;
|
||
void this.showModel(this.config.defaultModel)
|
||
.then((info) => {
|
||
if (!info?.parameters) return;
|
||
const match = /^num_ctx\s+(\d+)\s*$/m.exec(info.parameters);
|
||
if (match) {
|
||
const value = Number(match[1]);
|
||
if (Number.isFinite(value) && value > 0) {
|
||
this.cachedContextWindow = value;
|
||
log.info(`[Ollama] Context window (num_ctx) detected: ${value}`);
|
||
}
|
||
}
|
||
// v0.8.0 P0-3: 同一次探测顺带缓存思考能力(供 toNativeRequest 同步门控)
|
||
if (Array.isArray(info.capabilities)) {
|
||
this.cachedThinkingSupport = info.capabilities.map(String).includes('thinking');
|
||
}
|
||
})
|
||
.catch(() => {
|
||
/* 模型探测失败不阻塞对话 */
|
||
})
|
||
.finally(() => {
|
||
this.refreshingContextWindow = false;
|
||
});
|
||
}
|
||
|
||
/**
|
||
* v0.6.4 P4-1: 通过 /api/show 的 capabilities[] 动态探测模型真实能力。
|
||
* 此前 listModels 对所有本地模型硬编码 supportsToolCalling/supportsThinking:true
|
||
* (注释自知不准)—— 语言模型不支持 tools 时引擎仍下发工具定义,
|
||
* 造成"模型口头说调工具实际不调"的回归温床。探测失败返回 null 由调用方回退保守值。
|
||
*/
|
||
async probeCapabilities(model: string): Promise<{
|
||
supportsTools: boolean;
|
||
supportsVision: boolean;
|
||
supportsThinking: boolean;
|
||
} | null> {
|
||
const info = await this.showModel(model);
|
||
if (!info || !Array.isArray(info.capabilities)) return null;
|
||
const caps = new Set(info.capabilities.map((c) => String(c)));
|
||
return {
|
||
supportsTools: caps.has('tools'),
|
||
supportsVision: caps.has('vision'),
|
||
supportsThinking: caps.has('thinking'),
|
||
};
|
||
}
|
||
|
||
// ===== POST /api/show =====
|
||
|
||
async showModel(
|
||
model: string,
|
||
): Promise<{ parameters: string; template: string; capabilities: string[] } | null> {
|
||
try {
|
||
const response = await fetch(`${this.baseURL}/api/show`, {
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify({ model }),
|
||
signal: AbortSignal.timeout(10_000),
|
||
});
|
||
if (!response.ok) return null;
|
||
const data = (await response.json()) as {
|
||
parameters?: string;
|
||
template?: string;
|
||
capabilities?: string[];
|
||
};
|
||
return {
|
||
parameters: data.parameters ?? '',
|
||
template: data.template ?? '',
|
||
capabilities: data.capabilities ?? [],
|
||
};
|
||
} catch {
|
||
return null;
|
||
}
|
||
}
|
||
|
||
// ===== POST /api/pull =====
|
||
|
||
/**
|
||
* v0.6.4 P4-1 重构:pull 支持外部取消信号 —— 原实现固定 600s 超时会掐死
|
||
* 大模型下载(进度不能续命、无取消通道),大仓/慢网络场景必然失败。
|
||
* 现契约:调用方通过 AbortSignal 控制生命周期(UI 取消按钮即可触发);
|
||
* 超时语义交给用户取消或服务端断流(读循环结束即完成),不再人为设上限。
|
||
*/
|
||
async pullModel(
|
||
model: string,
|
||
onProgress?: (progress: { status: string; completed?: number; total?: number }) => void,
|
||
signal?: AbortSignal,
|
||
): Promise<void> {
|
||
const response = await fetch(`${this.baseURL}/api/pull`, {
|
||
method: 'POST',
|
||
headers: { 'Content-Type': 'application/json' },
|
||
body: JSON.stringify({ model, stream: true }),
|
||
signal,
|
||
});
|
||
|
||
if (!response.ok || !response.body) throw new Error(`Ollama pull error: ${response.status}`);
|
||
|
||
const reader = response.body.getReader();
|
||
const decoder = new TextDecoder();
|
||
let buffer = '';
|
||
|
||
while (true) {
|
||
const { done, value } = await reader.read();
|
||
if (done) break;
|
||
buffer += decoder.decode(value, { stream: true });
|
||
const lines = buffer.split('\n');
|
||
buffer = lines.pop() ?? '';
|
||
for (const line of lines) {
|
||
if (!line.trim()) continue;
|
||
try {
|
||
const chunk = JSON.parse(line);
|
||
onProgress?.({ status: chunk.status, completed: chunk.completed, total: chunk.total });
|
||
} catch {
|
||
// L-3 修复: 添加日志便于诊断非标准行(如进度通知、空行等)
|
||
log.debug('[Ollama] skipped non-JSON line during pull:', line.slice(0, 100));
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
// ===== GET /api/ps =====
|
||
|
||
async listRunning(): Promise<
|
||
Array<{ name: string; size: number; sizeVram: number; contextLength: number }>
|
||
> {
|
||
try {
|
||
const response = await fetch(`${this.baseURL}/api/ps`, {
|
||
signal: AbortSignal.timeout(10_000),
|
||
});
|
||
if (!response.ok) return [];
|
||
const data = (await response.json()) as {
|
||
models?: Array<{
|
||
name: string;
|
||
size?: number;
|
||
size_vram?: number;
|
||
context_length?: number;
|
||
}>;
|
||
};
|
||
return (data.models ?? []).map((m) => ({
|
||
name: m.name ?? '',
|
||
size: m.size ?? 0,
|
||
sizeVram: m.size_vram ?? 0,
|
||
contextLength: m.context_length ?? 0,
|
||
}));
|
||
} catch {
|
||
return [];
|
||
}
|
||
}
|
||
|
||
// ===== GET /api/version =====
|
||
|
||
async getVersion(): Promise<string> {
|
||
try {
|
||
const response = await fetch(`${this.baseURL}/api/version`, {
|
||
signal: AbortSignal.timeout(5_000),
|
||
});
|
||
if (!response.ok) return 'unknown';
|
||
const data = (await response.json()) as { version?: string };
|
||
return data.version ?? 'unknown';
|
||
} catch {
|
||
return 'unknown';
|
||
}
|
||
}
|
||
|
||
// ========== 私有转换方法 ==========
|
||
|
||
/**
|
||
* #2 修复: 下载 http(s) URL 图片并转为纯 base64 字符串(不含 data: 前缀)
|
||
*
|
||
* Ollama API 的 images 字段要求纯 base64 字符串数组。
|
||
* 当 MetonaMessage.images 中存储的是 URL 时,需先下载转为 base64。
|
||
* 下载失败时返回空字符串(Ollama 会忽略空图片),不阻断整个请求。
|
||
*/
|
||
private async resolveImageToBase64(url: string): Promise<string> {
|
||
try {
|
||
// 审查修复: 使用基类 fetchWithTimeout 合并 externalAbortSignal 和 30s 超时,
|
||
// 避免用户中断时图片下载最多阻塞 30s×N(externalAbortSignal 是 BaseAdapter 的
|
||
// private 属性,子类无法直接访问,故复用已合并 signal 的 fetchWithTimeout,
|
||
// 该方法同时处理了 listener 泄漏问题)
|
||
const res = await this.fetchWithTimeout(url, {}, 30_000);
|
||
if (!res.ok) {
|
||
throw new Error(`HTTP ${res.status}`);
|
||
}
|
||
const buf = Buffer.from(await res.arrayBuffer());
|
||
return buf.toString('base64');
|
||
} catch (error) {
|
||
log.warn(
|
||
`[Ollama] Failed to download image ${url.slice(0, 100)}: ${(error as Error).message}`,
|
||
);
|
||
return '';
|
||
}
|
||
}
|
||
|
||
private async toNativeRequest(request: MetonaRequest): Promise<Record<string, unknown>> {
|
||
// 说明:能力/num_ctx 探测由构造函数 fire-and-forget 发起(probeAttempts 上限
|
||
// 1 次)—— 此处不再重复探测,避免每次请求都打 /api/show(也保持测试环境的
|
||
// fetch 捕获不受污染)。探测失败时 cachedThinkingSupport=null → 门控 fail-open。
|
||
const messages: Record<string, unknown>[] = [
|
||
{
|
||
role: 'system',
|
||
content: [
|
||
request.systemPrompt.roleDefinition,
|
||
request.systemPrompt.outputConstraints,
|
||
request.systemPrompt.safetyGuidelines,
|
||
request.systemPrompt.dynamicReminders,
|
||
]
|
||
.filter(Boolean)
|
||
.join('\n\n'),
|
||
},
|
||
];
|
||
|
||
// #2 修复: 改为 for 循环以支持 async 图片下载(map 回调无法 await)
|
||
for (const m of request.messages) {
|
||
if (m.role === 'system') continue;
|
||
// C-6 修复: Ollama API 不支持 null content,assistant 仅有 tool_calls 时转为空字符串
|
||
const msg: Record<string, unknown> = { role: m.role, content: m.content ?? '' };
|
||
// Ollama 图片使用 images 字段(纯 base64 数组,不含 data: 前缀)
|
||
if (m.images?.length) {
|
||
// #2 修复: 支持公网 URL 图片,下载后转为纯 base64
|
||
// 之前直接将 URL 字符串传给 Ollama,导致 base64 解码错误
|
||
const resolvedImages: string[] = [];
|
||
for (const img of m.images) {
|
||
const url = img.url;
|
||
if (url.startsWith('data:')) {
|
||
// data:image/png;base64,iVBOR... → iVBOR...
|
||
const base64Part = url.split(',')[1];
|
||
resolvedImages.push(base64Part ?? url);
|
||
} else if (url.startsWith('http://') || url.startsWith('https://')) {
|
||
// #2 修复: 公网 URL → 下载 → 纯 base64
|
||
const base64 = await this.resolveImageToBase64(url);
|
||
if (base64) resolvedImages.push(base64);
|
||
} else {
|
||
// 已是纯 base64 字符串(无 data: 前缀)
|
||
resolvedImages.push(url);
|
||
}
|
||
}
|
||
msg.images = resolvedImages;
|
||
}
|
||
// 工具结果
|
||
if (m.role === 'tool' && m.toolResult) {
|
||
msg.tool_call_id = m.toolResult.toolCallId;
|
||
// CE-2 修复: 工具失败时 result 为 null,优先用 error 字段作为 content
|
||
msg.content = m.toolResult.error
|
||
? m.toolResult.error
|
||
: typeof m.toolResult.result === 'string'
|
||
? m.toolResult.result
|
||
: JSON.stringify(m.toolResult.result);
|
||
}
|
||
// assistant 工具调用(Ollama REST API 要求 arguments 为 JSON 字符串)
|
||
if (m.role === 'assistant' && m.toolCalls?.length) {
|
||
msg.tool_calls = m.toolCalls.map((tc) => ({
|
||
function: { name: tc.name, arguments: JSON.stringify(tc.args) },
|
||
}));
|
||
}
|
||
// 推理内容回传(保持多轮推理链完整)
|
||
if (m.role === 'assistant' && m.reasoningContent) {
|
||
(msg as Record<string, unknown>).reasoning_content = m.reasoningContent;
|
||
}
|
||
messages.push(msg);
|
||
}
|
||
|
||
const body: Record<string, unknown> = {
|
||
model: this.config.defaultModel,
|
||
messages,
|
||
options: {
|
||
temperature: request.params.temperature,
|
||
num_predict: request.params.maxTokens,
|
||
...(request.params.topP != null && { top_p: request.params.topP }),
|
||
...(request.params.stopSequences?.length && { stop: request.params.stopSequences }),
|
||
...(request.params.contextLength != null && { num_ctx: request.params.contextLength }),
|
||
},
|
||
};
|
||
|
||
// Tool Calling
|
||
if (request.tools?.length) {
|
||
body.tools = request.tools.map((t) => ({
|
||
type: 'function',
|
||
function: {
|
||
name: t.name,
|
||
description: t.description,
|
||
parameters: t.parameters,
|
||
},
|
||
}));
|
||
}
|
||
|
||
// Thinking 模式
|
||
// v0.8.0 P0-3: 模型能力门控 —— /api/show capabilities 探测为不支持思考
|
||
//(cachedThinkingSupport === false)时不发 think 参数(Ollama 服务端默认关闭);
|
||
// 未探测/探测失败(null)fail-open 放行,与 listModels 能力回退策略一致。
|
||
// 探测在适配器实例创建时 fire-and-forget 发起(refreshContextWindow)。
|
||
if (request.params.thinkingEnabled) {
|
||
const modelThinkingSupported = this.cachedThinkingSupport !== false;
|
||
if (modelThinkingSupported) {
|
||
const effortMap: Record<string, string | boolean> = {
|
||
low: 'low',
|
||
medium: 'medium',
|
||
high: 'high',
|
||
max: true,
|
||
};
|
||
body.think = effortMap[request.params.thinkingEffort ?? 'high'] ?? true;
|
||
// v0.8.0 P0-3: 思考占用 num_predict 输出预算 —— 预算过小时显式告警
|
||
const numPredict = request.params.maxTokens;
|
||
if (typeof numPredict === 'number' && numPredict < 8192) {
|
||
log.warn(
|
||
`[Ollama] thinking enabled with small num_predict budget (${numPredict}) — reasoning may consume the entire budget and truncate the answer`,
|
||
);
|
||
}
|
||
} else {
|
||
log.warn(
|
||
`[Ollama] model "${this.config.defaultModel}" does not support thinking (per /api/show capabilities) — omitting think parameter (P0-3 capability gate)`,
|
||
);
|
||
}
|
||
}
|
||
|
||
return body;
|
||
}
|
||
|
||
private toMetonaResponse(
|
||
data: Record<string, unknown>,
|
||
requestId: string,
|
||
iteration: number = 0,
|
||
): MetonaResponse {
|
||
const message = data.message as Record<string, unknown> | undefined;
|
||
const toolCalls = message?.tool_calls as Array<Record<string, unknown>> | undefined;
|
||
return {
|
||
meta: {
|
||
requestId,
|
||
provider: this.providerId,
|
||
model: (data.model as string) ?? this.config.defaultModel,
|
||
latencyMs: 0,
|
||
timestamp: Date.now(),
|
||
perfStats: {
|
||
loadDurationMs: data.load_duration ? (data.load_duration as number) / 1e6 : undefined,
|
||
promptEvalDurationMs: data.prompt_eval_duration
|
||
? (data.prompt_eval_duration as number) / 1e6
|
||
: undefined,
|
||
evalDurationMs: data.eval_duration ? (data.eval_duration as number) / 1e6 : undefined,
|
||
tokensPerSecond:
|
||
data.eval_count && data.eval_duration
|
||
? (data.eval_count as number) / ((data.eval_duration as number) / 1e9)
|
||
: undefined,
|
||
},
|
||
},
|
||
content: (message?.content as string) ?? '',
|
||
reasoningContent: message?.thinking as string | undefined,
|
||
toolCalls: toolCalls?.map((tc) => {
|
||
const fn = tc.function as Record<string, unknown>;
|
||
const rawArgs = fn?.arguments;
|
||
let args: Record<string, unknown> = {};
|
||
try {
|
||
args =
|
||
typeof rawArgs === 'string'
|
||
? JSON.parse(rawArgs)
|
||
: ((rawArgs as Record<string, unknown>) ?? {});
|
||
} catch (parseErr) {
|
||
// v0.6.4: 非流式路径截断自愈对齐 —— 原 catch 静默降级 {},与流式修复后的
|
||
// 行为不一致。统一转为 _truncatedArguments 错误参数。
|
||
const sample =
|
||
typeof rawArgs === 'string' ? rawArgs.slice(-120) : String(rawArgs).slice(-120);
|
||
log.warn(
|
||
`[Ollama] Non-stream tool call args truncated (unparseable JSON, ${(parseErr as Error).message}). Tail: ...${sample}`,
|
||
);
|
||
args = truncatedArgumentsPayload((parseErr as Error).message, sample);
|
||
}
|
||
return {
|
||
// L-9 修复(审计补充): 非流式路径统一使用 nanoid,与流式路径(sendStream)保持一致
|
||
id: `tc_${nanoid(8)}`,
|
||
name: (fn?.name as string) ?? '',
|
||
args,
|
||
iteration,
|
||
timestamp: Date.now(),
|
||
};
|
||
}),
|
||
usage: {
|
||
inputTokens: (data.prompt_eval_count as number) ?? 0,
|
||
outputTokens: (data.eval_count as number) ?? 0,
|
||
totalTokens: ((data.prompt_eval_count as number) ?? 0) + ((data.eval_count as number) ?? 0),
|
||
},
|
||
finishReason: mapOllamaDoneReason(
|
||
data.done_reason as string | undefined,
|
||
!!message?.tool_calls,
|
||
),
|
||
};
|
||
}
|
||
}
|
||
|
||
/**
|
||
* 映射 Ollama done_reason → MetonaFinishReason
|
||
*
|
||
* @see apis/ollama-api-docs-20260518.html — /api/chat 响应字段
|
||
*/
|
||
function mapOllamaDoneReason(
|
||
reason: string | undefined,
|
||
hasToolCalls: boolean,
|
||
): MetonaFinishReason {
|
||
if (hasToolCalls) return MetonaFinishReason.TOOL_CALLS;
|
||
switch (reason) {
|
||
case 'stop':
|
||
return MetonaFinishReason.STOP;
|
||
case 'length':
|
||
return MetonaFinishReason.LENGTH;
|
||
case 'load':
|
||
return MetonaFinishReason.STOP; // 冷启动加载完成,非错误
|
||
case 'unload':
|
||
return MetonaFinishReason.STOP;
|
||
default:
|
||
return MetonaFinishReason.STOP;
|
||
}
|
||
}
|