Files
searxng-use-cli/SKILL.md
T
thzxx 4df521dc9d feat(v2.3.0): fetch 结构化 JSON 契约 + 并发批量 + 真实并发门控
迭代 1 — 正确性修复:
- 修复 --pages N 多页聚合的 unresponsive-engine 警告误判: 原用循环末次
  cached 变量判断, 缓存命中时警告被错误跳过/误触发; 改用独立
  performed_live_query 标记
- UA 池单一来源: 删除 common.py 手工副本 _FALLBACK_UAS_BUILTIN,
  FALLBACK_UAS 直接引用 _config.UA_POOL, 消除双份漂移
- search_html 解码修复: 硬编码 utf-8 改为 detect_charset(header/meta
  自动检测), 新增 --encoding 强制覆盖, 贯穿 search_multi 全链
- AdaptiveThrottle 真实并发门控: acquire_slot()/release_slot() 槽位机制,
  退避降并发后新请求被快速拒绝(E_RATE_LIMIT), 实现持久降并发而非名义降并发

迭代 2 — fetch JSON 契约 + 批量并发:
- fetch.py --format json: 成功 {status,url,final_url,content_type,extract,
  truncated,text_length,user_agent}; 失败 {status,error,error_code,
  status_code,url}, 对齐 search.py 错误码体系
- fetch_page 采集 title + latency, 填充 --fetch-report json 空字段
- --queries-file --parallel-queries N (1-8): 并发批量, 输出保序, 受
  AdaptiveThrottle 门控; 并发模式禁用 --fetch(嵌套并行不安全)
- queries 文件编码自动检测 (UTF-8 → GBK 回退)

迭代 3 — 工程化:
- 新增 pyproject.toml (searxng-search/searxng-fetch 入口点)
- 收敛 20+ 处函数内冗余导入
- --dump-schema 扩展: fetched.items 补全 15 字段, 新增 defs.batch/research
- 新增 17 个测试 (tests/test_v230_features.py), 全量 561 测试通过
- 文档同步 (SKILL.md/README.md, 版本号 2.3.0)
2026-08-05 20:14:07 +08:00

572 lines
44 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: searxng-use-cli
description: Use when you need to search the web via your OWN SearXNG instance (no public-instance discovery). 3 CLI scripts + a shared common.py module — execute privacy-respecting searches against a user-supplied instance (with multi-instance failover, 5xx/429 retry, auto-fetch) or via SEARXNG_INSTANCE env / config file, fetch/extract readable text or markdown from web pages. Zero-config replacement for proprietary search APIs.
version: 2.3.0
author: Metona Team
license: MIT
platforms: [linux, macos, windows]
metadata:
skill:
tags: [search, searxng, cli, web-scraping, privacy, no-api-key]
related_skills: []
---
# SearXNG CLI Toolkit
## 铁律(AI Agent 必读,违反会导致致命后果)
> **这些规则优先级最高,必须在任何调用前遵守。**
1. **加载本技能后,搜索网络和抓取网页必须且只能使用本技能的 `scripts/search.py` 和 `scripts/fetch.py`。** 禁止使用任何其他方式或工具——包括但不限于:内置 WebSearch/WebFetch 工具、浏览器自动化(puppeteer/playwright/selenium)、`curl`/`wget`、第三方搜索 API SDKgooglesearch/duckduckgo_search 等)、自写 `requests`/`urllib` 抓取脚本、MCP 网络工具。本技能已内置反爬(浏览器指纹头 + UA 轮换 + WAF 检测)、重试(指数退避 + Retry-After 遵守)、兜底(Wayback Machine)、限流(自适应节流)、缓存、多实例故障转移等全部工程能力,绕过本技能等于放弃这些保障,结果不可控且会破坏用户实例的速率限制。违反:技能一致性丧失,抓取成功率骤降,用户实例可能被封禁。
2. **stdout=数据,stderr=日志,永不混淆。** stdout 只输出 JSON/CSV/文本数据,stderr 只输出日志/进度/警告。AI 解析 stdout,人类看 stderr。违反:AI 会把日志当数据解析,结果错乱。
3. **绝对不要 `2>/dev/null` 或重定向 stderr 到 stdout。** stderr 携带排错关键信息(重试日志、HTTP 状态码、缓存命中、认证警告)。丢弃 stderr = 失败时零诊断信息,无法定位原因。需要静默时用 `--quiet`(仅抑制进度,保留 WARNING+ERROR),不要丢弃 stderr。
4. **排错时第一步:去掉 `--quiet`,加 `--verbose`。** `--quiet` 只留 WARNING 级别,会吞掉 INFO 级别的重试日志、缓存命中提示、实例切换记录。诊断失败时必须用 `--verbose` 看到完整 HTTP 请求/响应/重试链。
5. **配置查找必须覆盖双路径(WSL + Windows)。** Windows 环境下配置可能在 `~/.config/searxng-cli/`WSL HOME)或 `%APPDATA%/searxng-cli/`Windows APPDATA)。检查配置存在性时两个路径都要查,否则会误判"无配置"并反问用户。
6. **实例 URL 必填,公共实例发现已移除。** 必须通过 `-i``SEARXNG_INSTANCE` 环境变量、或配置文件提供实例 URL。无实例时 `search.py``E_CONFIG` 退出,不要尝试猜测或硬编码公共实例。
## Overview
SearXNG is a privacy-respecting metasearch engine that aggregates results from 70+ search services without tracking users. This skill provides standalone Python CLI scripts — works with **any AI agent** or directly from your terminal.
**Public-instance discovery has been removed.** You must supply your own SearXNG instance URL (self-hosted or one you trust). This makes behavior deterministic and avoids depending on volatile public instances.
**Key capabilities:**
**Search & results**
- Multi-instance failover with parallel probing (faster failover, deterministic output order)
- Exponential-backoff retry on transient errors (403/429/5xx/connection) via shared retry-policy components in `common.py` (backoff/retryable-status/Retry-After), with the retry loop in each script
- `--verify` health-check mode (reachability / JSON-API / latency / POST / engine list / auth status)
- Cross-engine result deduplication (default on; `--no-dedup` disables) — collapses duplicate URLs ignoring tracking params (`utm_*`, `gclid`, etc.) and fragments
- Similarity deduplication (`--similarity-dedup`; off by default) — title-based SimHash + Jaccard, auto-skipped when results > 500
- Result sorting (`--sort-by {score,date,engine,none}`; default: score descending) — applied after dedup, before `--max-results`
- Domain allowlist/blocklist (`--include-domain` / `--exclude-domain`) — case-insensitive, ignores leading `www.`, exclude wins on conflict
- Batch mode (`--queries-file`) — run multiple queries from a file in sequence, combined output; v2.3.0 adds `--parallel-queries N` (1-8 concurrent, output order preserved, gated by AdaptiveThrottle)
- Pagination aggregation (`--pages N`) — fetch N pages in one run, merge with cross-page dedup
**Output formats**
- JSON (default, rich metadata), brief (title+URL+snippet), urls (plain list), CSV (spreadsheet-ready)
- Enhanced Markdown conversion — nested ordered lists, mixed `ul`/`ol` nesting, `<dl>`/`<dt>`/`<dd>` definition lists, GFM tables, fenced code blocks, blockquotes
- Structured JSON error output (in `--format json` mode) with `error_code` field for machine-readable failure reporting
- JSON Lines streaming (`--stream`) — each result emitted as a separate JSON line to stdout, enabling incremental processing
- Progress events (`--progress`) — structured JSON Lines events to stderr for real-time execution tracking, includes `request_id` for correlation
- `fetch.py --format json` (v2.3.0) — structured JSON contract for page fetches: success `{status, url, final_url, content_type, extract, truncated, text_length, user_agent}`, failure `{status, error, error_code, status_code, url}` (aligns with search.py's error-code system)
**Caching & config**
- SQLite result caching (`--cache-ttl`) with size cap + LRU eviction (`--cache-max-size`, default 100MB) — identical queries within a TTL skip the network entirely; `--clear-cache` / `--cache-stats` manage it
- Config file (`searxng.toml`) pre-sets most flags; `--config FILE` loads a non-default config; `instances.txt` for plain URL lists; `--save-config FILE` writes current args to a config file
- Instance resolution priority: `-i``SEARXNG_INSTANCE` env → config file (`./searxng.toml``~/.config/searxng-cli/searxng.toml``%APPDATA%/searxng-cli/searxng.toml` on Windows) → `instances.txt` (`./instances.txt``~/.config/searxng-cli/instances.txt``%APPDATA%/searxng-cli/instances.txt`)
**Network & auth**
- Proxy support (`--proxy`) for both search and fetch (sets `HTTP_PROXY`/`HTTPS_PROXY`/`NO_PROXY`)
- Auth via CLI flag, file, config file, or env var (`--auth-bearer` / `--auth-basic` + `*-file` variants) to avoid leaking secrets in shell history
- Config-file auth — `auth_basic` and `auth_bearer` fields in `searxng.toml` let AI agents set credentials once (priority: CLI > file > config > env)
- Credentials-file permission warning — `--auth-*-file` warns on stderr if the file is group/other-readable (POSIX only)
**Anti-bot & fetch stability**
- **Browser fingerprint headers** — `build_browser_headers()` sends full Sec-Ch-Ua / Sec-Fetch-* / Accept-Language / Accept-Encoding, not just User-Agent. Bypasses 80%+ of lightweight WAFs (Sec-Ch-Ua/Sec-Fetch-* only sent for Chrome/Edge; Firefox omits them). **Smart Accept-Encoding**: only advertises `br` when the brotli decompressor is actually installed, preventing servers from returning Brotli-compressed bytes that requests can't auto-decompress (which would produce garbled output)
- **15-UA pool** — Chrome 138-140 (Win/mac/Linux), Edge 138 (Win/mac), Firefox 140 (Win/mac/Linux), Safari 18 (mac). `get_ua_for_domain()` deterministically assigns one UA per domain (stable within a session, reproducible across processes)
- **requests.Session reuse** — module-level Session with connection pool + cookie persistence + TLS session resumption
- **Split timeouts** — `timeout=(connect, read)` tuple (connect capped at 10s, read defaults to 15s)
- **Retry-After compliance** — 429/503 responses read the `Retry-After` header (numeric seconds or HTTP date) and wait at least that long before retrying
- **Capped backoff** — `compute_backoff_delay()` caps at 60s
- **WAF fingerprint library** — `_detect_anti_bot()` (in `search.py`, used by `--fetch` flow) identifies Cloudflare / Imperva / PerimeterX / DataDome / Akamai / Anubis / generic challenges via `<title>` tag matching + full-document scan. Returns `waf_type` for AI-agent decisioning
- **Wayback Machine fallback** — 404/403/timeout AND anti-bot-blocked pages automatically retry via `https://web.archive.org/web/2/<url>` (latest snapshot). Default ON; `--no-fallback` disables. Wayback timeout = `min(--timeout, 10s)`; Wayback retries = `min(--retries, 2)`
- **Hard-blocked domain fallback** — `is_hard_blocked_domain()` (in `common.py`) detects known strong-anti-bot sites (www.baidu.com, baike.baidu.com, zhihu.com, weibo.com, mp.weixin.qq.com, douban.com, etc.) that almost always return 403. These trigger Wayback immediately on failure. List maintained in `HARD_BLOCKED_DOMAINS`. Note: baidu uses exact subdomain matching (www/baike/zhidao/tieba/wenku) to avoid over-blocking pan.baidu.com; zhihu/weibo/douban use subdomain wildcard matching
- **Adaptive throttling** — `AdaptiveThrottle` (in `search.py`) state machine: consecutive failures → double delay + halve concurrency; consecutive successes → gradual recovery; 429 → global pause. Thread-safe. Parameters configurable via `--throttle-failure-threshold` / `--throttle-pause-seconds` / `--throttle-max-delay`. v2.3.0: real concurrency gating via `acquire_slot()`/`release_slot()` — after backoff reduces concurrency, new requests are rejected until in-flight drops below the target (persistent throttling, not just nominal)
- **readability-lite extraction** — when `<article>`/`<main>`/`role="main"`/content-class `<div>` are all missing and only `<body>` remains, picks the highest text-density `<div>`/`<section>`/`<article>` node (scored by text density + `<p>` count weighting). Text-density threshold is language-aware — CJK content uses 100 chars, other languages use 200 chars
- **PDF/document parsing** — `fetch.py` parses PDF (via `pdftotext` subprocess) and `.docx`/`.xlsx` (via stdlib `zipfile`). Unsupported binary types return `E_UNSUPPORTED_MEDIA`
- **`--fetch-report [json]`** — structured per-URL report to stderr after `--fetch N`: status, WAF type, fallback used, char count, plus throttle stats. `--fetch-report json` outputs a full JSON report (items array + summary)
- **`--referer`** — set Referer header for fetch requests (defaults to the instance URL when fetching result pages)
- **`--request-delay`** — configurable delay between fetch requests (default 0.3s; adaptive throttling may increase this on failures)
- **`search.py --fetch` output fields** (in `fetched` array): `anti_bot_detected` (bool), `waf_type` (str|null), `fallback_used` (str|null), `error_code` (str|null — structured error code on fetch failure); v2.3.0 also `title` (str|null) and `latency` (float|null seconds) — consumed by `--fetch-report json`
- **`fetch.py` FetchResult fields** (standalone script output): `content` (str), `content_type` (str), `final_url` (str), `truncated` (bool), `user_agent` (str), `error_code` (str|null), `error_message` (str|null); with `--format json` (v2.3.0) these map to a structured JSON contract
**Research mode**
- **`--research <topic>`** — given a research topic, auto-expands into 5 multi-angle queries (overview/profile/background/works/review) and runs them in sequence. Results include `research_topic` and `research_queries` metadata so AI agents can structure their final report by angle
- **Cross-angle merge & dedup** — JSON output includes a top-level `merged_results` field: all per-angle results are combined, deduplicated, and sorted. Brief/urls formats append a `[MERGED]` section after the per-angle blocks
- **Bilingual suffixes** — detects whether the topic contains CJK characters. Chinese topics use Chinese suffixes (overview uses empty suffix = topic itself, then 简介/经历/作品/评价); English topics use English suffixes (overview uses empty suffix, then profile/background/works/reviews)
- **`--research-angles`** — custom angles overriding the default 5 (e.g. `"overview,profile,timeline,controversy"`); each angle appended directly as a query suffix
- Mutually exclusive with `--query` and `--queries-file`. Supports all output formats. Deterministic expansion (no AI judgment) — same topic always produces same queries
**Observability**
- **Structured logging** (`--log-format {text,json}`, default text) — `json` outputs one JSON object per line `{ts, level, logger, msg, request_id}` for programmatic parsing
- **Request ID propagation** — each run auto-generates an 8-char hex `request_id` threaded through all logs and progress events
- **`--dry-run`** — preview mode: no HTTP requests sent; prints `{dry_run, instances, headers_count, action, ...}` JSON to stdout (`instances` is a list of instance URLs)
- **`--progress`** — JSON Lines events to stderr with `request_id` for real-time tracking
**Engineering**
- Shared `common.py` module — unified retry policy/charset/auth resolution/logging/UA-pool/browser-headers/backoff/error classification/error codes/recovery hints/Wayback fallback/hard-blocked domains/similarity dedup/UTF-8 enforcement/proxy/Retry-After parsing across both scripts
- `search.py --fetch` reuses `fetch.py`'s higher-quality text extractor (no code duplication)
- Structured logging (`--verbose` / `--quiet`) — three levels: default INFO, `--verbose` DEBUG, `--quiet` WARNING. All log output to stderr; stdout reserved for data
- UTF-8 stdout enforcement (`force_utf8_stdout()`) — Windows Python defaults to GBK and crashes on non-ASCII chars; both scripts force UTF-8 at startup
- Windows config discovery — also checks `%APPDATA%/searxng-cli/` in addition to `~/.config/searxng-cli/`
- Fetch failure diagnostics — error output includes `status_code=`, `cause=`, `url=` fields for programmatic distinction of 404 vs 403 vs DNS failure
- Engine/category whitespace normalization (`"google, bing"``"google,bing"`)
**Scripts + shared module:**
1. `search.py` — execute searches against a user-supplied instance, with multi-instance failover + retry + auto-fetch + caching + batch + domain filtering
2. `fetch.py` — download and extract readable text or markdown from web pages
3. `common.py` — shared utilities (auth resolution + headers, charset detection, retry policy, UA pool, browser headers, backoff, logging, progress events, error classification + codes, recovery hints, Wayback fallback, hard-blocked domains, similarity dedup, UTF-8 enforcement, proxy, Retry-After parsing) used by both scripts
4. `cache.py` — SQLite-backed result cache (SHA-256 key, TTL, WAL mode, size cap + LRU eviction, schema migration, module-level + class API)
5. `_config.py` — package constants (version, schema version, default user agent, UA pool)
## Default settings
`search.py` ships with opinionated defaults tuned for AI research:
| Setting | Default | Flag to override |
|---------|---------|------------------|
| Instance | **required** — via `-i`, `SEARXNG_INSTANCE` env var, or config file | `-i / --instance` |
| Safe search | **0 (off)** | `-s / --safesearch {0,1,2}` |
| Time range | **year** | `-t / --time-range {day,week,month,year,none}` (none = disabled) |
| Output format | **json** | `-f / --format {json,brief,urls,csv}` |
| Engines | **google,bing,brave,duckduckgo,startpage,wikipedia,wikidata** | `--engines <list>` |
## Quick Start
```bash
# Prerequisites: Python 3.8+
# Optional but recommended:
pip install requests beautifulsoup4
# 1. Search against YOUR instance (instance URL is required)
python scripts/search.py -q "python asyncio tutorial" -i https://my-searxng.example.com
# 2. Multiple instances for failover (comma-separated)
python scripts/search.py -q "rust memory safety" \
-i https://a.example.com,https://b.example.com --format brief
# 3. Search + auto-fetch top 3 result pages in one command
python scripts/search.py -q "climate policy" -i https://my-searxng.example.com --fetch 3
# 4. Fetch a result page
python scripts/fetch.py -u "https://example.com" --extract text
# 5. Skip -i by configuring the instance once (env var, current shell)
export SEARXNG_INSTANCE="https://my-searxng.example.com,https://backup.example.com"
python scripts/search.py -q "python asyncio tutorial" # -i not needed
# 6. Or use a config file (./searxng.toml or ~/.config/searxng-cli/searxng.toml
# or %APPDATA%/searxng-cli/searxng.toml on Windows)
# [searxng]
# instance = "https://my-searxng.example.com"
# # or: instances = ["https://a.example.com", "https://b.example.com"]
# # Auth (optional, for private instances):
# auth_basic = "user:password" # Basic auth
# auth_bearer = "sk-token-123" # Bearer token (auth_bearer wins if both set)
# # Any flag below can also be pre-set here (engines, categories, language,
# # safesearch, time_range, method, format, timeout, max_retries, proxy,
# # cache_ttl, cache_max_size, fetch, fetch_timeout, fetch_retries, max_size).
# # Explicit CLI flags always override config values.
# # Auth priority: --auth-* > --auth-*-file > searxng.toml > env var
# Plain list also works in ./instances.txt (one URL per line, # for comments)
# 7. Cache results for 30 minutes (identical queries skip the network)
python scripts/search.py -q "python asyncio" -i https://s.example.com --cache-ttl 30
# 7b. Sort by date (newest first) or disable dedup for raw engine output
python scripts/search.py -q "ai news" -i https://s.example.com --sort-by date --no-dedup
# 8. Batch: run queries from a file (one per line; blank/# lines skipped)
python scripts/search.py --queries-file queries.txt -i https://s.example.com --format json > batch.json
# 9. Domain allowlist + blocklist (applied after search)
python scripts/search.py -q "rust async" -i https://s.example.com \
--include-domain doc.rust-lang.org,wikipedia.org --exclude-domain pinterest.com
# 10. Route through a corporate proxy (applies to search and fetch)
python scripts/search.py -q "ai news" -i https://s.example.com --proxy http://corp-proxy:8080
# 11. Auth from a file (avoids leaking tokens in shell history)
python scripts/search.py -q "test" -i https://private.example.com --auth-bearer-file ~/.searxng_token
# 12. Cache management (no search performed)
python scripts/search.py --cache-stats # entries, oldest/newest, size, total, max, evicted, utilization
python scripts/search.py --clear-cache # delete all entries
# 13. Export results as CSV (great for spreadsheets / data analysis)
python scripts/search.py -q "rust async" -i https://s.example.com --format csv > results.csv
# 14. Use a specific config file (overrides auto-discovered searxng.toml)
python scripts/search.py --config ./my-config.toml -q "test"
# 15. Control log verbosity on stderr
python scripts/search.py -q "test" -i https://s.example.com --verbose # debug detail
python scripts/search.py -q "test" -i https://s.example.com --quiet # errors only
```
**Dependency levels:**
| Level | Scripts | What you get |
|-------|---------|-------------|
| Zero deps (stdlib only) | `search.py`, `_config.py` | Full search + auto-fetch |
| `pip install requests` | `fetch.py` | Better HTTP (session reuse, redirect handling) |
| `pip install beautifulsoup4` | `fetch.py` | Higher-quality text extraction |
## When to Use
- **Web search without API keys** — programmatic search results against your own SearXNG instance
- **Privacy-conscious research** — queries routed through your instance, not ad-tech infrastructure
- **Scraping search results** — batch query multiple terms and collect structured results as JSON
- **Fetching search result pages** — follow links from search results and extract clean text
**Don't use for:**
- High-frequency production search without a properly-scaled instance — respect your instance's rate limits
- Guaranteed uptime/accuracy — depends entirely on the instance you supply
## AI Agent Integration Guide
This section documents the structured interfaces that AI agents can rely on
for programmatic integration. All features are designed to be machine-readable
and machine-actionable.
### Output Channels
| Channel | Content | Description |
|---------|---------|-------------|
| stdout | Data | JSON/CSV/text — the only source AI should parse |
| stderr | Logs + Progress | Human-readable logs (default) or JSON Lines events (`--progress`) |
| exit 0 | Success | Results available on stdout |
| exit 1 | Fatal error | Error JSON on stdout (in `--format json` mode) or stderr |
| exit 2 | Empty results | Search succeeded but returned no results |
### Error Code System
In `--format json` mode, errors are emitted as structured JSON on stdout:
```json
{
"error": "All 2 instances failed. Last error: connection refused",
"error_code": "E_NETWORK",
"recovery_hint": "Retry with backoff, or try a different SearXNG instance. Check network connectivity, proxy settings, and instance uptime.",
"exit_code": 1,
"query": "search term"
}
```
AI agents can use `error_code` to programmatically decide recovery strategy, and `recovery_hint` for a ready-to-use actionable suggestion:
| Code | Meaning | recovery_hint (abridged) |
|------|---------|--------------------------|
| `E_CONFIG` | Configuration error (no instance resolved) | Provide `-i`/`SEARXNG_INSTANCE`/config file |
| `E_AUTH` | Authentication failed (401/403) | Verify credentials, check token expiry & permissions |
| `E_NETWORK` | Network error (connection refused, timeout, 5xx, all instances failed) | Retry with backoff, switch instance, check proxy |
| `E_RATE_LIMIT` | Rate limited (429) | Wait and retry, reduce frequency, distribute load |
| `E_PARSE` | Parse error (JSON/HTML parsing failed) | Try different instance, switch `--method` |
| `E_EMPTY` | Empty results (exit code 2) | Refine query, broaden `--time-range`, add `--categories` |
| `E_INPUT` | Input error (bad parameters, file not found) | Check query syntax, flag combinations, file paths |
| `E_INTERNAL` | Internal error (unexpected exception) | Re-run with `--verbose`, report bug |
| `E_UNSUPPORTED_MEDIA` | Unsupported binary media type (PDF/docx/xlsx parse failure or unsupported content type) | No auto hint — try a different URL, or install `pdftotext` for PDF parsing |
### JSON Lines Streaming (`--stream`)
For large result sets, `--stream` outputs results as JSON Lines (one JSON
object per line) to stdout, allowing AI agents to process results
incrementally without waiting for the full response:
```bash
python scripts/search.py -q "large topic" -i https://your-instance --stream
```
Output format (each line is a separate JSON object):
```
{"type": "result", "result": {"title": "...", "url": "...", "content": "..."}}
{"type": "result", "result": {"title": "...", "url": "...", "content": "..."}}
{"type": "done", "schema_version": "1.0", "count": 2, "query": "large topic"}
```
- `type: "result"` — one per search result, emitted as soon as available
- `type: "done"` — terminal event with total count and `schema_version`, always emitted last
- `type: "error"` — emitted when the search fails, includes `error_code` and `recovery_hint`:
```
{"type": "error", "error": "All 2 instances failed. Last error: HTTP Error 403", "error_code": "E_AUTH", "recovery_hint": "Verify credentials...", "query": "..."}
```
- Exit code 0 on success, 1 on error, 2 on empty results (done event still emitted)
Only valid with `--format json` (single query mode). Using `--stream` with
`--queries-file`, `--research`, or non-json formats raises `E_INPUT` immediately.
### Progress Events (`--progress`)
For long-running operations, `--progress` emits structured JSON Lines events
to stderr, enabling AI agents to track execution progress in real time.
Every event includes a `request_id` field (8-char hex, auto-generated per run)
for correlating all events from a single execution.
```bash
python scripts/search.py -q "research topic" -i https://your-instance --progress --fetch 3
```
Event types (each on its own line, JSON Lines format on stderr):
```jsonl
{"event": "start", "query": "research topic", "instances": 2, "request_id": "a1b2c3d4"}
{"event": "instance_try", "url": "https://instance1.example.com", "attempt": 1, "request_id": "a1b2c3d4"}
{"event": "instance_ok", "url": "https://instance1.example.com", "latency": 0.342, "results": 10, "request_id": "a1b2c3d4"}
{"event": "instance_fail", "url": "https://instance2.example.com", "error": "HTTP 503", "error_code": "E_NETWORK", "request_id": "a1b2c3d4"}
{"event": "cache_hit", "query": "research topic", "ttl": 30, "request_id": "a1b2c3d4"}
{"event": "cache_store", "query": "research topic", "ttl": 30, "request_id": "a1b2c3d4"}
{"event": "page_ok", "pageno": 1, "results": 10, "request_id": "a1b2c3d4"}
{"event": "page_fail", "pageno": 2, "error": "HTTP 503", "request_id": "a1b2c3d4"}
{"event": "fetch_start", "count": 3, "request_id": "a1b2c3d4"}
{"event": "fetch_ok", "url": "https://example.com/page", "chars": 12345, "request_id": "a1b2c3d4"}
{"event": "fetch_fail", "url": "https://bad.example.com", "error": "HTTP 503", "request_id": "a1b2c3d4"}
{"event": "done", "results": 10, "query": "research topic", "request_id": "a1b2c3d4"}
{"event": "error", "error": "connection refused", "error_code": "E_NETWORK", "query": "...", "request_id": "a1b2c3d4"}
```
AI agents can parse these events to:
- Show progress indicators to users
- Detect cache hits (skip waiting)
- Monitor fetch failures and retry strategies
- Correlate errors with specific queries in batch mode
- Track per-page progress via `page_ok`/`page_fail` events
- Correlate all events of a single run via `request_id`
`--progress` and `--verbose` can be used together (progress events on stderr,
debug logs also on stderr). `--progress` events are JSON Lines; `--verbose`
logs are human-readable text.
### Output JSON Schema
The default `--format json` output includes a `schema_version` field so AI
agents can detect breaking changes. Run `--dump-schema` to get the full JSON
Schema document programmatically:
```bash
python scripts/search.py --dump-schema
```
Single query output shape:
```json
{
"schema_version": "1.0",
"query": "search term",
"number_of_results": 10,
"results": [
{
"title": "Result title",
"url": "https://example.com/page",
"content": "Snippet text...",
"engine": "google",
"score": 1.0,
"category": "general",
"published_date": "2024-01-15T10:30:00"
}
],
"answers": ["Direct answer if available"],
"corrections": [],
"suggestions": ["related suggestion"],
"infoboxes": [],
"unresponsive_engines": [["engine_name", "error reason"]],
"fetched": [
{
"url": "https://example.com/page",
"status": "ok",
"text": "Extracted page content...",
"text_length": 12345,
"truncated": false,
"final_url": "https://example.com/final",
"user_agent_used": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
"anti_bot_detected": false,
"waf_type": null,
"fallback_used": null,
"error_code": null
}
],
"fetched_source": "json"
}
```
Batch mode (`--queries-file`) wraps results in a unified schema:
```json
{
"schema_version": "1.0",
"queries": [
{"query": "term1", "status": "ok", "results": {...}},
{"query": "term2", "status": "error", "error": "...", "error_code": "E_NETWORK"}
]
}
```
Each batch entry has a `status` field (`"ok"` or `"error"`). Successful entries
contain `results`; failed entries contain `error` and `error_code`.
Fields marked as optional may be absent. The `fetched` and `fetched_source`
fields only appear when `--fetch N` is used.
## Cross-Agent Compatibility
These scripts are **agent-agnostic** — they work with any AI agent that can invoke terminal commands, or directly from a terminal.
**Key design decisions for universal compatibility:**
- Zero external dependencies (stdlib-only for `search.py`)
- Scripts inject their own directory into `sys.path`, so they run from **any** working directory
- Stdout carries data (JSON/text), stderr carries progress/warnings
- Exit codes: 0=success, 1=fatal error, 2=no results/empty
- NO agent-specific API calls or tool dependencies — purely CLI-based, portable across all agent platforms
## Scripts
All scripts live in `scripts/`; run with `python scripts/<name>.py` from any directory. They import `_config.py` for shared constants and self-inject their own directory into `sys.path`.
### 1. `search.py` — Execute SearXNG Search
```
usage: search.py [-h] [--query QUERY] [--instance URL]
[--categories CATS] [--language LANG] [--pageno N] [--pages N]
[--time-range {day,week,month,year,none}] [--safesearch {0,1,2}]
[--engines E] [--method {GET,POST}] [--max-results N]
[--format {json,brief,urls,csv}] [--snippet-len N] [--stream]
[--progress] [--fetch N] [--fetch-timeout SEC] [--fetch-retries N]
[--fetch-report [json]] [--no-fallback] [--referer URL]
[--request-delay SEC] [--max-size BYTES] [--output FILE]
[--timeout SEC] [--retry N] [--fail-fast] [--serial]
[--verify] [--auth-bearer TOKEN] [--auth-bearer-file FILE]
[--auth-basic USER:PASS] [--auth-basic-file FILE]
[--proxy URL] [--include-domain DOMAINS]
[--exclude-domain DOMAINS] [--queries-file FILE]
[--parallel-queries N] [--research TOPIC] [--research-angles ANGLES]
[--cache-ttl MINUTES] [--cache-max-size MB] [--clear-cache]
[--cache-stats] [--sort-by {score,date,engine,none}] [--no-dedup]
[--similarity-dedup] [--similarity-threshold FLOAT]
[--throttle-failure-threshold N] [--throttle-pause-seconds SEC]
[--throttle-max-delay SEC] [--log-format {text,json}]
[--save-config FILE] [--dry-run]
[--config FILE] [--dump-schema] [--verbose] [--quiet] [--version]
```
**What it does:**
1. Takes a search query and resolves one or more instance URLs (`-i`, `SEARXNG_INSTANCE`, or config file — comma-separated for failover)
2. **Multi-instance failover:** if an instance fails (429/5xx/timeout/captcha), automatically tries the next one
3. **Exponential backoff:** retries each instance up to 3 times with jitter on transient errors (403, 429, 502, 503, 504, and connection errors)
4. **Parallel probing (default for 2+ instances):** queries every instance concurrently and returns the first *successful* result in your original instance order — this keeps output deterministic while drastically speeding up failover when an early instance is down/slow. Use `--serial` to disable.
5. Calls the SearXNG API (GET or POST) with `format=json`
6. Falls back to HTML scraping if JSON is blocked
7. HTML parser extracts results + suggestions + answers + infoboxes (full content, never truncated)
8. **Auto-fetch:** `--fetch 3` concurrently downloads top 3 result pages with retry, full browser fingerprint headers, 15-UA deterministic per-domain pool, Retry-After compliance, WAF fingerprint detection, Wayback Machine fallback for 404/403/timeout, adaptive throttling, and `fetch.py`'s readability-lite extractor. `--fetch-report` prints a structured per-URL report to stderr
9. **Health-check mode:** `--verify` probes each instance (reachability / JSON-API support / latency / POST support / engine list / auth status) and prints a report, then exits without searching — use it to validate your instance list
10. **Result caching:** `--cache-ttl 30` stores results for 30 min; identical queries within the TTL skip the network entirely. Cache lives at `$SEARXNG_CACHE_DIR` or `~/.cache/searxng-cli/cache.db` (SQLite, WAL mode, LRU eviction). `--clear-cache` / `--cache-stats` manage it without searching
11. **Batch mode:** `--queries-file FILE` reads one query per line (blank/`#` lines skipped) and runs them in sequence; JSON output is `{"schema_version": "1.0", "queries": [{"query":..., "status": "ok"|"error", ...}]}` (or one block per query in brief/urls). A failed query is recorded but does not abort the batch. v2.3.0: `--parallel-queries N` (1-8) runs queries concurrently with output order preserved; concurrent mode disables `--fetch` (nested parallel fetch is unsafe). Queries-file encoding is auto-detected (UTF-8, GBK fallback — v2.3.0). Exit codes: 0 if any query returned results, 1 if all errored, 2 if all empty
12. **Dedup + sort + domain filter:** After search (and cache), duplicate URLs are collapsed (default; `--no-dedup` disables), results are sorted (`--sort-by`; default: score descending), and then `--include-domain`/`--exclude-domain` filter by domain. Matching is case-insensitive and ignores a leading `www.`; when a domain is in both lists, exclude wins
13. **Proxy & auth:** `--proxy URL` routes both search and fetch through a proxy; `--auth-bearer` / `--auth-basic` (plus `*-file` variants, `searxng.toml` `auth_basic`/`auth_bearer` fields, and `SEARXNG_BEARER_TOKEN` / `SEARXNG_BASIC_AUTH` env vars) supply credentials. Priority: CLI flag > file > config file > env var
14. **Config defaults:** `searxng.toml` may pre-set most flags; explicit CLI flags always win
15. **Structured errors:** in `--format json` mode, failures print a JSON object `{"error": "...", "error_code": "E_*", "recovery_hint": "...", "exit_code": N, "query": "..."}` to stdout so agents can parse them and decide recovery strategy
**Key options:**
- `--query "your search"` — **required unless** `--verify`, `--queries-file`, `--research`, `--clear-cache`, `--cache-stats`, `--save-config`, or `--dump-schema` is used
- `--instance https://searx.example.org` — **required unless** `SEARXNG_INSTANCE` env var or a config file supplies it; comma-separated list enables failover
- `--queries-file FILE` — read queries from a file (one per line; blank/`#` skipped) and run them in sequence; overrides `--query`
- `--engines google,duckduckgo` — restrict to specific search engines (whitespace around commas is auto-stripped; see default list above)
- `--method POST` — use POST instead of GET (better for long queries)
- `--categories general,news` — comma-separated categories (whitespace auto-stripped)
- `--language zh-CN` — language filter
- `--pageno 1` — page number
- `--pages N` — fetch N pages of results in one run and merge with cross-page dedup; each page cached independently (cache key includes `pageno`)
- `--time-range {day,week,month,year,none}` — time filter (default: `year`; `none` disables time filtering)
- `--safesearch {0,1,2}` — safe search (default: `0` = off)
- `--max-results N` — limit number of results (applied AFTER dedup+sort, so the highest-scoring/newest items are kept)
- `--sort-by {score,date,engine,none}` — sort results (default: `score` descending; `none` preserves instance order). Applied after dedup, before `--max-results`. HTML-fallback results have no score and keep their order
- `--no-dedup` — disable cross-engine deduplication (by default, duplicate URLs — same page ignoring tracking params/fragment — are collapsed, keeping the first occurrence's engine/score)
- `--similarity-dedup` — enable title-based similarity deduplication (SimHash + Jaccard); off by default; auto-skipped when results > 500
- `--similarity-threshold FLOAT` — similarity threshold for `--similarity-dedup` (default: 0.85; higher = stricter)
- `--config FILE` — path to a `searxng.toml` config file; overrides the default auto-discovery (`./searxng.toml` → `~/.config/searxng-cli/searxng.toml` → `%APPDATA%/searxng-cli/searxng.toml` on Windows). Must be the first flag so its values can set defaults for other flags
- `--verbose` / `-v` — show debug-level diagnostics on stderr (HTTP request URLs, response codes, cache keys, retry detail)
- `--quiet` — suppress progress messages and retry notices on stderr; only warnings and errors are shown (no short flag: `-q` is `--query`)
- `--include-domain a.com,b.org` — allowlist; only results from these domains are kept (applied after search)
- `--exclude-domain pinterest.com` — blocklist; results from these domains are dropped (applied after search; wins over include on conflict)
- `--serial` — disable parallel multi-instance probing; search instances strictly one at a time
- `--verify` — health-check mode: verify instances and exit (no search); combine with `--format brief` for a table or `--format json` for machine-readable output
- `--format json` — full JSON (default); includes `fetched` array when `--fetch` is used; errors are emitted as JSON to stdout
- `--format brief` — title + URL + full snippet (no truncation by default; use `--snippet-len 200` to cap)
- `--format urls` — only result URLs
- `--format csv` — CSV export (title,url,engine,score,published_date,content); in batch mode (`--queries-file`), all queries merge into one CSV with a `query` column
- `--fetch N` — after search, auto-fetch full text of top N result pages (concurrent, stdlib only)
- `--fetch-timeout 10` — timeout per page fetch (default: 10s)
- `--fetch-retries 3` — max retries per page fetch
- `--max-size BYTES` — cap page size (default: unlimited; e.g. `5242880` for 5MB)
- `--retry 5` — max retries per instance (default: 3)
- `--timeout 15` — request timeout in seconds
- `--fail-fast` — use only the first instance, don't fail over to the rest
- `--proxy URL` — HTTP/HTTPS proxy for both search and fetch (e.g. `http://corp-proxy:8080`)
- `--cache-ttl MINUTES` — cache results for N minutes (default: `0` = disabled); identical queries within the TTL skip the network
- `--cache-max-size MB` — cache size cap in MB with LRU eviction (default: `0` = defer to `$SEARXNG_CACHE_MAX_SIZE_BYTES` env var or 100MB built-in default; when cap exceeded, least-recently-accessed entries evicted; set a very large value for effectively unlimited)
- `--clear-cache` — delete all cached entries and exit (no search)
- `--cache-stats` — print cache statistics (entries, oldest_created_at, newest_created_at, path, size_bytes, total_bytes, max_size_bytes, evicted_count, utilization_pct) and exit
- `--dump-schema` — print the JSON Schema for `--format json` output and exit; lets AI agents programmatically discover field names and types without parsing prose docs
- `--auth-bearer TOKEN` — `Authorization: Bearer` header for private instances
- `--auth-bearer-file FILE` — read Bearer token from a file (first non-empty, non-`#` line); also honors `SEARXNG_BEARER_TOKEN` env var
- `--auth-basic USER:PASS` — `Authorization: Basic` header (auto base64-encoded)
- `--auth-basic-file FILE` — read `user:pass` from a file (first non-empty, non-`#` line); also honors `SEARXNG_BASIC_AUTH` env var
- `--log-format {text,json}` — structured logging format (default: `text`); `json` outputs one JSON object per line `{ts, level, logger, msg, request_id}`
- `--dry-run` — preview mode: no HTTP requests; prints `{dry_run, instances, headers_count, action, ...}` JSON to stdout (`instances` is a list); supports search/research/batch/verify
- `--throttle-failure-threshold N` — consecutive failures before doubling delay + halving concurrency (default: 3)
- `--throttle-pause-seconds SEC` — global pause on 429 (default: 30)
- `--throttle-max-delay SEC` — max adaptive delay (default: 10)
- `--research-angles "a,b,c"` — custom research angles overriding the default 5; each angle appended directly as a query suffix
- `--save-config FILE` — save current CLI args as a `searxng.toml` config file and exit
- `--fetch-report [json]` — `--fetch-report json` outputs a full JSON report (items + summary); original `--fetch-report` (no arg) keeps text format
- `--version` — print version and exit
**Completion criterion:** Outputs valid JSON with `results` array. Non-zero exit on total failure (all instances exhausted).
### 2. `fetch.py` — Fetch & Extract Web Page Content
```
usage: fetch.py [-h] --url URL [--extract {text,html,markdown}]
[--format {text,json}]
[--timeout SEC] [--retries N] [--max-size BYTES]
[--user-agent STR] [--encoding CHARSET]
[--no-redirect] [--no-fallback] [--referer URL] [--proxy URL]
[--output FILE] [--auth-bearer TOKEN] [--auth-bearer-file FILE]
[--auth-basic USER:PASS] [--auth-basic-file FILE]
[--verbose] [--quiet] [--version]
```
**What it does:**
1. Downloads a web page via HTTP GET with retry + exponential backoff
2. Retries on 429/5xx/connection errors (3x default) with Retry-After header compliance and capped 60s backoff; falls back through the 15-UA deterministic pool with full browser fingerprint headers; reuses a requests.Session for connection pooling + cookie persistence. **Both paths handle Content-Encoding decompression**: stdlib urllib path handles gzip/deflate/br (urllib doesn't auto-decompress any of them); requests path handles br manually (requests auto-decompresses gzip/deflate but not Brotli unless the brotli package is installed)
3. **No size limit by default** — full page content returned; use `--max-size` for a cap
4. Detects charset from HTTP headers, HTML meta tags, or UTF-8 fallback
5. Extracts readable content using tree-based parsers (stdlib or BeautifulSoup); readability-lite text-density fallback when `<article>`/`<main>`/`role="main"`/content-class `<div>` are all missing and only `<body>` remains
6. Outputs clean text, raw HTML, or properly-converted Markdown (with correct nested-link handling)
7. **PDF/document parsing** (regardless of `--extract` mode): PDF via `pdftotext` subprocess, `.docx`/`.xlsx` via stdlib `zipfile`. Unsupported binary types return `E_UNSUPPORTED_MEDIA`
**Key options:**
- `--url https://...` — required
- `--format {text,json}` — output format (default: `text`; `json` = structured contract, v2.3.0). Success: `{status: ok, url, final_url, content_type, extract, truncated, text_length, user_agent}`; failure: `{status: error, url, error, error_code, status_code}`. stdout is pure JSON; logs stay on stderr; exit 0 on success / 1 on failure
- `--extract text` — clean readable text (default); PDF/docx/xlsx parsed automatically regardless of extract mode (see "What it does" #7)
- `--extract html` — raw HTML
- `--extract markdown` — Markdown conversion (tree-based; handles nested tags, GFM tables, fenced code blocks, blockquotes, inline code, ordered/unordered/nested lists, definition lists, images, emphasis)
- `--encoding gbk` — force charset for non-UTF-8 pages (else auto-detected from HTTP header / HTML meta / UTF-8 fallback)
- `--timeout 15` — request timeout in seconds (split into connect/read tuple internally)
- `--retries 3` — max retries on transient errors (429/5xx/connection); falls back through the 15-UA pool when blocked
- `--max-size BYTES` — cap page size (default: unlimited; e.g. `5242880` for 5MB)
- `--user-agent STR` — custom User-Agent header (overrides per-domain UA selection)
- `--no-redirect` — do **not** follow HTTP 3xx redirects (implemented for both the requests and stdlib paths)
- `--no-fallback` — disable Wayback Machine fallback for 404/403/timeout and anti-bot-blocked pages (hard-blocked domains like baike.baidu.com, zhihu.com still prefer Wayback)
- `--referer URL` — set Referer header to disguise traffic source
- `--proxy URL` — HTTP/HTTPS proxy (e.g. `http://corp-proxy:8080`); respects existing `HTTP_PROXY`/`HTTPS_PROXY` env vars when omitted
- `--output FILE` — save to file instead of stdout
- `--auth-bearer TOKEN` / `--auth-bearer-file FILE` — `Authorization: Bearer` header; file variant reads first non-empty, non-`#` line; also honors `SEARXNG_BEARER_TOKEN` env var
- `--auth-basic USER:PASS` / `--auth-basic-file FILE` — `Authorization: Basic` header (auto base64-encoded); file variant + `SEARXNG_BASIC_AUTH` env var also supported
- `--verbose` / `-v` — show debug-level diagnostics on stderr
- `--quiet` — suppress progress messages on stderr; only warnings and errors are shown
- `--version` — print version and exit
**Extraction strategy (text mode):**
1. Strip non-content elements (script, style, nav, footer, header)
2. Extract main content in priority order: `<article>` → `<main>` → `role="main"` → content-class `<div>` (class~=`content`/`article`/`post`/`entry`) → `<body>`
3. If only `<body>` matched (all higher-priority elements missing), run readability-lite to pick the highest text-density `<div>`/`<section>`/`<article>` node (scored by text density + `<p>` count weighting; filters out nav/sidebar/footer by class/id)
4. Collapse whitespace, output clean UTF-8
5. Warn if extracted text < 500 chars (likely JS-heavy or bot-blocked)
**Completion criterion:** Outputs page content (or the structured JSON object with `--format json`). Non-zero exit on HTTP failure. Stderr carries warnings for low-confidence extraction.