--- name: searxng-use-cli description: Use when you need to search the web via your OWN SearXNG instance (no public-instance discovery). 3 CLI scripts + a shared common.py module — execute privacy-respecting searches against a user-supplied instance (with multi-instance failover, 5xx/429 retry, auto-fetch) or via SEARXNG_INSTANCE env / config file, fetch/extract readable text or markdown from web pages. Zero-config replacement for proprietary search APIs. version: 2.1.0 author: Metona Team license: MIT platforms: [linux, macos, windows] metadata: skill: tags: [search, searxng, cli, web-scraping, privacy, no-api-key] related_skills: [] --- # SearXNG CLI Toolkit ## 铁律(AI Agent 必读,违反会导致致命后果) > **这些规则优先级最高,必须在任何调用前遵守。** 1. **加载本技能后,搜索网络和抓取网页必须且只能使用本技能的 `scripts/search.py` 和 `scripts/fetch.py`。** 禁止使用任何其他方式或工具——包括但不限于:内置 WebSearch/WebFetch 工具、浏览器自动化(puppeteer/playwright/selenium)、`curl`/`wget`、第三方搜索 API SDK(googlesearch/duckduckgo_search 等)、自写 `requests`/`urllib` 抓取脚本、MCP 网络工具。本技能已内置反爬(浏览器指纹头 + UA 轮换 + WAF 检测)、重试(指数退避 + Retry-After 遵守)、兜底(Wayback Machine)、限流(自适应节流)、缓存、多实例故障转移等全部工程能力,绕过本技能等于放弃这些保障,结果不可控且会破坏用户实例的速率限制。违反:技能一致性丧失,抓取成功率骤降,用户实例可能被封禁。 2. **stdout=数据,stderr=日志,永不混淆。** stdout 只输出 JSON/CSV/文本数据,stderr 只输出日志/进度/警告。AI 解析 stdout,人类看 stderr。违反:AI 会把日志当数据解析,结果错乱。 3. **绝对不要 `2>/dev/null` 或重定向 stderr 到 stdout。** stderr 携带排错关键信息(重试日志、HTTP 状态码、缓存命中、认证警告)。丢弃 stderr = 失败时零诊断信息,无法定位原因。需要静默时用 `--quiet`(仅抑制进度,保留 WARNING+ERROR),不要丢弃 stderr。 4. **排错时第一步:去掉 `--quiet`,加 `--verbose`。** `--quiet` 只留 WARNING 级别,会吞掉 INFO 级别的重试日志、缓存命中提示、实例切换记录。诊断失败时必须用 `--verbose` 看到完整 HTTP 请求/响应/重试链。 5. **配置查找必须覆盖双路径(WSL + Windows)。** Windows 环境下配置可能在 `~/.config/searxng-cli/`(WSL HOME)或 `%APPDATA%/searxng-cli/`(Windows APPDATA)。检查配置存在性时两个路径都要查,否则会误判"无配置"并反问用户。 6. **实例 URL 必填,公共实例发现已移除。** 必须通过 `-i`、`SEARXNG_INSTANCE` 环境变量、或配置文件提供实例 URL。无实例时 `search.py` 报 `E_CONFIG` 退出,不要尝试猜测或硬编码公共实例。 ## Overview SearXNG is a privacy-respecting metasearch engine that aggregates results from 70+ search services without tracking users. This skill provides three standalone Python CLI scripts — works with **any AI agent** (Hermes, Claude Code, Codex, OpenCode, Cursor, Trae, etc.) or directly from your terminal. **Public-instance discovery has been removed.** You must supply your own SearXNG instance URL (self-hosted or one you trust). This makes behavior deterministic and avoids depending on volatile public instances. **Key capabilities:** **Search & results** - Multi-instance failover with parallel probing (faster failover, deterministic output order) - Exponential-backoff retry on transient errors (429/5xx/connection) via shared `common.py` - `--verify` health-check mode (reachability / JSON-API / latency / POST / engine list / auth status) - Cross-engine result deduplication (default on; `--no-dedup` disables) — collapses duplicate URLs ignoring tracking params (`utm_*`, `gclid`, etc.) and fragments - Result sorting (`--sort-by {score,date,engine,none}`; default: score descending) — applied after dedup, before `--max-results` - Domain allowlist/blocklist (`--include-domain` / `--exclude-domain`) — case-insensitive, ignores leading `www.`, exclude wins on conflict - Batch mode (`--queries-file`) — run multiple queries from a file in sequence, combined output **Output formats** - JSON (default, rich metadata), brief (title+URL+snippet), urls (plain list), CSV (spreadsheet-ready) - Enhanced Markdown conversion — nested ordered lists (numbered), mixed `ul`/`ol` nesting, `
`/`
`/`
` definition lists, GFM tables, fenced code blocks, blockquotes - Structured JSON error output (in `--format json` mode) with `error_code` field for machine-readable failure reporting - JSON Lines streaming (`--stream`) — each result emitted as a separate JSON line to stdout, enabling incremental processing by AI agents - Progress events (`--progress`) — structured JSON Lines events to stderr for real-time execution tracking **Caching & config** - SQLite result caching (`--cache-ttl`) — identical queries within a TTL skip the network entirely; `--clear-cache` / `--cache-stats` manage it - Config file (`searxng.toml`) pre-sets most flags; `--config FILE` loads a non-default config; `instances.txt` for plain URL lists - Instance resolution priority: `-i` → `SEARXNG_INSTANCE` env → config file (`./searxng.toml` → `~/.config/searxng-cli/searxng.toml` → `%APPDATA%/searxng-cli/searxng.toml` on Windows) **Network & auth** - Proxy support (`--proxy`) for both search and fetch (sets `HTTP_PROXY`/`HTTPS_PROXY`/`NO_PROXY`) - Auth via CLI flag, file, config file, or env var (`--auth-bearer` / `--auth-basic` + `*-file` variants) to avoid leaking secrets in shell history - Config-file auth — `auth_basic` and `auth_bearer` fields in `searxng.toml` let AI agents set credentials once (priority: CLI > file > config > env) - Credentials-file permission warning — `--auth-*-file` warns on stderr if the file is group/other-readable (POSIX only) **Anti-bot & fetch stability (v2.0.0)** - **Browser fingerprint headers** — `build_browser_headers()` sends full Sec-Ch-Ua / Sec-Fetch-* / Accept-Language / Accept-Encoding, not just User-Agent. Bypasses 80%+ of lightweight WAFs (Cloudflare basic, Nginx UA blocks) - **12-UA pool** — Chrome/Edge/Firefox × Windows/macOS/Linux × versions 129-131. `get_ua_for_domain()` uses SHA-256 to deterministically assign one UA per domain (stable within a session, reproducible across processes) - **requests.Session reuse** — module-level Session with connection pool (10 conns/host) + cookie persistence + TLS session resumption. Cuts TLS handshake overhead for multi-page fetches - **Split timeouts** — `timeout=(connect, read)` tuple (5s connect, 15s read by default). Previously a single 15s total timeout wasted already-established connections on large pages - **Retry-After compliance** — 429/503 responses read the `Retry-After` header (numeric seconds or HTTP date) and wait at least that long before retrying. Non-compliance triggers harsher rate limits - **Capped backoff** — `compute_backoff_delay()` caps at 60s (was uncapped: 1.5*2^10 = 1536s would hang the process) - **WAF fingerprint library** — `_detect_anti_bot()` identifies Cloudflare / Imperva / PerimeterX / DataDome / Akamai / Anubis / generic challenges via `` tag matching (most precise) + full-document scan of technical identifiers (cookie/header/JS variable names, not bare vendor names). v2.0.1 narrowed broad keywords (e.g. bare `cloudflare`/`captcha`/`challenge`) that caused false positives on normal articles, and added title-tag detection. Returns `waf_type` for AI-agent decisioning - **Wayback Machine fallback** — 404/403/timeout AND anti-bot-blocked pages automatically retry via `https://web.archive.org/web/2/<url>` (latest snapshot). v2.0.1 fixed a bug where HTTP 200 anti-bot challenge pages (Cloudflare returns 200 for JS challenges) bypassed the fallback trigger. v2.1.0: `fetch.py` standalone calls now also get Wayback fallback (was only in `search.py --fetch`). Default ON; `--no-fallback` disables. Independent 10s timeout so Wayback slowness never blocks the main flow - **Hard-blocked domain fallback** (v2.1.0) — `is_hard_blocked_domain()` detects known strong-anti-bot sites (baike.baidu.com, zhihu.com, weibo.com, mp.weixin.qq.com, douban.com, etc.) that almost always return 403 regardless of UA. These sites bypass the normal `should_try_wayback()` check and trigger Wayback immediately on failure. List maintained in `common.py` `HARD_BLOCKED_DOMAINS` - **Adaptive throttling** — `AdaptiveThrottle` state machine: 3 consecutive failures → double delay + halve concurrency; 5 consecutive successes → gradual recovery; 429 → global pause 30s. Thread-safe - **readability-lite extraction** — when `<article>`/`<main>`/content-class `<div>` are all missing, `_readability_lite()` picks the highest text-density node (text chars / tag count + `<p>` weighting), avoiding nav/sidebar/footer noise - **`--fetch-report`** — structured per-URL report to stderr after `--fetch N`: status, WAF type, fallback used, char count, plus adaptive throttle stats and a JSON summary line - **`--referer`** — set Referer header for fetch requests (defaults to the instance URL when fetching result pages, disguising traffic source) - **`--request-delay`** — configurable delay between fetch requests (default 0.3s; adaptive throttling may increase this on failures) - New fetch result fields: `anti_bot_detected` (bool), `waf_type` (str|null), `fallback_used` (str|null) **Research mode (v2.1.0)** - **`--research <topic>`** — given a research topic, auto-expands into 5 multi-angle queries (overview/profile/background/works/review) and runs them in sequence. Results include `research_topic` and `research_queries` metadata so AI agents can structure their final report by angle. Mutually exclusive with `--query` and `--queries-file`. Supports all output formats (json/brief/urls/csv). Deterministic expansion (no AI judgment) — same topic always produces same queries, reproducible across processes **Engineering** - Shared `common.py` module — unified retry/charset/auth/logging/UA-pool/browser-headers/backoff logic across both scripts - `search.py --fetch` reuses `fetch.py`'s higher-quality text extractor (no code duplication) - Structured logging (`--verbose` / `--quiet`) — three levels: default INFO (progress + warnings), `--verbose` DEBUG (HTTP detail, cache keys), `--quiet` WARNING (errors only). All log output to stderr; stdout reserved for data - UTF-8 stdout enforcement (`force_utf8_stdout()`) — Windows Python defaults to GBK and crashes on non-ASCII chars; both scripts force UTF-8 + `errors='replace'` at startup so `print('\xa0')` never raises - Windows config discovery — `resolve_instances` / `load_config` also check `%APPDATA%/searxng-cli/` (Windows per-user app convention) in addition to `~/.config/searxng-cli/` (POSIX convention) - fetch.py failure diagnostics — error output includes `status_code=`, `cause=`, `url=` fields so AI agents can programmatically distinguish 404 vs 403 vs DNS failure without parsing English prose - Engine/category whitespace normalization (`"google, bing"` → `"google,bing"`) - `--time-range none` option to disable time filtering **Scripts + shared module:** 1. `search.py` — execute searches against a user-supplied instance, with multi-instance failover + exponential-backoff retry (429/5xx/connection) + auto-fetch + caching + batch + domain filtering 2. `fetch.py` — download and extract readable text or markdown from web pages 3. `common.py` — shared utilities (auth headers, charset detection, retry policy, fallback UAs, retry constants) used by both scripts 4. `cache.py` — SQLite-backed result cache (SHA-256 key, TTL, WAL mode) 5. `_config.py` — package constants (version, User-Agent) ## Default settings `search.py` ships with opinionated defaults tuned for AI research: | Setting | Default | Flag to override | |---------|---------|------------------| | Instance | **required** — via `-i`, `SEARXNG_INSTANCE` env var, or config file | `-i / --instance` | | Safe search | **0 (off)** | `-s / --safesearch {0,1,2}` | | Time range | **year** | `-t / --time-range {day,month,year,none}` (none = disabled) | | Output format | **json** | `-f / --format {json,brief,urls,csv}` | | Engines | **google,bing,brave,duckduckgo,startpage,wikipedia,wikidata** | `--engines <list>` | ## Quick Start ```bash # Prerequisites: Python 3.8+ # Optional but recommended: pip install requests beautifulsoup4 # 1. Search against YOUR instance (instance URL is required) python scripts/search.py -q "python asyncio tutorial" -i https://my-searxng.example.com # 2. Multiple instances for failover (comma-separated) python scripts/search.py -q "rust memory safety" \ -i https://a.example.com,https://b.example.com --format brief # 3. Search + auto-fetch top 3 result pages in one command python scripts/search.py -q "climate policy" -i https://my-searxng.example.com --fetch 3 # 4. Fetch a result page python scripts/fetch.py -u "https://example.com" --extract text # 5. Skip -i by configuring the instance once (env var, current shell) export SEARXNG_INSTANCE="https://my-searxng.example.com,https://backup.example.com" python scripts/search.py -q "python asyncio tutorial" # -i not needed # 6. Or use a config file (./searxng.toml or ~/.config/searxng-cli/searxng.toml # or %APPDATA%/searxng-cli/searxng.toml on Windows) # [searxng] # instance = "https://my-searxng.example.com" # # or: instances = ["https://a.example.com", "https://b.example.com"] # # Auth (optional, for private instances): # auth_basic = "user:password" # Basic auth # auth_bearer = "sk-token-123" # Bearer token (auth_bearer wins if both set) # # Any flag below can also be pre-set here (engines, categories, language, # # safesearch, time_range, method, format, timeout, max_retries, proxy, # # cache_ttl, fetch, fetch_timeout, fetch_retries, max_size). # # Explicit CLI flags always override config values. # # Auth priority: --auth-* > --auth-*-file > searxng.toml > env var # Plain list also works in ./instances.txt (one URL per line, # for comments) # 7. Cache results for 30 minutes (identical queries skip the network) python scripts/search.py -q "python asyncio" -i https://s.example.com --cache-ttl 30 # 7b. Sort by date (newest first) or disable dedup for raw engine output python scripts/search.py -q "ai news" -i https://s.example.com --sort-by date --no-dedup # 8. Batch: run queries from a file (one per line; blank/# lines skipped) python scripts/search.py --queries-file queries.txt -i https://s.example.com --format json > batch.json # 9. Domain allowlist + blocklist (applied after search) python scripts/search.py -q "rust async" -i https://s.example.com \ --include-domain doc.rust-lang.org,wikipedia.org --exclude-domain pinterest.com # 10. Route through a corporate proxy (applies to search and fetch) python scripts/search.py -q "ai news" -i https://s.example.com --proxy http://corp-proxy:8080 # 11. Auth from a file (avoids leaking tokens in shell history) python scripts/search.py -q "test" -i https://private.example.com --auth-bearer-file ~/.searxng_token # 12. Cache management (no search performed) python scripts/search.py --cache-stats # entry count, age, size, path python scripts/search.py --clear-cache # delete all entries # 13. Export results as CSV (great for spreadsheets / data analysis) python scripts/search.py -q "rust async" -i https://s.example.com --format csv > results.csv # 14. Use a specific config file (overrides auto-discovered searxng.toml) python scripts/search.py --config ./my-config.toml -q "test" # 15. Control log verbosity on stderr python scripts/search.py -q "test" -i https://s.example.com --verbose # debug detail python scripts/search.py -q "test" -i https://s.example.com --quiet # errors only ``` **Dependency levels:** | Level | Scripts | What you get | |-------|---------|-------------| | Zero deps (stdlib only) | `search.py`, `_config.py` | Full search + auto-fetch | | `pip install requests` | `fetch.py` | Better HTTP (session reuse, redirect handling) | | `pip install beautifulsoup4` | `fetch.py` | Higher-quality text extraction | ## When to Use - **Web search without API keys** — programmatic search results against your own SearXNG instance - **Privacy-conscious research** — queries routed through your instance, not ad-tech infrastructure - **Scraping search results** — batch query multiple terms and collect structured results as JSON - **Fetching search result pages** — follow links from search results and extract clean text **Don't use for:** - High-frequency production search without a properly-scaled instance — respect your instance's rate limits - Guaranteed uptime/accuracy — depends entirely on the instance you supply ## AI Agent Integration Guide This section documents the structured interfaces that AI agents can rely on for programmatic integration. All features are designed to be machine-readable and machine-actionable. ### Output Channels | Channel | Content | Description | |---------|---------|-------------| | stdout | Data | JSON/CSV/text — the only source AI should parse | | stderr | Logs + Progress | Human-readable logs (default) or JSON Lines events (`--progress`) | | exit 0 | Success | Results available on stdout | | exit 1 | Fatal error | Error JSON on stdout (in `--format json` mode) or stderr | | exit 2 | Empty results | Search succeeded but returned no results | ### Error Code System In `--format json` mode, errors are emitted as structured JSON on stdout: ```json { "error": "All 2 instances failed. Last error: connection refused", "error_code": "E_NETWORK", "recovery_hint": "Retry with backoff, or try a different SearXNG instance. Check network connectivity, proxy settings, and instance uptime.", "exit_code": 1, "query": "search term" } ``` AI agents can use `error_code` to programmatically decide recovery strategy, and `recovery_hint` for a ready-to-use actionable suggestion: | Code | Meaning | recovery_hint (abridged) | |------|---------|--------------------------| | `E_CONFIG` | Configuration error (no instance resolved) | Provide `-i`/`SEARXNG_INSTANCE`/config file | | `E_AUTH` | Authentication failed (401/403) | Verify credentials, check token expiry & permissions | | `E_NETWORK` | Network error (connection refused, timeout, 5xx, all instances failed) | Retry with backoff, switch instance, check proxy | | `E_RATE_LIMIT` | Rate limited (429) | Wait and retry, reduce frequency, distribute load | | `E_PARSE` | Parse error (JSON/HTML parsing failed) | Try different instance, switch `--method` | | `E_EMPTY` | Empty results (exit code 2) | Refine query, broaden `--time-range`, add `--categories` | | `E_INPUT` | Input error (bad parameters, file not found) | Check query syntax, flag combinations, file paths | | `E_INTERNAL` | Internal error (unexpected exception) | Re-run with `--verbose`, report bug | ### JSON Lines Streaming (`--stream`) For large result sets, `--stream` outputs results as JSON Lines (one JSON object per line) to stdout, allowing AI agents to process results incrementally without waiting for the full response: ```bash python scripts/search.py -q "large topic" -i https://your-instance --stream ``` Output format (each line is a separate JSON object): ``` {"type": "result", "result": {"title": "...", "url": "...", "content": "..."}} {"type": "result", "result": {"title": "...", "url": "...", "content": "..."}} {"type": "done", "schema_version": "1.0", "count": 2, "query": "large topic"} ``` - `type: "result"` — one per search result, emitted as soon as available - `type: "done"` — terminal event with total count and `schema_version`, always emitted last - `type: "error"` — emitted when the search fails, includes `error_code` and `recovery_hint`: ``` {"type": "error", "error": "All 2 instances failed. Last error: HTTP Error 403", "error_code": "E_AUTH", "recovery_hint": "Verify credentials...", "query": "..."} ``` - Exit code 0 on success, 1 on error, 2 on empty results (done event still emitted) Only valid with `--format json` (single query mode). Using `--stream` with `--queries-file` or non-json formats raises `E_INPUT` immediately. ### Progress Events (`--progress`) For long-running operations, `--progress` emits structured JSON Lines events to stderr, enabling AI agents to track execution progress in real time: ```bash python scripts/search.py -q "research topic" -i https://your-instance --progress --fetch 3 ``` Event types (each on its own line, JSON Lines format on stderr): ```jsonl {"event": "start", "query": "research topic", "instances": 2} {"event": "instance_try", "url": "https://instance1.example.com", "attempt": 1} {"event": "instance_ok", "url": "https://instance1.example.com", "latency": 0.342, "results": 10} {"event": "instance_fail", "url": "https://instance2.example.com", "error": "HTTP 503", "error_code": "E_NETWORK"} {"event": "cache_hit", "query": "research topic", "ttl": 30} {"event": "cache_store", "query": "research topic", "ttl": 30} {"event": "fetch_start", "count": 3} {"event": "fetch_ok", "url": "https://example.com/page", "chars": 12345} {"event": "fetch_fail", "url": "https://bad.example.com", "error": "HTTP 503"} {"event": "done", "results": 10, "query": "research topic"} {"event": "error", "error": "connection refused", "error_code": "E_NETWORK", "query": "..."} ``` AI agents can parse these events to: - Show progress indicators to users - Detect cache hits (skip waiting) - Monitor fetch failures and retry strategies - Correlate errors with specific queries in batch mode `--progress` and `--verbose` can be used together (progress events on stderr, debug logs also on stderr). `--progress` events are JSON Lines; `--verbose` logs are human-readable text. ### Output JSON Schema The default `--format json` output includes a `schema_version` field so AI agents can detect breaking changes. Run `--dump-schema` to get the full JSON Schema document programmatically: ```bash python scripts/search.py --dump-schema ``` Single query output shape: ```json { "schema_version": "1.0", "query": "search term", "number_of_results": 10, "results": [ { "title": "Result title", "url": "https://example.com/page", "content": "Snippet text...", "engine": "google", "score": 1.0, "category": "general", "published_date": "2024-01-15T10:30:00" } ], "answers": ["Direct answer if available"], "corrections": [], "suggestions": ["related suggestion"], "infoboxes": [], "unresponsive_engines": [["engine_name", "error reason"]], "fetched": [ { "url": "https://example.com/page", "status": "ok", "text": "Extracted page content...", "text_length": 12345, "truncated": false, "final_url": "https://example.com/final", "user_agent_used": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36", "anti_bot_detected": false, "waf_type": null, "fallback_used": null } ], "fetched_source": "json" } ``` Batch mode (`--queries-file`) wraps results in a unified schema: ```json { "schema_version": "1.0", "queries": [ {"query": "term1", "status": "ok", "results": {...}}, {"query": "term2", "status": "error", "error": "...", "error_code": "E_NETWORK"} ] } ``` Each batch entry has a `status` field (`"ok"` or `"error"`). Successful entries contain `results`; failed entries contain `error` and `error_code`. Fields marked as optional may be absent. The `fetched` and `fetched_source` fields only appear when `--fetch N` is used. ## Cross-Agent Compatibility These scripts are **agent-agnostic** — they work with any AI agent that can invoke terminal commands (Hermes, Claude Code, Codex, OpenCode, Cursor, Trae, etc.), or directly from a terminal. **Key design decisions for universal compatibility:** - Zero external dependencies (stdlib-only for `search.py`) - Scripts inject their own directory into `sys.path`, so they run from **any** working directory - Stdout carries data (JSON/text), stderr carries progress/warnings - Exit codes: 0=success, 1=fatal error, 2=no results/empty - NO agent-specific API calls or tool dependencies — purely CLI-based, portable across all agent platforms ## Scripts All scripts live in `scripts/`; run with `python scripts/<name>.py` from any directory. They import `_config.py` for shared constants and self-inject their own directory into `sys.path`. ### 1. `search.py` — Execute SearXNG Search ``` usage: search.py [-h] [--query QUERY] [--instance URL] [--categories CATS] [--language LANG] [--pageno N] [--time-range {day,month,year,none}] [--safesearch {0,1,2}] [--engines E] [--method {GET,POST}] [--max-results N] [--format {json,brief,urls,csv}] [--snippet-len N] [--fetch N] [--fetch-timeout SEC] [--fetch-retries N] [--fetch-report] [--no-fallback] [--referer URL] [--request-delay SEC] [--max-size BYTES] [--output FILE] [--timeout SEC] [--retry N] [--fail-fast] [--serial] [--verify] [--auth-bearer TOKEN] [--auth-bearer-file FILE] [--auth-basic USER:PASS] [--auth-basic-file FILE] [--proxy URL] [--include-domain DOMAINS] [--exclude-domain DOMAINS] [--queries-file FILE] [--cache-ttl MINUTES] [--clear-cache] [--cache-stats] [--sort-by {score,date,engine,none}] [--no-dedup] [--config FILE] [--dump-schema] [--verbose] [--quiet] [--version] ``` **What it does:** 1. Takes a search query and resolves one or more instance URLs (`-i`, `SEARXNG_INSTANCE`, or config file — comma-separated for failover) 2. **Multi-instance failover:** if an instance fails (429/5xx/timeout/captcha), automatically tries the next one 3. **Exponential backoff:** retries each instance up to 3 times with jitter on transient errors (429, 502, 503, 504, and connection errors) 4. **Parallel probing (default for 2+ instances):** queries every instance concurrently and returns the first *successful* result in your original instance order — this keeps output deterministic while drastically speeding up failover when an early instance is down/slow. Use `--serial` to disable. 5. Calls the SearXNG API (GET or POST) with `format=json` 6. Falls back to HTML scraping if JSON is blocked 7. HTML parser extracts results + suggestions + answers + infoboxes (full content, never truncated) 8. **Stable auto-fetch (v2.0.0 enhanced):** `--fetch 3` concurrently downloads top 3 result pages with retry, **full browser fingerprint headers** (Sec-Ch-Ua/Sec-Fetch-*), **12-UA deterministic per-domain pool**, **Retry-After compliance**, **WAF fingerprint detection** (Cloudflare/Imperva/PerimeterX/DataDome/Akamai), **Wayback Machine fallback** for 404/403/timeout, **adaptive throttling** (auto backoff + concurrency reduction on failures), and `fetch.py`'s readability-lite extractor. `--fetch-report` prints a structured per-URL report to stderr 9. **Health-check mode:** `--verify` probes each instance (reachability / JSON-API support / latency / POST support / engine list / auth status) and prints a report, then exits without searching — use it to validate your instance list 10. **Result caching:** `--cache-ttl 30` stores results for 30 min; identical queries within the TTL skip the network entirely. Cache lives at `$SEARXNG_CACHE_DIR` or `~/.cache/searxng-cli/cache.db` (SQLite, WAL mode). `--clear-cache` / `--cache-stats` manage it without searching 11. **Batch mode:** `--queries-file FILE` reads one query per line (blank/`#` lines skipped) and runs them in sequence; JSON output is `{"schema_version": "1.0", "queries": [{"query":..., "status": "ok"|"error", ...}]}` (or one block per query in brief/urls). A failed query is recorded but does not abort the batch. Exit codes: 0 if any query returned results, 1 if all errored, 2 if all empty 12. **Dedup + sort + domain filter:** After search (and cache), duplicate URLs are collapsed (default; `--no-dedup` disables), results are sorted (`--sort-by`; default: score descending), and then `--include-domain`/`--exclude-domain` filter by domain. Matching is case-insensitive and ignores a leading `www.`; when a domain is in both lists, exclude wins 13. **Proxy & auth:** `--proxy URL` routes both search and fetch through a proxy; `--auth-bearer` / `--auth-basic` (plus `*-file` variants, `searxng.toml` `auth_basic`/`auth_bearer` fields, and `SEARXNG_BEARER_TOKEN` / `SEARXNG_BASIC_AUTH` env vars) supply credentials. Priority: CLI flag > file > config file > env var 14. **Config defaults:** `searxng.toml` may pre-set most flags (engines, categories, language, safesearch, time_range, method, format, sort_by, timeout, max_retries, proxy, cache_ttl, fetch, fetch_timeout, fetch_retries, max_size, request_delay, auth_basic, auth_bearer); explicit CLI flags always win 15. **Structured errors:** in `--format json` mode, failures print a JSON object `{"error": "...", "error_code": "E_*", "recovery_hint": "...", "exit_code": N, "query": "..."}` to stdout so agents can parse them and decide recovery strategy **Key options:** - `--query "your search"` — **required unless** `--verify`, `--queries-file`, `--clear-cache`, `--cache-stats`, or `--dump-schema` is used - `--instance https://searx.example.org` — **required unless** `SEARXNG_INSTANCE` env var or a config file supplies it; comma-separated list enables failover - `--queries-file FILE` — read queries from a file (one per line; blank/`#` skipped) and run them in sequence; overrides `--query` - `--engines google,duckduckgo` — restrict to specific search engines (whitespace around commas is auto-stripped; see default list above) - `--method POST` — use POST instead of GET (better for long queries) - `--categories general,news` — comma-separated categories (whitespace auto-stripped) - `--language zh-CN` — language filter - `--pageno 1` — page number - `--time-range {day,month,year,none}` — time filter (default: `year`; `none` disables time filtering) - `--safesearch {0,1,2}` — safe search (default: `0` = off) - `--max-results N` — limit number of results (applied AFTER dedup+sort, so the highest-scoring/newest items are kept) - `--sort-by {score,date,engine,none}` — sort results (default: `score` descending; `none` preserves instance order). Applied after dedup, before `--max-results`. HTML-fallback results have no score and keep their order - `--no-dedup` — disable cross-engine deduplication (by default, duplicate URLs — same page ignoring tracking params/fragment — are collapsed, keeping the first occurrence's engine/score) - `--config FILE` — path to a `searxng.toml` config file; overrides the default auto-discovery (`./searxng.toml` → `~/.config/searxng-cli/searxng.toml` → `%APPDATA%/searxng-cli/searxng.toml` on Windows). Must be the first flag so its values can set defaults for other flags - `--verbose` / `-v` — show debug-level diagnostics on stderr (HTTP request URLs, response codes, cache keys, retry detail) - `--quiet` — suppress progress messages and retry notices on stderr; only warnings and errors are shown (no short flag: `-q` is `--query`) - `--include-domain a.com,b.org` — allowlist; only results from these domains are kept (applied after search) - `--exclude-domain pinterest.com` — blocklist; results from these domains are dropped (applied after search; wins over include on conflict) - `--serial` — disable parallel multi-instance probing; search instances strictly one at a time - `--verify` — health-check mode: verify instances and exit (no search); combine with `--format brief` for a table or `--format json` for machine-readable output - `--format json` — full JSON (default); includes `fetched` array when `--fetch` is used; errors are emitted as JSON to stdout - `--format brief` — title + URL + full snippet (no truncation by default; use `--snippet-len 200` to cap) - `--format urls` — only result URLs - `--format csv` — CSV export (title,url,engine,score,published_date,content); in batch mode (`--queries-file`), all queries merge into one CSV with a `query` column - `--fetch N` — after search, auto-fetch full text of top N result pages (concurrent, stdlib only) - `--fetch-timeout 10` — timeout per page fetch (default: 10s) - `--fetch-retries 3` — max retries per page fetch - `--max-size BYTES` — cap page size (default: unlimited; e.g. `5242880` for 5MB) - `--retry 5` — max retries per instance (default: 3) - `--timeout 15` — request timeout in seconds - `--fail-fast` — use only the first instance, don't fail over to the rest - `--proxy URL` — HTTP/HTTPS proxy for both search and fetch (e.g. `http://corp-proxy:8080`) - `--cache-ttl MINUTES` — cache results for N minutes (default: `0` = disabled); identical queries within the TTL skip the network - `--clear-cache` — delete all cached entries and exit (no search) - `--cache-stats` — print cache statistics (entries, age, size, path) and exit - `--dump-schema` — print the JSON Schema for `--format json` output and exit; lets AI agents programmatically discover field names and types without parsing prose docs - `--auth-bearer TOKEN` — `Authorization: Bearer` header for private instances - `--auth-bearer-file FILE` — read Bearer token from a file (first non-empty, non-`#` line); also honors `SEARXNG_BEARER_TOKEN` env var - `--auth-basic USER:PASS` — `Authorization: Basic` header (auto base64-encoded) - `--auth-basic-file FILE` — read `user:pass` from a file (first non-empty, non-`#` line); also honors `SEARXNG_BASIC_AUTH` env var - `--version` — print version and exit **Completion criterion:** Outputs valid JSON with `results` array. Non-zero exit on total failure (all instances exhausted). ### 2. `fetch.py` — Fetch & Extract Web Page Content ``` usage: fetch.py [-h] --url URL [--extract {text,html,markdown}] [--timeout SEC] [--retries N] [--max-size BYTES] [--user-agent STR] [--encoding CHARSET] [--no-redirect] [--referer URL] [--proxy URL] [--output FILE] [--auth-bearer TOKEN] [--auth-bearer-file FILE] [--auth-basic USER:PASS] [--auth-basic-file FILE] [--verbose] [--quiet] [--version] ``` **What it does:** 1. Downloads a web page via HTTP GET with retry + exponential backoff 2. **Stable fetching (v2.0.0 enhanced):** retries on 429/5xx/connection errors (3x default) with **Retry-After header compliance** and **capped 60s backoff**; falls back through a **12-UA deterministic pool** with **full browser fingerprint headers** (Sec-Ch-Ua/Sec-Fetch-*/Accept-Language); reuses a **requests.Session** for connection pooling + cookie persistence 3. **No size limit by default** — full page content returned; use `--max-size` for a cap 4. Detects charset from HTTP headers, HTML meta tags, or UTF-8 fallback 5. Extracts readable content using tree-based parsers (stdlib or BeautifulSoup); v2.0.0 adds **readability-lite** text-density fallback when `<article>`/`<main>` are missing 6. Outputs clean text, raw HTML, or properly-converted Markdown (with correct nested-link handling) **Key options:** - `--url https://...` — required - `--extract text` — clean readable text (default) - `--extract html` — raw HTML - `--extract markdown` — Markdown conversion (tree-based; handles nested tags, GFM tables, fenced code blocks, blockquotes, inline code, ordered/unordered/nested lists, definition lists, images, emphasis) - `--encoding gbk` — force charset for non-UTF-8 pages (else auto-detected from HTTP header / HTML meta / UTF-8 fallback) - `--timeout 15` — request timeout in seconds (v2.0.0: split into connect/read tuple internally) - `--retries 3` — max retries on transient errors (429/5xx/connection); falls back through the 12-UA pool when blocked - `--max-size BYTES` — cap page size (default: unlimited; e.g. `5242880` for 5MB) - `--user-agent STR` — custom User-Agent header (overrides per-domain UA selection) - `--no-redirect` — do **not** follow HTTP 3xx redirects (implemented for both the requests and stdlib paths) - `--referer URL` — set Referer header to disguise traffic source (v2.0.0 anti-bot measure) - `--proxy URL` — HTTP/HTTPS proxy (e.g. `http://corp-proxy:8080`); respects existing `HTTP_PROXY`/`HTTPS_PROXY` env vars when omitted - `--output FILE` — save to file instead of stdout - `--auth-bearer TOKEN` / `--auth-bearer-file FILE` — `Authorization: Bearer` header; file variant reads first non-empty, non-`#` line; also honors `SEARXNG_BEARER_TOKEN` env var - `--auth-basic USER:PASS` / `--auth-basic-file FILE` — `Authorization: Basic` header (auto base64-encoded); file variant + `SEARXNG_BASIC_AUTH` env var also supported - `--verbose` / `-v` — show debug-level diagnostics on stderr - `--quiet` — suppress progress messages on stderr; only warnings and errors are shown - `--version` — print version and exit **Extraction strategy (text mode):** 1. Strip non-content elements (script, style, nav, footer, header) 2. Extract `<article>`, `<main>`, or `<body>` content 3. v2.0.0: if only `<body>` matched, run **readability-lite** to pick the highest text-density `<div>`/`<section>` (filters out nav/sidebar/footer by class/id) 4. Collapse whitespace, output clean UTF-8 5. Warn if extracted text < 500 chars (likely JS-heavy or bot-blocked) **Completion criterion:** Outputs page content. Non-zero exit on HTTP failure. Stderr carries warnings for low-confidence extraction.