Files
searxng-use-cli/SKILL.md
T
thzxx 657af0a221 docs: 精简 SKILL.md,移除 AI Agent 无关章节
删除:SearXNG Search API Quick Reference、Common Workflow、Troubleshooting、Common Pitfalls、Verification Checklist

精简:Cross-Agent Compatibility 表格压缩为段落

SKILL.md: 43KB -> 30KB(-32%),聚焦 AI Agent 核心需求
2026-08-01 19:36:23 +08:00

29 KiB

name, description, version, author, license, platforms, metadata
name description version author license platforms metadata
searxng-use-cli Use when you need to search the web via your OWN SearXNG instance (no public-instance discovery). 3 CLI scripts + a shared common.py module — execute privacy-respecting searches against a user-supplied instance (with multi-instance failover, 5xx/429 retry, auto-fetch) or via SEARXNG_INSTANCE env / config file, fetch/extract readable text or markdown from web pages. Zero-config replacement for proprietary search APIs. 1.8.0 Metona Team MIT
linux
macos
windows
skill
tags related_skills
search
searxng
cli
web-scraping
privacy
no-api-key

SearXNG CLI Toolkit

Overview

SearXNG is a privacy-respecting metasearch engine that aggregates results from 70+ search services without tracking users. This skill provides three standalone Python CLI scripts — works with any AI agent (Hermes, Claude Code, Codex, OpenCode, Cursor, Trae, etc.) or directly from your terminal.

Public-instance discovery has been removed. You must supply your own SearXNG instance URL (self-hosted or one you trust). This makes behavior deterministic and avoids depending on volatile public instances.

Key capabilities:

Search & results

  • Multi-instance failover with parallel probing (faster failover, deterministic output order)
  • Exponential-backoff retry on transient errors (429/5xx/connection) via shared common.py
  • --verify health-check mode (reachability / JSON-API / latency / POST / engine list / auth status)
  • Cross-engine result deduplication (default on; --no-dedup disables) — collapses duplicate URLs ignoring tracking params (utm_*, gclid, etc.) and fragments
  • Result sorting (--sort-by {score,date,engine,none}; default: score descending) — applied after dedup, before --max-results
  • Domain allowlist/blocklist (--include-domain / --exclude-domain) — case-insensitive, ignores leading www., exclude wins on conflict
  • Batch mode (--queries-file) — run multiple queries from a file in sequence, combined output

Output formats

  • JSON (default, rich metadata), brief (title+URL+snippet), urls (plain list), CSV (spreadsheet-ready)
  • Enhanced Markdown conversion — nested ordered lists (numbered), mixed ul/ol nesting, <dl>/<dt>/<dd> definition lists, GFM tables, fenced code blocks, blockquotes
  • Structured JSON error output (in --format json mode) with error_code field for machine-readable failure reporting
  • JSON Lines streaming (--stream) — each result emitted as a separate JSON line to stdout, enabling incremental processing by AI agents
  • Progress events (--progress) — structured JSON Lines events to stderr for real-time execution tracking

Caching & config

  • SQLite result caching (--cache-ttl) — identical queries within a TTL skip the network entirely; --clear-cache / --cache-stats manage it
  • Config file (searxng.toml) pre-sets most flags; --config FILE loads a non-default config; instances.txt for plain URL lists
  • Instance resolution priority: -iSEARXNG_INSTANCE env → config file

Network & auth

  • Proxy support (--proxy) for both search and fetch (sets HTTP_PROXY/HTTPS_PROXY/NO_PROXY)
  • Auth via CLI flag, file, config file, or env var (--auth-bearer / --auth-basic + *-file variants) to avoid leaking secrets in shell history
  • Config-file auth — auth_basic and auth_bearer fields in searxng.toml let AI agents set credentials once (priority: CLI > file > config > env)
  • Credentials-file permission warning — --auth-*-file warns on stderr if the file is group/other-readable (POSIX only)

Engineering

  • Shared common.py module — unified retry/charset/auth/logging logic across both scripts
  • search.py --fetch reuses fetch.py's higher-quality text extractor (no code duplication)
  • Structured logging (--verbose / --quiet) — three levels: default INFO (progress + warnings), --verbose DEBUG (HTTP detail, cache keys), --quiet WARNING (errors only). All log output to stderr; stdout reserved for data
  • Engine/category whitespace normalization ("google, bing""google,bing")
  • --time-range none option to disable time filtering

Scripts + shared module:

  1. search.py — execute searches against a user-supplied instance, with multi-instance failover + exponential-backoff retry (429/5xx/connection) + auto-fetch + caching + batch + domain filtering
  2. fetch.py — download and extract readable text or markdown from web pages
  3. common.py — shared utilities (auth headers, charset detection, retry policy, fallback UAs, retry constants) used by both scripts
  4. cache.py — SQLite-backed result cache (SHA-256 key, TTL, WAL mode)
  5. _config.py — package constants (version, User-Agent)

Default settings

search.py ships with opinionated defaults tuned for AI research:

Setting Default Flag to override
Instance required — via -i, SEARXNG_INSTANCE env var, or config file -i / --instance
Safe search 0 (off) -s / --safesearch {0,1,2}
Time range year -t / --time-range {day,month,year,none} (none = disabled)
Output format json -f / --format {json,brief,urls,csv}
Engines google,bing,brave,duckduckgo,startpage,wikipedia,wikidata --engines <list>

Quick Start

# Prerequisites: Python 3.8+
# Optional but recommended:
pip install requests beautifulsoup4

# 1. Search against YOUR instance (instance URL is required)
python scripts/search.py -q "python asyncio tutorial" -i https://my-searxng.example.com

# 2. Multiple instances for failover (comma-separated)
python scripts/search.py -q "rust memory safety" \
  -i https://a.example.com,https://b.example.com --format brief

# 3. Search + auto-fetch top 3 result pages in one command
python scripts/search.py -q "climate policy" -i https://my-searxng.example.com --fetch 3

# 4. Fetch a result page
python scripts/fetch.py -u "https://example.com" --extract text

# 5. Skip -i by configuring the instance once (env var, current shell)
export SEARXNG_INSTANCE="https://my-searxng.example.com,https://backup.example.com"
python scripts/search.py -q "python asyncio tutorial"          # -i not needed

# 6. Or use a config file (./searxng.toml or ~/.config/searxng-cli/searxng.toml)
#    [searxng]
#    instance = "https://my-searxng.example.com"
#    # or: instances = ["https://a.example.com", "https://b.example.com"]
#    # Auth (optional, for private instances):
#    auth_basic = "user:password"        # Basic auth
#    auth_bearer = "sk-token-123"        # Bearer token (auth_bearer wins if both set)
#    # Any flag below can also be pre-set here (engines, categories, language,
#    # safesearch, time_range, method, format, timeout, max_retries, proxy,
#    # cache_ttl, fetch, fetch_timeout, fetch_retries, max_size).
#    # Explicit CLI flags always override config values.
#    # Auth priority: --auth-* > --auth-*-file > searxng.toml > env var
#    Plain list also works in ./instances.txt (one URL per line, # for comments)

# 7. Cache results for 30 minutes (identical queries skip the network)
python scripts/search.py -q "python asyncio" -i https://s.example.com --cache-ttl 30

# 7b. Sort by date (newest first) or disable dedup for raw engine output
python scripts/search.py -q "ai news" -i https://s.example.com --sort-by date --no-dedup

# 8. Batch: run queries from a file (one per line; blank/# lines skipped)
python scripts/search.py --queries-file queries.txt -i https://s.example.com --format json > batch.json

# 9. Domain allowlist + blocklist (applied after search)
python scripts/search.py -q "rust async" -i https://s.example.com \
  --include-domain doc.rust-lang.org,wikipedia.org --exclude-domain pinterest.com

# 10. Route through a corporate proxy (applies to search and fetch)
python scripts/search.py -q "ai news" -i https://s.example.com --proxy http://corp-proxy:8080

# 11. Auth from a file (avoids leaking tokens in shell history)
python scripts/search.py -q "test" -i https://private.example.com --auth-bearer-file ~/.searxng_token

# 12. Cache management (no search performed)
python scripts/search.py --cache-stats        # entry count, age, size, path
python scripts/search.py --clear-cache        # delete all entries

# 13. Export results as CSV (great for spreadsheets / data analysis)
python scripts/search.py -q "rust async" -i https://s.example.com --format csv > results.csv

# 14. Use a specific config file (overrides auto-discovered searxng.toml)
python scripts/search.py --config ./my-config.toml -q "test"

# 15. Control log verbosity on stderr
python scripts/search.py -q "test" -i https://s.example.com --verbose   # debug detail
python scripts/search.py -q "test" -i https://s.example.com --quiet     # errors only

Dependency levels:

Level Scripts What you get
Zero deps (stdlib only) search.py, _config.py Full search + auto-fetch
pip install requests fetch.py Better HTTP (session reuse, redirect handling)
pip install beautifulsoup4 fetch.py Higher-quality text extraction

When to Use

  • Web search without API keys — programmatic search results against your own SearXNG instance
  • Privacy-conscious research — queries routed through your instance, not ad-tech infrastructure
  • Scraping search results — batch query multiple terms and collect structured results as JSON
  • Fetching search result pages — follow links from search results and extract clean text

Don't use for:

  • High-frequency production search without a properly-scaled instance — respect your instance's rate limits
  • Guaranteed uptime/accuracy — depends entirely on the instance you supply

AI Agent Integration Guide

This section documents the structured interfaces that AI agents can rely on for programmatic integration. All features are designed to be machine-readable and machine-actionable.

Output Channels

Channel Content Description
stdout Data JSON/CSV/text — the only source AI should parse
stderr Logs + Progress Human-readable logs (default) or JSON Lines events (--progress)
exit 0 Success Results available on stdout
exit 1 Fatal error Error JSON on stdout (in --format json mode) or stderr
exit 2 Empty results Search succeeded but returned no results

Error Code System

In --format json mode, errors are emitted as structured JSON on stdout:

{
  "error": "All 2 instances failed. Last error: connection refused",
  "error_code": "E_NETWORK",
  "recovery_hint": "Retry with backoff, or try a different SearXNG instance. Check network connectivity, proxy settings, and instance uptime.",
  "exit_code": 1,
  "query": "search term"
}

AI agents can use error_code to programmatically decide recovery strategy, and recovery_hint for a ready-to-use actionable suggestion:

Code Meaning recovery_hint (abridged)
E_CONFIG Configuration error (no instance resolved) Provide -i/SEARXNG_INSTANCE/config file
E_AUTH Authentication failed (401/403) Verify credentials, check token expiry & permissions
E_NETWORK Network error (connection refused, timeout, 5xx, all instances failed) Retry with backoff, switch instance, check proxy
E_RATE_LIMIT Rate limited (429) Wait and retry, reduce frequency, distribute load
E_PARSE Parse error (JSON/HTML parsing failed) Try different instance, switch --method
E_EMPTY Empty results (exit code 2) Refine query, broaden --time-range, add --categories
E_INPUT Input error (bad parameters, file not found) Check query syntax, flag combinations, file paths
E_INTERNAL Internal error (unexpected exception) Re-run with --verbose, report bug

JSON Lines Streaming (--stream)

For large result sets, --stream outputs results as JSON Lines (one JSON object per line) to stdout, allowing AI agents to process results incrementally without waiting for the full response:

python scripts/search.py -q "large topic" -i https://your-instance --stream

Output format (each line is a separate JSON object):

{"type": "result", "result": {"title": "...", "url": "...", "content": "..."}}
{"type": "result", "result": {"title": "...", "url": "...", "content": "..."}}
{"type": "done", "schema_version": "1.0", "count": 2, "query": "large topic"}
  • type: "result" — one per search result, emitted as soon as available
  • type: "done" — terminal event with total count and schema_version, always emitted last
  • type: "error" — emitted when the search fails, includes error_code and recovery_hint:
    {"type": "error", "error": "All 2 instances failed. Last error: HTTP Error 403", "error_code": "E_AUTH", "recovery_hint": "Verify credentials...", "query": "..."}
    
  • Exit code 0 on success, 1 on error, 2 on empty results (done event still emitted)

Only valid with --format json (single query mode). Using --stream with --queries-file or non-json formats raises E_INPUT immediately.

Progress Events (--progress)

For long-running operations, --progress emits structured JSON Lines events to stderr, enabling AI agents to track execution progress in real time:

python scripts/search.py -q "research topic" -i https://your-instance --progress --fetch 3

Event types (each on its own line, JSON Lines format on stderr):

{"event": "start", "query": "research topic", "instances": 2}
{"event": "instance_try", "url": "https://instance1.example.com", "attempt": 1}
{"event": "instance_ok", "url": "https://instance1.example.com", "latency": 0.342, "results": 10}
{"event": "instance_fail", "url": "https://instance2.example.com", "error": "HTTP 503", "error_code": "E_NETWORK"}
{"event": "cache_hit", "query": "research topic", "ttl": 30}
{"event": "cache_store", "query": "research topic", "ttl": 30}
{"event": "fetch_start", "count": 3}
{"event": "fetch_ok", "url": "https://example.com/page", "chars": 12345}
{"event": "fetch_fail", "url": "https://bad.example.com", "error": "HTTP 503"}
{"event": "done", "results": 10, "query": "research topic"}
{"event": "error", "error": "connection refused", "error_code": "E_NETWORK", "query": "..."}

AI agents can parse these events to:

  • Show progress indicators to users
  • Detect cache hits (skip waiting)
  • Monitor fetch failures and retry strategies
  • Correlate errors with specific queries in batch mode

--progress and --verbose can be used together (progress events on stderr, debug logs also on stderr). --progress events are JSON Lines; --verbose logs are human-readable text.

Output JSON Schema

The default --format json output includes a schema_version field so AI agents can detect breaking changes. Run --dump-schema to get the full JSON Schema document programmatically:

python scripts/search.py --dump-schema

Single query output shape:

{
  "schema_version": "1.0",
  "query": "search term",
  "number_of_results": 10,
  "results": [
    {
      "title": "Result title",
      "url": "https://example.com/page",
      "content": "Snippet text...",
      "engine": "google",
      "score": 1.0,
      "category": "general",
      "published_date": "2024-01-15T10:30:00"
    }
  ],
  "answers": ["Direct answer if available"],
  "corrections": [],
  "suggestions": ["related suggestion"],
  "infoboxes": [],
  "unresponsive_engines": [["engine_name", "error reason"]],
  "fetched": [
    {
      "url": "https://example.com/page",
      "status": "ok",
      "text": "Extracted page content...",
      "text_length": 12345,
      "truncated": false,
      "final_url": "https://example.com/final",
      "user_agent_used": "searxng-cli/1.8.0"
    }
  ],
  "fetched_source": "json"
}

Batch mode (--queries-file) wraps results in a unified schema:

{
  "schema_version": "1.0",
  "queries": [
    {"query": "term1", "status": "ok", "results": {...}},
    {"query": "term2", "status": "error", "error": "...", "error_code": "E_NETWORK"}
  ]
}

Each batch entry has a status field ("ok" or "error"). Successful entries contain results; failed entries contain error and error_code.

Fields marked as optional may be absent. The fetched and fetched_source fields only appear when --fetch N is used.

Cross-Agent Compatibility

These scripts are agent-agnostic — they work with any AI agent that can invoke terminal commands (Hermes, Claude Code, Codex, OpenCode, Cursor, Trae, etc.), or directly from a terminal.

Key design decisions for universal compatibility:

  • Zero external dependencies (stdlib-only for search.py)
  • Scripts inject their own directory into sys.path, so they run from any working directory
  • Stdout carries data (JSON/text), stderr carries progress/warnings
  • Exit codes: 0=success, 1=fatal error, 2=no results/empty
  • NO agent-specific API calls or tool dependencies — purely CLI-based, portable across all agent platforms

Scripts

All scripts live in scripts/; run with python scripts/<name>.py from any directory. They import _config.py for shared constants and self-inject their own directory into sys.path.

usage: search.py [-h] [--query QUERY] [--instance URL]
                 [--categories CATS] [--language LANG] [--pageno N]
                 [--time-range {day,month,year,none}] [--safesearch {0,1,2}]
                 [--engines E] [--method {GET,POST}] [--max-results N]
                 [--format {json,brief,urls,csv}] [--snippet-len N]
                 [--fetch N] [--fetch-timeout SEC] [--fetch-retries N]
                 [--max-size BYTES] [--output FILE] [--timeout SEC]
                 [--retry N] [--fail-fast] [--serial] [--verify]
                 [--auth-bearer TOKEN] [--auth-bearer-file FILE]
                 [--auth-basic USER:PASS] [--auth-basic-file FILE]
                 [--proxy URL] [--include-domain DOMAINS]
                 [--exclude-domain DOMAINS] [--queries-file FILE]
                 [--cache-ttl MINUTES] [--clear-cache] [--cache-stats]
                 [--sort-by {score,date,engine,none}] [--no-dedup]
                 [--config FILE] [--dump-schema] [--verbose] [--quiet] [--version]

What it does:

  1. Takes a search query and resolves one or more instance URLs (-i, SEARXNG_INSTANCE, or config file — comma-separated for failover)
  2. Multi-instance failover: if an instance fails (429/5xx/timeout/captcha), automatically tries the next one
  3. Exponential backoff: retries each instance up to 3 times with jitter on transient errors (429, 502, 503, 504, and connection errors)
  4. Parallel probing (default for 2+ instances): queries every instance concurrently and returns the first successful result in your original instance order — this keeps output deterministic while drastically speeding up failover when an early instance is down/slow. Use --serial to disable.
  5. Calls the SearXNG API (GET or POST) with format=json
  6. Falls back to HTML scraping if JSON is blocked
  7. HTML parser extracts results + suggestions + answers + infoboxes (full content, never truncated)
  8. Stable auto-fetch: --fetch 3 concurrently downloads top 3 result pages with retry, browser-UA fallback, CAPTCHA detection, and fetch.py's higher-quality text extractor (the same engine fetch.py uses)
  9. Health-check mode: --verify probes each instance (reachability / JSON-API support / latency / POST support / engine list / auth status) and prints a report, then exits without searching — use it to validate your instance list
  10. Result caching: --cache-ttl 30 stores results for 30 min; identical queries within the TTL skip the network entirely. Cache lives at $SEARXNG_CACHE_DIR or ~/.cache/searxng-cli/cache.db (SQLite, WAL mode). --clear-cache / --cache-stats manage it without searching
  11. Batch mode: --queries-file FILE reads one query per line (blank/# lines skipped) and runs them in sequence; JSON output is {"schema_version": "1.0", "queries": [{"query":..., "status": "ok"|"error", ...}]} (or one block per query in brief/urls). A failed query is recorded but does not abort the batch. Exit codes: 0 if any query returned results, 1 if all errored, 2 if all empty
  12. Dedup + sort + domain filter: After search (and cache), duplicate URLs are collapsed (default; --no-dedup disables), results are sorted (--sort-by; default: score descending), and then --include-domain/--exclude-domain filter by domain. Matching is case-insensitive and ignores a leading www.; when a domain is in both lists, exclude wins
  13. Proxy & auth: --proxy URL routes both search and fetch through a proxy; --auth-bearer / --auth-basic (plus *-file variants, searxng.toml auth_basic/auth_bearer fields, and SEARXNG_BEARER_TOKEN / SEARXNG_BASIC_AUTH env vars) supply credentials. Priority: CLI flag > file > config file > env var
  14. Config defaults: searxng.toml may pre-set most flags (engines, categories, language, safesearch, time_range, method, format, sort_by, timeout, max_retries, proxy, cache_ttl, fetch, fetch_timeout, fetch_retries, max_size, auth_basic, auth_bearer); explicit CLI flags always win
  15. Structured errors: in --format json mode, failures print a JSON object {"error": "...", "error_code": "E_*", "recovery_hint": "...", "exit_code": N, "query": "..."} to stdout so agents can parse them and decide recovery strategy

Key options:

  • --query "your search"required unless --verify, --queries-file, --clear-cache, --cache-stats, or --dump-schema is used
  • --instance https://searx.example.orgrequired unless SEARXNG_INSTANCE env var or a config file supplies it; comma-separated list enables failover
  • --queries-file FILE — read queries from a file (one per line; blank/# skipped) and run them in sequence; overrides --query
  • --engines google,duckduckgo — restrict to specific search engines (whitespace around commas is auto-stripped; see default list above)
  • --method POST — use POST instead of GET (better for long queries)
  • --categories general,news — comma-separated categories (whitespace auto-stripped)
  • --language zh-CN — language filter
  • --pageno 1 — page number
  • --time-range {day,month,year,none} — time filter (default: year; none disables time filtering)
  • --safesearch {0,1,2} — safe search (default: 0 = off)
  • --max-results N — limit number of results (applied AFTER dedup+sort, so the highest-scoring/newest items are kept)
  • --sort-by {score,date,engine,none} — sort results (default: score descending; none preserves instance order). Applied after dedup, before --max-results. HTML-fallback results have no score and keep their order
  • --no-dedup — disable cross-engine deduplication (by default, duplicate URLs — same page ignoring tracking params/fragment — are collapsed, keeping the first occurrence's engine/score)
  • --config FILE — path to a searxng.toml config file; overrides the default auto-discovery (./searxng.toml~/.config/searxng-cli/searxng.toml). Must be the first flag so its values can set defaults for other flags
  • --verbose / -v — show debug-level diagnostics on stderr (HTTP request URLs, response codes, cache keys, retry detail)
  • --quiet — suppress progress messages and retry notices on stderr; only warnings and errors are shown (no short flag: -q is --query)
  • --include-domain a.com,b.org — allowlist; only results from these domains are kept (applied after search)
  • --exclude-domain pinterest.com — blocklist; results from these domains are dropped (applied after search; wins over include on conflict)
  • --serial — disable parallel multi-instance probing; search instances strictly one at a time
  • --verify — health-check mode: verify instances and exit (no search); combine with --format brief for a table or --format json for machine-readable output
  • --format json — full JSON (default); includes fetched array when --fetch is used; errors are emitted as JSON to stdout
  • --format brief — title + URL + full snippet (no truncation by default; use --snippet-len 200 to cap)
  • --format urls — only result URLs
  • --format csv — CSV export (title,url,engine,score,published_date,content); in batch mode (--queries-file), all queries merge into one CSV with a query column
  • --fetch N — after search, auto-fetch full text of top N result pages (concurrent, stdlib only)
  • --fetch-timeout 10 — timeout per page fetch (default: 10s)
  • --fetch-retries 3 — max retries per page fetch
  • --max-size BYTES — cap page size (default: unlimited; e.g. 5242880 for 5MB)
  • --retry 5 — max retries per instance (default: 3)
  • --timeout 15 — request timeout in seconds
  • --fail-fast — use only the first instance, don't fail over to the rest
  • --proxy URL — HTTP/HTTPS proxy for both search and fetch (e.g. http://corp-proxy:8080)
  • --cache-ttl MINUTES — cache results for N minutes (default: 0 = disabled); identical queries within the TTL skip the network
  • --clear-cache — delete all cached entries and exit (no search)
  • --cache-stats — print cache statistics (entries, age, size, path) and exit
  • --dump-schema — print the JSON Schema for --format json output and exit; lets AI agents programmatically discover field names and types without parsing prose docs
  • --auth-bearer TOKENAuthorization: Bearer header for private instances
  • --auth-bearer-file FILE — read Bearer token from a file (first non-empty, non-# line); also honors SEARXNG_BEARER_TOKEN env var
  • --auth-basic USER:PASSAuthorization: Basic header (auto base64-encoded)
  • --auth-basic-file FILE — read user:pass from a file (first non-empty, non-# line); also honors SEARXNG_BASIC_AUTH env var
  • --version — print version and exit

Completion criterion: Outputs valid JSON with results array. Non-zero exit on total failure (all instances exhausted).

2. fetch.py — Fetch & Extract Web Page Content

usage: fetch.py [-h] --url URL [--extract {text,html,markdown}]
                [--timeout SEC] [--retries N] [--max-size BYTES]
                [--user-agent STR] [--encoding CHARSET]
                [--no-redirect] [--proxy URL] [--output FILE]
                [--auth-bearer TOKEN] [--auth-bearer-file FILE]
                [--auth-basic USER:PASS] [--auth-basic-file FILE]
                [--verbose] [--quiet] [--version]

What it does:

  1. Downloads a web page via HTTP GET with retry + exponential backoff
  2. Stable fetching: retries on 429/5xx/connection errors (3x default), falls back to browser User-Agent if blocked
  3. No size limit by default — full page content returned; use --max-size for a cap
  4. Detects charset from HTTP headers, HTML meta tags, or UTF-8 fallback
  5. Extracts readable content using tree-based parsers (stdlib or BeautifulSoup)
  6. Outputs clean text, raw HTML, or properly-converted Markdown (with correct nested-link handling)

Key options:

  • --url https://... — required
  • --extract text — clean readable text (default)
  • --extract html — raw HTML
  • --extract markdown — Markdown conversion (tree-based; handles nested tags, GFM tables, fenced code blocks, blockquotes, inline code, ordered/unordered/nested lists, definition lists, images, emphasis)
  • --encoding gbk — force charset for non-UTF-8 pages (else auto-detected from HTTP header / HTML meta / UTF-8 fallback)
  • --timeout 15 — request timeout in seconds
  • --retries 3 — max retries on transient errors (429/5xx/connection); falls back to a browser User-Agent when blocked
  • --max-size BYTES — cap page size (default: unlimited; e.g. 5242880 for 5MB)
  • --user-agent STR — custom User-Agent header
  • --no-redirect — do not follow HTTP 3xx redirects (implemented for both the requests and stdlib paths)
  • --proxy URL — HTTP/HTTPS proxy (e.g. http://corp-proxy:8080); respects existing HTTP_PROXY/HTTPS_PROXY env vars when omitted
  • --output FILE — save to file instead of stdout
  • --auth-bearer TOKEN / --auth-bearer-file FILEAuthorization: Bearer header; file variant reads first non-empty, non-# line; also honors SEARXNG_BEARER_TOKEN env var
  • --auth-basic USER:PASS / --auth-basic-file FILEAuthorization: Basic header (auto base64-encoded); file variant + SEARXNG_BASIC_AUTH env var also supported
  • --verbose / -v — show debug-level diagnostics on stderr
  • --quiet — suppress progress messages on stderr; only warnings and errors are shown
  • --version — print version and exit

Extraction strategy (text mode):

  1. Strip non-content elements (script, style, nav, footer, header)
  2. Extract <article>, <main>, or <body> content
  3. Collapse whitespace, output clean UTF-8
  4. Warn if extracted text < 500 chars (likely JS-heavy or bot-blocked)

Completion criterion: Outputs page content. Non-zero exit on HTTP failure. Stderr carries warnings for low-confidence extraction.