fix(v2.1.1): 修复执行问题记录中的真实 bug + 文档对齐

源码修复(5 项):
1. search.py --time-range choices 加入 week(对齐 SearXNG API 四档)
2. fetch.py stdlib 路径处理 gzip/deflate 解压(被沙箱伪响应掩盖的真实 bug,
   无 requests 环境抓取压缩服务器会全页 U+FFFD 乱码)
3. search.py --research 模式实现跨角度合并去重,输出 merged_results 字段
   (兑现文档承诺 "Results are merged and deduplicated")
4. search.py fetch_page 返回 error_code 字段 + AdaptiveThrottle 用
   E_RATE_LIMIT 结构化检测 429(原字符串匹配 "429" 会漏判
   "Too Many Requests")
5. search.py _retry_with_backoff 复用 compute_backoff_delay(60s 封顶)
   + 处理 Retry-After header,与 fetch.py 保持一致

增强(3 项):
- common.py 精确化 baidu 子域列表(pan.baidu.com/cloud.baidu.com 不再误伤)
- search.py expand_research_queries 根据主题语言切换中英文后缀
- search.py 新增 _warn_unresponsive_engines,识别实例侧引擎挂起并提示

文档/版本:
- _config.py VERSION 2.1.0 → 2.1.1
- SKILL.md 同步更新(time-range week、merged_results、error_code、baidu 精确化)
- README.md 同步更新 + 测试数量 503 → 539

测试: 539 个全部通过,含 6 个新增验证测试
This commit is contained in:
2026-08-03 12:54:27 +08:00
parent e94cbe0783
commit 0c8fdc1e45
9 changed files with 257 additions and 41 deletions
+9 -6
View File
@@ -154,7 +154,7 @@ auth_basic = "user:password" # Basic 认证
python scripts/search.py -q "查询词" -i https://your-instance \ python scripts/search.py -q "查询词" -i https://your-instance \
[--format json|brief|urls|csv] \ [--format json|brief|urls|csv] \
[--engines google,bing,brave] \ [--engines google,bing,brave] \
[--time-range day|month|year|none] \ [--time-range day|week|month|year|none] \
[--language zh-CN] \ [--language zh-CN] \
[--sort-by score|date|engine|none] \ [--sort-by score|date|engine|none] \
[--no-dedup] \ [--no-dedup] \
@@ -183,7 +183,7 @@ python scripts/search.py -q "查询词" -i https://your-instance \
| `-i / --instance` | SearXNG 实例 URL,逗号分隔实现故障转移 | — | | `-i / --instance` | SearXNG 实例 URL,逗号分隔实现故障转移 | — |
| `-f / --format` | 输出格式:json/brief/urls/csv | json | | `-f / --format` | 输出格式:json/brief/urls/csv | json |
| `--engines` | 搜索引擎列表 | google,bing,brave,duckduckgo,startpage,wikipedia,wikidata | | `--engines` | 搜索引擎列表 | google,bing,brave,duckduckgo,startpage,wikipedia,wikidata |
| `-t / --time-range` | 时间范围:day/month/year/none | year | | `-t / --time-range` | 时间范围:day/week/month/year/none | year |
| `-s / --safesearch` | 安全搜索:0/1/2 | 0(关闭) | | `-s / --safesearch` | 安全搜索:0/1/2 | 0(关闭) |
| `-l / --language` | 语言代码 | — | | `-l / --language` | 语言代码 | — |
| `-p / --pageno` | 页码 | 1 | | `-p / --pageno` | 页码 | 1 |
@@ -250,6 +250,7 @@ python scripts/fetch.py -u https://example.com \
- 域名白名单/黑名单 - 域名白名单/黑名单
- 批量查询(`--queries-file` - 批量查询(`--queries-file`
- v2.1.0 研究模式(`--research`):给定主题自动扩展 5 个多角度查询(overview/profile/background/works/review),输出含 `research_topic``research_queries` 元数据 - v2.1.0 研究模式(`--research`):给定主题自动扩展 5 个多角度查询(overview/profile/background/works/review),输出含 `research_topic``research_queries` 元数据
- v2.1.1 增强:跨角度合并去重,输出 `merged_results` 字段;根据主题语言自动切换中英文后缀(中文主题用"简介/经历/作品/评价",英文主题用"profile/background/works/reviews"
- 实例健康检查(`--verify` - 实例健康检查(`--verify`
**输出** **输出**
@@ -276,11 +277,13 @@ python scripts/fetch.py -u https://example.com \
- WAF 指纹库:识别 Cloudflare / Imperva / PerimeterX / DataDome / Akamai / 通用反爬页,全文档扫描(非仅前 2000 字符) - WAF 指纹库:识别 Cloudflare / Imperva / PerimeterX / DataDome / Akamai / 通用反爬页,全文档扫描(非仅前 2000 字符)
- Wayback Machine 兜底:404/403/超时自动尝试 `https://web.archive.org/web/2/<url>`,默认启用,`--no-fallback` 关闭 - Wayback Machine 兜底:404/403/超时自动尝试 `https://web.archive.org/web/2/<url>`,默认启用,`--no-fallback` 关闭
- v2.1.0 fetch.py 独立调用也支持 Wayback 兜底(之前仅 search.py --fetch 路径有) - v2.1.0 fetch.py 独立调用也支持 Wayback 兜底(之前仅 search.py --fetch 路径有)
- v2.1.0 被墙站点智能回退:`is_hard_blocked_domain()` 识别百度百科/知乎/微博/微信公众号/豆瓣等强反爬站点,403 时自动优先 Wayback - v2.1.0 被墙站点智能回退:`is_hard_blocked_domain()` 识别百度搜索/百度百科/知乎/微博/微信公众号/豆瓣等强反爬站点,403 时自动优先 Wayback
- v2.1.1 精确化:`baidu.com` 从子域匹配改为精确子域列表(www/baike/zhidao/tieba/wenku),`pan.baidu.com`(网盘)/`cloud.baidu.com`(智能云)不再被误伤
- 自适应限流:`AdaptiveThrottle` 状态机,连续 3 次失败自动翻倍延迟 + 减半并发,429 触发全局暂停 30s - 自适应限流:`AdaptiveThrottle` 状态机,连续 3 次失败自动翻倍延迟 + 减半并发,429 触发全局暂停 30s
- v2.1.1`report_failure` 新增 `error_code` 参数,优先用结构化 `E_RATE_LIMIT` 检测 429(原字符串匹配"429"会漏判"Too Many Requests");`fetch_page` 返回结果新增 `error_code` 字段
- `--fetch-report`:结构化抓取报告到 stderr(每 URL 状态/WAF 类型/兜底方式/字符数 + JSON 摘要) - `--fetch-report`:结构化抓取报告到 stderr(每 URL 状态/WAF 类型/兜底方式/字符数 + JSON 摘要)
- `--referer` / `--request-delay`:精细控制 Referer 头和请求间隔 - `--referer` / `--request-delay`:精细控制 Referer 头和请求间隔
- fetch 结果新增字段:`anti_bot_detected`bool)、`waf_type`str|null)、`fallback_used`str|null - fetch 结果新增字段:`anti_bot_detected`bool)、`waf_type`str|null)、`fallback_used`str|null`error_code`str|nullv2.1.1
**缓存** **缓存**
- SQLite 缓存(`--cache-ttl`),相同查询在 TTL 内跳过网络 - SQLite 缓存(`--cache-ttl`),相同查询在 TTL 内跳过网络
@@ -303,7 +306,7 @@ python scripts/fetch.py -u https://example.com \
- JSON Lines 流式输出(`--stream`,含 `error` 事件类型) - JSON Lines 流式输出(`--stream`,含 `error` 事件类型)
- 进度事件(`--progress`,含 `instance_try`/`instance_ok`/`instance_fail`JSON Lines 到 stderr - 进度事件(`--progress`,含 `instance_try`/`instance_ok`/`instance_fail`JSON Lines 到 stderr
- batch 模式统一 schema`status` 字段区分成功/失败) - batch 模式统一 schema`status` 字段区分成功/失败)
- 503 个单元+集成测试 - 539 个单元+集成测试
## 跨 Agent 兼容性 ## 跨 Agent 兼容性
@@ -331,7 +334,7 @@ pip install pytest
pytest -q pytest -q
``` ```
503 个测试覆盖:缓存操作、认证解析、域名过滤、Markdown 转换、搜索逻辑、集成流程、日志配置、HTML 回退、自动抓取、健康检查、输出格式化、实例解析、并行搜索、CLI 端到端、错误码分类、流式输出、进度事件、配置文件认证、schema_version、recovery_hint、batch 统一 schema、--dump-schema、UTF-8 stdout 强制、Windows APPDATA 路径、v2.0.0 浏览器指纹头、WAF 反爬检测(v2.0.1 收窄误判 + title 精准检测)、Wayback Machine 兜底(v2.0.1 修复 HTTP 200 反爬页触发)、自适应限流、v2.1.0 Wayback 共享逻辑(should_try_wayback/build_wayback_url)、被墙站点智能回退(is_hard_blocked_domain)、--research 研究模式(多角度查询扩展)。 539 个测试覆盖:缓存操作、认证解析、域名过滤、Markdown 转换、搜索逻辑、集成流程、日志配置、HTML 回退、自动抓取、健康检查、输出格式化、实例解析、并行搜索、CLI 端到端、错误码分类、流式输出、进度事件、配置文件认证、schema_version、recovery_hint、batch 统一 schema、--dump-schema、UTF-8 stdout 强制、Windows APPDATA 路径、v2.0.0 浏览器指纹头、WAF 反爬检测(v2.0.1 收窄误判 + title 精准检测)、Wayback Machine 兜底(v2.0.1 修复 HTTP 200 反爬页触发)、自适应限流v2.1.1 结构化 error_code 检测 429、v2.1.0 Wayback 共享逻辑(should_try_wayback/build_wayback_url)、被墙站点智能回退(v2.1.1 精确化 baidu 子域,pan/cloud 不再误伤)、--research 研究模式(v2.1.1 跨角度合并去重 + 中英文双语后缀)、--time-range week 支持(v2.1.1)、stdlib gzip/deflate 解压(v2.1.1)、实例引擎挂起检测(v2.1.1 _warn_unresponsive_engines)。
## 项目结构 ## 项目结构
+11 -8
View File
@@ -1,7 +1,7 @@
--- ---
name: searxng-use-cli name: searxng-use-cli
description: Use when you need to search the web via your OWN SearXNG instance (no public-instance discovery). 3 CLI scripts + a shared common.py module — execute privacy-respecting searches against a user-supplied instance (with multi-instance failover, 5xx/429 retry, auto-fetch) or via SEARXNG_INSTANCE env / config file, fetch/extract readable text or markdown from web pages. Zero-config replacement for proprietary search APIs. description: Use when you need to search the web via your OWN SearXNG instance (no public-instance discovery). 3 CLI scripts + a shared common.py module — execute privacy-respecting searches against a user-supplied instance (with multi-instance failover, 5xx/429 retry, auto-fetch) or via SEARXNG_INSTANCE env / config file, fetch/extract readable text or markdown from web pages. Zero-config replacement for proprietary search APIs.
version: 2.1.0 version: 2.1.1
author: Metona Team author: Metona Team
license: MIT license: MIT
platforms: [linux, macos, windows] platforms: [linux, macos, windows]
@@ -73,16 +73,19 @@ SearXNG is a privacy-respecting metasearch engine that aggregates results from 7
- **Capped backoff** — `compute_backoff_delay()` caps at 60s (was uncapped: 1.5*2^10 = 1536s would hang the process) - **Capped backoff** — `compute_backoff_delay()` caps at 60s (was uncapped: 1.5*2^10 = 1536s would hang the process)
- **WAF fingerprint library** — `_detect_anti_bot()` identifies Cloudflare / Imperva / PerimeterX / DataDome / Akamai / Anubis / generic challenges via `<title>` tag matching (most precise) + full-document scan of technical identifiers (cookie/header/JS variable names, not bare vendor names). v2.0.1 narrowed broad keywords (e.g. bare `cloudflare`/`captcha`/`challenge`) that caused false positives on normal articles, and added title-tag detection. Returns `waf_type` for AI-agent decisioning - **WAF fingerprint library** — `_detect_anti_bot()` identifies Cloudflare / Imperva / PerimeterX / DataDome / Akamai / Anubis / generic challenges via `<title>` tag matching (most precise) + full-document scan of technical identifiers (cookie/header/JS variable names, not bare vendor names). v2.0.1 narrowed broad keywords (e.g. bare `cloudflare`/`captcha`/`challenge`) that caused false positives on normal articles, and added title-tag detection. Returns `waf_type` for AI-agent decisioning
- **Wayback Machine fallback** — 404/403/timeout AND anti-bot-blocked pages automatically retry via `https://web.archive.org/web/2/<url>` (latest snapshot). v2.0.1 fixed a bug where HTTP 200 anti-bot challenge pages (Cloudflare returns 200 for JS challenges) bypassed the fallback trigger. v2.1.0: `fetch.py` standalone calls now also get Wayback fallback (was only in `search.py --fetch`). Default ON; `--no-fallback` disables. Independent 10s timeout so Wayback slowness never blocks the main flow - **Wayback Machine fallback** — 404/403/timeout AND anti-bot-blocked pages automatically retry via `https://web.archive.org/web/2/<url>` (latest snapshot). v2.0.1 fixed a bug where HTTP 200 anti-bot challenge pages (Cloudflare returns 200 for JS challenges) bypassed the fallback trigger. v2.1.0: `fetch.py` standalone calls now also get Wayback fallback (was only in `search.py --fetch`). Default ON; `--no-fallback` disables. Independent 10s timeout so Wayback slowness never blocks the main flow
- **Hard-blocked domain fallback** (v2.1.0) — `is_hard_blocked_domain()` detects known strong-anti-bot sites (baike.baidu.com, zhihu.com, weibo.com, mp.weixin.qq.com, douban.com, etc.) that almost always return 403 regardless of UA. These sites bypass the normal `should_try_wayback()` check and trigger Wayback immediately on failure. List maintained in `common.py` `HARD_BLOCKED_DOMAINS` - **Hard-blocked domain fallback** (v2.1.0, refined in v2.1.1) — `is_hard_blocked_domain()` detects known strong-anti-bot sites (www.baidu.com, baike.baidu.com, zhihu.com, weibo.com, mp.weixin.qq.com, douban.com, etc.) that almost always return 403 regardless of UA. These sites bypass the normal `should_try_wayback()` check and trigger Wayback immediately on failure. v2.1.1 refined the list: moved `baidu.com` from broad subdomain matching to a precise per-subdomain list (www/baike/zhidao/tieba/wenku), so `pan.baidu.com` (netdisk) and `cloud.baidu.com` (cloud) are no longer false-positively blocked. List maintained in `common.py` `HARD_BLOCKED_DOMAINS`
- **Adaptive throttling** — `AdaptiveThrottle` state machine: 3 consecutive failures → double delay + halve concurrency; 5 consecutive successes → gradual recovery; 429 → global pause 30s. Thread-safe - **Adaptive throttling** — `AdaptiveThrottle` state machine: 3 consecutive failures → double delay + halve concurrency; 5 consecutive successes → gradual recovery; 429 → global pause 30s. Thread-safe
- **readability-lite extraction** — when `<article>`/`<main>`/content-class `<div>` are all missing, `_readability_lite()` picks the highest text-density node (text chars / tag count + `<p>` weighting), avoiding nav/sidebar/footer noise - **readability-lite extraction** — when `<article>`/`<main>`/content-class `<div>` are all missing, `_readability_lite()` picks the highest text-density node (text chars / tag count + `<p>` weighting), avoiding nav/sidebar/footer noise
- **`--fetch-report`** — structured per-URL report to stderr after `--fetch N`: status, WAF type, fallback used, char count, plus adaptive throttle stats and a JSON summary line - **`--fetch-report`** — structured per-URL report to stderr after `--fetch N`: status, WAF type, fallback used, char count, plus adaptive throttle stats and a JSON summary line
- **`--referer`** — set Referer header for fetch requests (defaults to the instance URL when fetching result pages, disguising traffic source) - **`--referer`** — set Referer header for fetch requests (defaults to the instance URL when fetching result pages, disguising traffic source)
- **`--request-delay`** — configurable delay between fetch requests (default 0.3s; adaptive throttling may increase this on failures) - **`--request-delay`** — configurable delay between fetch requests (default 0.3s; adaptive throttling may increase this on failures)
- New fetch result fields: `anti_bot_detected` (bool), `waf_type` (str|null), `fallback_used` (str|null) - New fetch result fields: `anti_bot_detected` (bool), `waf_type` (str|null), `fallback_used` (str|null), `error_code` (str|null, v2.1.1 — structured error code on fetch failure, e.g. `E_RATE_LIMIT` for 429, lets AI agents programmatically distinguish rate-limit from auth/network errors)
**Research mode (v2.1.0)** **Research mode (v2.1.0, enhanced in v2.1.1)**
- **`--research <topic>`** — given a research topic, auto-expands into 5 multi-angle queries (overview/profile/background/works/review) and runs them in sequence. Results include `research_topic` and `research_queries` metadata so AI agents can structure their final report by angle. Mutually exclusive with `--query` and `--queries-file`. Supports all output formats (json/brief/urls/csv). Deterministic expansion (no AI judgment) — same topic always produces same queries, reproducible across processes - **`--research <topic>`** — given a research topic, auto-expands into 5 multi-angle queries (overview/profile/background/works/review) and runs them in sequence. Results include `research_topic` and `research_queries` metadata so AI agents can structure their final report by angle. v2.1.1 enhancements:
- **Cross-angle merge & dedup** — JSON output now includes a top-level `merged_results` field: all per-angle results are combined, deduplicated (same URL collapsing), and sorted, so AI agents can get a unified overview without re-deduplicating themselves. Brief/urls formats append a `[MERGED]` section after the per-angle blocks.
- **Bilingual suffixes** — `expand_research_queries()` now detects whether the topic contains CJK characters. Chinese topics use Chinese suffixes (简介/经历/作品/评价); English topics use English suffixes (profile/background/works/reviews). Avoids low-relevance cross-language combinations like "Python asyncio 经历".
- Mutually exclusive with `--query` and `--queries-file`. Supports all output formats (json/brief/urls/csv). Deterministic expansion (no AI judgment) — same topic always produces same queries, reproducible across processes
**Engineering** **Engineering**
- Shared `common.py` module — unified retry/charset/auth/logging/UA-pool/browser-headers/backoff logic across both scripts - Shared `common.py` module — unified retry/charset/auth/logging/UA-pool/browser-headers/backoff logic across both scripts
@@ -109,7 +112,7 @@ SearXNG is a privacy-respecting metasearch engine that aggregates results from 7
|---------|---------|------------------| |---------|---------|------------------|
| Instance | **required** — via `-i`, `SEARXNG_INSTANCE` env var, or config file | `-i / --instance` | | Instance | **required** — via `-i`, `SEARXNG_INSTANCE` env var, or config file | `-i / --instance` |
| Safe search | **0 (off)** | `-s / --safesearch {0,1,2}` | | Safe search | **0 (off)** | `-s / --safesearch {0,1,2}` |
| Time range | **year** | `-t / --time-range {day,month,year,none}` (none = disabled) | | Time range | **year** | `-t / --time-range {day,week,month,year,none}` (none = disabled; `week` added in v2.1.1 to align with SearXNG API standard) |
| Output format | **json** | `-f / --format {json,brief,urls,csv}` | | Output format | **json** | `-f / --format {json,brief,urls,csv}` |
| Engines | **google,bing,brave,duckduckgo,startpage,wikipedia,wikidata** | `--engines <list>` | | Engines | **google,bing,brave,duckduckgo,startpage,wikipedia,wikidata** | `--engines <list>` |
@@ -400,7 +403,7 @@ All scripts live in `scripts/`; run with `python scripts/<name>.py` from any dir
``` ```
usage: search.py [-h] [--query QUERY] [--instance URL] usage: search.py [-h] [--query QUERY] [--instance URL]
[--categories CATS] [--language LANG] [--pageno N] [--categories CATS] [--language LANG] [--pageno N]
[--time-range {day,month,year,none}] [--safesearch {0,1,2}] [--time-range {day,week,month,year,none}] [--safesearch {0,1,2}]
[--engines E] [--method {GET,POST}] [--max-results N] [--engines E] [--method {GET,POST}] [--max-results N]
[--format {json,brief,urls,csv}] [--snippet-len N] [--format {json,brief,urls,csv}] [--snippet-len N]
[--fetch N] [--fetch-timeout SEC] [--fetch-retries N] [--fetch N] [--fetch-timeout SEC] [--fetch-retries N]
@@ -442,7 +445,7 @@ usage: search.py [-h] [--query QUERY] [--instance URL]
- `--categories general,news` — comma-separated categories (whitespace auto-stripped) - `--categories general,news` — comma-separated categories (whitespace auto-stripped)
- `--language zh-CN` — language filter - `--language zh-CN` — language filter
- `--pageno 1` — page number - `--pageno 1` — page number
- `--time-range {day,month,year,none}` — time filter (default: `year`; `none` disables time filtering) - `--time-range {day,week,month,year,none}` — time filter (default: `year`; `none` disables time filtering; v2.1.1 adds `week` to align with SearXNG API's standard four tiers)
- `--safesearch {0,1,2}` — safe search (default: `0` = off) - `--safesearch {0,1,2}` — safe search (default: `0` = off)
- `--max-results N` — limit number of results (applied AFTER dedup+sort, so the highest-scoring/newest items are kept) - `--max-results N` — limit number of results (applied AFTER dedup+sort, so the highest-scoring/newest items are kept)
- `--sort-by {score,date,engine,none}` — sort results (default: `score` descending; `none` preserves instance order). Applied after dedup, before `--max-results`. HTML-fallback results have no score and keep their order - `--sort-by {score,date,engine,none}` — sort results (default: `score` descending; `none` preserves instance order). Applied after dedup, before `--max-results`. HTML-fallback results have no score and keep their order
+1 -1
View File
@@ -7,6 +7,6 @@ Retry settings and shared HTTP utilities now live in ``common.py`` so that
both ``search.py`` and ``fetch.py`` share one consistent implementation. both ``search.py`` and ``fetch.py`` share one consistent implementation.
""" """
VERSION = "2.1.0" VERSION = "2.1.1"
SCHEMA_VERSION = "1.0" SCHEMA_VERSION = "1.0"
USER_AGENT = f"searxng-cli/{VERSION}" USER_AGENT = f"searxng-cli/{VERSION}"
+5 -3
View File
@@ -751,6 +751,7 @@ def build_wayback_url(url: str) -> str:
# 3. 列表按域名匹配(子域名也算命中) # 3. 列表按域名匹配(子域名也算命中)
HARD_BLOCKED_DOMAINS = frozenset([ HARD_BLOCKED_DOMAINS = frozenset([
"www.baidu.com", # 百度搜索:强反爬 + Cookie 检测
"baike.baidu.com", # 百度百科:强反爬 + Cookie 检测 "baike.baidu.com", # 百度百科:强反爬 + Cookie 检测
"zhidao.baidu.com", # 百度知道:同上 "zhidao.baidu.com", # 百度知道:同上
"tieba.baidu.com", # 百度贴吧:同上 "tieba.baidu.com", # 百度贴吧:同上
@@ -764,15 +765,16 @@ HARD_BLOCKED_DOMAINS = frozenset([
"www.douban.com", # 豆瓣主站 "www.douban.com", # 豆瓣主站
"book.douban.com", # 豆瓣读书 "book.douban.com", # 豆瓣读书
"movie.douban.com", # 豆瓣电影 "movie.douban.com", # 豆瓣电影
"tieba.baidu.com", # 百度贴吧(重复,确保子域匹配)
]) ])
# 部分域名需要子域匹配(如 *.zhihu.com, *.weibo.com, *.douban.com # 需要子域匹配的域名(如 *.zhihu.com, *.weibo.com, *.douban.com
# 注意:baidu.com 不在此列——其反爬子域(www/baike/zhidao/wenku/tieba)已在
# 上方精确列表中,而 pan.baidu.com(网盘)/cloud.baidu.com(智能云)等
# 子域可正常抓取,整体匹配会误伤。zhihu/weibo/douban 的子域基本都被墙。
_SUBDOMAIN_BLOCKED = frozenset([ _SUBDOMAIN_BLOCKED = frozenset([
"zhihu.com", "zhihu.com",
"weibo.com", "weibo.com",
"douban.com", "douban.com",
"baidu.com",
]) ])
+24
View File
@@ -9,6 +9,7 @@ for improved extraction quality (optional, falls back to stdlib).
""" """
import argparse import argparse
import gzip
import logging import logging
import random import random
import re import re
@@ -16,6 +17,7 @@ import sys
import time import time
import urllib.error import urllib.error
import urllib.request import urllib.request
import zlib
from collections import namedtuple from collections import namedtuple
from html.parser import HTMLParser from html.parser import HTMLParser
from pathlib import Path from pathlib import Path
@@ -737,6 +739,28 @@ def fetch_url(url: str, timeout=15, user_agent: str = None,
content_type = resp.headers.get("Content-Type", "") content_type = resp.headers.get("Content-Type", "")
final_url = resp.geturl() final_url = resp.geturl()
# v2.1.1 修复:stdlib urllib 不自动解压 gzip/deflate
# 服务器返回压缩字节流时 raw.decode() 会失败 →
# errors="replace" → 全页 U+FFFD 乱码。
# requests 库会自动处理 Content-Encoding,但 stdlib 不会。
# 此前该 bug 被沙箱伪响应掩盖(两者都产生 U+FFFD),
# 实际在无 requests 的真实环境中会复现。
content_encoding = (resp.headers.get("Content-Encoding", "")
.lower().strip())
if content_encoding and raw:
try:
if "gzip" in content_encoding:
raw = gzip.decompress(raw)
elif "deflate" in content_encoding:
# deflate 可能是 zlib 包装或裸 deflate
try:
raw = zlib.decompress(raw)
except zlib.error:
raw = zlib.decompress(raw, -zlib.MAX_WBITS)
except (OSError, zlib.error) as e:
logger.debug(f" decompress failed ({content_encoding}): {e}")
# 解压失败保留原 raw,让下游 decode 兜底
if encoding: if encoding:
charset = encoding charset = encoding
else: else:
+148 -17
View File
@@ -36,9 +36,11 @@ from common import (
build_auth_headers, build_auth_headers,
build_wayback_url, build_wayback_url,
classify_error, classify_error,
compute_backoff_delay,
emit_progress, emit_progress,
force_utf8_stdout, force_utf8_stdout,
is_hard_blocked_domain, is_hard_blocked_domain,
parse_retry_after,
resolve_auth_basic, resolve_auth_basic,
resolve_auth_bearer, resolve_auth_bearer,
set_progress_enabled, set_progress_enabled,
@@ -476,7 +478,12 @@ def _cfg_float(config: dict, key: str, default: float) -> float:
# ----- Retry logic ----- # ----- Retry logic -----
def _retry_with_backoff(fn, max_retries: int = MAX_RETRIES, base_delay: float = RETRY_BACKOFF_BASE): def _retry_with_backoff(fn, max_retries: int = MAX_RETRIES, base_delay: float = RETRY_BACKOFF_BASE):
"""Call fn with exponential backoff + jitter on transient failures.""" """Call fn with exponential backoff + jitter on transient failures.
v2.1.0 修复:复用 common.compute_backoff_delay(带 60s 封顶),
避免高重试次数(如 --retry 10)时 1.5*2^10=1536s 卡死进程。
同时遵守 Retry-After header429/503),与 fetch.py 保持一致。
"""
last_error = None last_error = None
for attempt in range(max_retries + 1): for attempt in range(max_retries + 1):
try: try:
@@ -485,15 +492,23 @@ def _retry_with_backoff(fn, max_retries: int = MAX_RETRIES, base_delay: float =
if e.code in RETRYABLE_STATUS: # 403 (UA block) + 429 + 5xx if e.code in RETRYABLE_STATUS: # 403 (UA block) + 429 + 5xx
last_error = e last_error = e
if attempt < max_retries: if attempt < max_retries:
delay = base_delay * (2 ** attempt) + random.uniform(0, 1) # 429/503:遵守 Retry-After header,避免触发更严厉限流
logger.info(f" HTTP {e.code}, retrying in {delay:.1f}s... (attempt {attempt+1}/{max_retries})") retry_after_sec = 0.0
if e.code in (429, 503) and e.headers:
retry_after_sec = parse_retry_after(
e.headers.get("Retry-After", ""))
delay = max(retry_after_sec,
compute_backoff_delay(attempt, base=base_delay))
logger.info(f" HTTP {e.code}, retrying in {delay:.1f}s... "
f"(attempt {attempt+1}/{max_retries})"
f"{f' Retry-After={retry_after_sec:.1f}s' if retry_after_sec > 0 else ''}")
time.sleep(delay) time.sleep(delay)
continue continue
raise raise
except (urllib.error.URLError, OSError) as e: except (urllib.error.URLError, OSError) as e:
last_error = e last_error = e
if attempt < max_retries: if attempt < max_retries:
delay = base_delay * (2 ** attempt) + random.uniform(0, 1) delay = compute_backoff_delay(attempt, base=base_delay)
logger.info(f" Connection error ({e}), retrying in {delay:.1f}s...") logger.info(f" Connection error ({e}), retrying in {delay:.1f}s...")
time.sleep(delay) time.sleep(delay)
continue continue
@@ -850,6 +865,7 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
# 主抓取 # 主抓取
result = None result = None
error_msg = None error_msg = None
error_code = None
try: try:
result = fetch_url( result = fetch_url(
url, timeout=timeout, auth_headers=auth_headers, url, timeout=timeout, auth_headers=auth_headers,
@@ -858,6 +874,9 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
) )
except Exception as e: except Exception as e:
error_msg = str(e) if str(e) else e.__class__.__name__ error_msg = str(e) if str(e) else e.__class__.__name__
# v2.1.0:结构化错误码,让 AdaptiveThrottle 能用 error_code 检测 429
# 而非字符串匹配("Too Many Requests" 不含 "429" 会漏判)
error_code = classify_error(e)
# 反爬检测(v2.0.0 增强:全文档扫描 + WAF 指纹库) # 反爬检测(v2.0.0 增强:全文档扫描 + WAF 指纹库)
# 必须在 Wayback 兜底判断之前执行:Cloudflare 质询页常返回 HTTP 200 # 必须在 Wayback 兜底判断之前执行:Cloudflare 质询页常返回 HTTP 200
@@ -914,6 +933,7 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
return { return {
"url": url, "status": "error", "url": url, "status": "error",
"error": error_msg or "unknown error", "error": error_msg or "unknown error",
"error_code": error_code,
"text": "", "text_length": 0, "truncated": False, "text": "", "text_length": 0, "truncated": False,
"anti_bot_detected": False, "waf_type": None, "anti_bot_detected": False, "waf_type": None,
"fallback_used": None, "fallback_used": None,
@@ -924,6 +944,7 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
return { return {
"url": url, "final_url": result.final_url, "status": "error", "url": url, "final_url": result.final_url, "status": "error",
"error": f"Bot protection detected ({waf_type})", "error": f"Bot protection detected ({waf_type})",
"error_code": E_PARSE,
"text": "", "text_length": 0, "truncated": False, "text": "", "text_length": 0, "truncated": False,
"anti_bot_detected": True, "waf_type": waf_type, "anti_bot_detected": True, "waf_type": waf_type,
"fallback_used": fallback_used, "fallback_used": fallback_used,
@@ -1157,12 +1178,23 @@ class AdaptiveThrottle:
self._concurrency = min(self._initial_concurrency, self._concurrency = min(self._initial_concurrency,
self._concurrency * 2) self._concurrency * 2)
def report_failure(self, error_msg: str = "") -> None: def report_failure(self, error_msg: str = "",
error_code: str = None) -> None:
"""报告一次失败,触发自适应退避。
v2.1.0:优先用结构化 error_code 检测 429/限流(E_RATE_LIMIT),
回退到字符串匹配兼容旧调用方。原代码仅检查 "429" 字面量,
"Too Many Requests" 会漏判。
"""
with self._lock: with self._lock:
self._consecutive_successes = 0 self._consecutive_successes = 0
self._consecutive_failures += 1 self._consecutive_failures += 1
# 429 → 全局暂停(调用方会从 error_msg 提取秒数,这里只标记) # 429/限流 → 全局暂停 30s
if "429" in error_msg.lower(): # 优先用 error_code,回退到字符串匹配(兼容无 error_code 的旧调用)
is_rate_limit = (error_code == E_RATE_LIMIT or
"429" in error_msg.lower() or
"rate limit" in error_msg.lower())
if is_rate_limit:
self._global_pause_until = time.monotonic() + 30.0 self._global_pause_until = time.monotonic() + 30.0
# 连续 3 次失败 → 退避 + 降并发 # 连续 3 次失败 → 退避 + 降并发
if self._consecutive_failures >= 3: if self._consecutive_failures >= 3:
@@ -1263,7 +1295,8 @@ def fetch_top_results(results: dict, count: int, timeout: int = 10,
f"{trunc}{ua_note}{fb_note})") f"{trunc}{ua_note}{fb_note})")
else: else:
err_count[0] += 1 err_count[0] += 1
throttle.report_failure(result.get("error", "")) throttle.report_failure(result.get("error", ""),
error_code=result.get("error_code"))
# 统计反爬拦截 # 统计反爬拦截
if result.get("anti_bot_detected"): if result.get("anti_bot_detected"):
anti_bot_count[0] += 1 anti_bot_count[0] += 1
@@ -1279,6 +1312,7 @@ def fetch_top_results(results: dict, count: int, timeout: int = 10,
except Exception as e: except Exception as e:
u = future_map[future] u = future_map[future]
fetched.append({"url": u, "status": "error", "error": str(e), fetched.append({"url": u, "status": "error", "error": str(e),
"error_code": classify_error(e),
"text": "", "text_length": 0, "truncated": False, "text": "", "text_length": 0, "truncated": False,
"anti_bot_detected": False, "waf_type": None, "anti_bot_detected": False, "waf_type": None,
"fallback_used": None}) "fallback_used": None})
@@ -1581,6 +1615,45 @@ def _build_params(query: str, args) -> dict:
return params return params
def _warn_unresponsive_engines(results: dict, query: str,
result_count: int = None) -> None:
"""检测并提示实例侧引擎挂起/限流(v2.1.1)。
SearXNG JSON API 返回的 ``unresponsive_engines`` 字段格式为::
[["brave", "Suspended: too many requests"],
["duckduckgo", "CAPTCHA"]]
当该字段非空时,说明实例内多个引擎被上游限流挂起。此时:
1. 用 logger.warning 输出挂起的引擎列表及原因(到 stderr,
不污染 stdout 数据流)
2. 如果结果数较少,建议用 --engines 限定未挂起引擎
这是真实运营问题(见执行问题记录 #3):连续查询后 brave/duckduckgo/
startpage 等引擎会被上游限流挂起,导致结果骤减或全空。让用户及时
感知引擎状态,避免误判为"无结果"而反复重试触发更严厉限流。
"""
unresponsive = results.get("unresponsive_engines", [])
if not unresponsive:
return
# 格式化引擎列表:兼容 [engine, reason] 和 [engine] 两种格式
parts = []
for entry in unresponsive:
if isinstance(entry, (list, tuple)) and len(entry) >= 2:
parts.append(f"{entry[0]} ({entry[1]})")
elif isinstance(entry, (list, tuple)) and len(entry) == 1:
parts.append(str(entry[0]))
else:
parts.append(str(entry))
engine_list = ", ".join(parts)
logger.warning(f"Instance engines unresponsive: {engine_list}")
# 结果数少 + 引擎挂起 → 建议规避
if result_count is not None and result_count < 3 and len(unresponsive) >= 2:
# 找出可能未挂起的常见引擎提示
logger.warning("Hint: multiple engines suspended — consider using "
"--engines to target responsive ones, or wait before retrying")
def _run_single_query(query: str, args, instance_urls: list, def _run_single_query(query: str, args, instance_urls: list,
auth_headers: dict, ttl_seconds: int): auth_headers: dict, ttl_seconds: int):
"""Run one query end-to-end: search → limit → domain-filter → fetch. """Run one query end-to-end: search → limit → domain-filter → fetch.
@@ -1618,6 +1691,12 @@ def _run_single_query(query: str, args, instance_urls: list,
logger.info(f"[cache stored] q={query!r} TTL={args.cache_ttl}min") logger.info(f"[cache stored] q={query!r} TTL={args.cache_ttl}min")
emit_progress("cache_store", query=query, ttl=args.cache_ttl) emit_progress("cache_store", query=query, ttl=args.cache_ttl)
# v2.1.1:检测实例侧引擎挂起/限流(仅在实时查询后提示,cache hit 时
# unresponsive_engines 信息可能已过期)
if cached is None:
_warn_unresponsive_engines(results, query,
result_count=len(results.get("results", [])))
# Dedup (default on; --no-dedup disables) then sort, both BEFORE limit # Dedup (default on; --no-dedup disables) then sort, both BEFORE limit
# so --max-results keeps the highest-scoring / newest items. # so --max-results keeps the highest-scoring / newest items.
if not args.no_dedup: if not args.no_dedup:
@@ -1821,17 +1900,31 @@ def _read_queries_file(path: str) -> list:
# 给定一个主题,自动扩展多角度查询词,复用批量搜索逻辑。 # 给定一个主题,自动扩展多角度查询词,复用批量搜索逻辑。
# 扩展策略是确定性规则(不做 AI 判断),覆盖人物/主题/事件的通用研究维度。 # 扩展策略是确定性规则(不做 AI 判断),覆盖人物/主题/事件的通用研究维度。
# 研究角度定义:(角度标识, 后缀) # 研究角度定义:(角度标识, 中文后缀, 英文后缀)
# 顺序代表搜索优先级——基本信息优先,评价争议最后。 # 顺序代表搜索优先级——基本信息优先,评价争议最后。
# v2.1.0:支持中英文双语后缀,根据主题语言自动选择。
# 中文主题用中文后缀("简介"/"经历"等),英文主题用英文后缀
# "profile"/"background"等),避免 "Python asyncio 经历" 这类
# 跨语言组合在英文引擎上匹配度低的问题。angle 标识符保持英文,
# 便于 AI Agent 程序化处理。
_RESEARCH_ANGLES = [ _RESEARCH_ANGLES = [
("overview", ""), # 主题本身:最直接的搜索 ("overview", "", ""), # 主题本身:最直接的搜索
("profile", "简介"), # 基本信息:百科式介绍 ("profile", "简介", "profile"), # 基本信息:百科式介绍
("background", "经历"), # 背景经历:生平/历史 ("background", "经历", "background"), # 背景经历:生平/历史
("works", "作品"), # 作品成就:产出物 ("works", "作品", "works"), # 作品成就:产出物
("review", "评价"), # 评价争议:外界看法 ("review", "评价", "reviews"), # 评价争议:外界看法
] ]
def _is_chinese_topic(topic: str) -> bool:
"""检测主题是否包含中文字符(CJK 统一表意文字范围)。
用于 expand_research_queries 选择中文还是英文后缀。
纯英文主题(如 "Python asyncio")返回 False,用英文后缀。
"""
return bool(re.search(r'[\u4e00-\u9fff]', topic))
def expand_research_queries(topic: str) -> list: def expand_research_queries(topic: str) -> list:
"""将研究主题扩展为多角度查询词列表。 """将研究主题扩展为多角度查询词列表。
@@ -1843,6 +1936,10 @@ def expand_research_queries(topic: str) -> list:
4. works — 作品成就 4. works — 作品成就
5. review — 评价争议 5. review — 评价争议
v2.1.0 修复:根据主题语言自动切换后缀。含中文字符的主题用中文后缀
"七森莉莉 简介"),纯英文主题用英文后缀("Python asyncio profile"),
避免跨语言组合在搜索引擎上匹配度低。
返回 [(angle, query), ...] 列表,angle 用于结果标注。 返回 [(angle, query), ...] 列表,angle 用于结果标注。
确定性规则,不依赖 AI 判断——确保跨进程可复现,AI Agent 可预期。 确定性规则,不依赖 AI 判断——确保跨进程可复现,AI Agent 可预期。
@@ -1850,8 +1947,10 @@ def expand_research_queries(topic: str) -> list:
topic = topic.strip() topic = topic.strip()
if not topic: if not topic:
return [] return []
use_chinese = _is_chinese_topic(topic)
queries = [] queries = []
for angle, suffix in _RESEARCH_ANGLES: for angle, cn_suffix, en_suffix in _RESEARCH_ANGLES:
suffix = cn_suffix if use_chinese else en_suffix
query = f"{topic} {suffix}".strip() query = f"{topic} {suffix}".strip()
queries.append((angle, query)) queries.append((angle, query))
return queries return queries
@@ -1927,9 +2026,11 @@ Use --config FILE to load a non-default config file (overrides the auto-discover
help="Language code (e.g. en, zh-CN, de)") help="Language code (e.g. en, zh-CN, de)")
parser.add_argument("--pageno", "-p", type=int, default=1, parser.add_argument("--pageno", "-p", type=int, default=1,
help="Page number (default: 1)") help="Page number (default: 1)")
parser.add_argument("--time-range", "-t", choices=["day", "month", "year", "none"], parser.add_argument("--time-range", "-t",
choices=["day", "week", "month", "year", "none"],
default=config.get("time_range", "year"), default=config.get("time_range", "year"),
help="Time range filter (default: year; 'none' disables filtering)") help="Time range filter (default: year; 'none' disables filtering). "
"SearXNG API standard four tiers: day/week/month/year")
parser.add_argument("--safesearch", "-s", type=int, choices=[0, 1, 2], parser.add_argument("--safesearch", "-s", type=int, choices=[0, 1, 2],
default=_cfg_int(config, "safesearch", 0), default=_cfg_int(config, "safesearch", 0),
help="Safe search: 0=off, 1=moderate, 2=strict (default: 0=off)") help="Safe search: 0=off, 1=moderate, 2=strict (default: 0=off)")
@@ -2239,6 +2340,23 @@ Use --config FILE to load a non-default config file (overrides the auto-discover
batch.append({"query": q, "angle": angle, "status": "ok", batch.append({"query": q, "angle": angle, "status": "ok",
"results": results}) "results": results})
# v2.1.0 修复:跨角度合并去重
# 文档承诺 "Results are merged and deduplicated",原代码只输出 per-angle
# 结果,同一 URL 可能出现在多个角度中。此处合并所有成功角度的 results,
# 去重后作为 merged_results 字段输出,让 AI Agent 既能按角度组织报告,
# 也能获得去重后的总览。
merged = {"query": topic, "results": []}
for br in batch:
if br.get("status") == "ok" and "results" in br:
merged["results"].extend(br["results"].get("results", []))
if merged["results"]:
deduplicate_results(merged)
sort_results(merged, args.sort_by)
if args.max_results:
merged["results"] = merged["results"][:args.max_results]
merged_count = len(merged["results"])
logger.info(f"Research merged: {merged_count} unique results after dedup")
# 输出 # 输出
if args.format == "json": if args.format == "json":
output = json.dumps({ output = json.dumps({
@@ -2248,6 +2366,7 @@ Use --config FILE to load a non-default config file (overrides the auto-discover
{"angle": a, "query": q} for a, q in research_queries {"angle": a, "query": q} for a, q in research_queries
], ],
"queries": batch, "queries": batch,
"merged_results": merged,
}, indent=2, ensure_ascii=False) }, indent=2, ensure_ascii=False)
elif args.format == "csv": elif args.format == "csv":
import csv as csv_mod import csv as csv_mod
@@ -2276,12 +2395,18 @@ Use --config FILE to load a non-default config file (overrides the auto-discover
output = out.getvalue().rstrip() output = out.getvalue().rstrip()
elif args.format == "urls": elif args.format == "urls":
parts = [] parts = []
# 先输出 per-angle 结果
for br in batch: for br in batch:
parts.append(f"# [{br.get('angle', '?')}] {br['query']}") parts.append(f"# [{br.get('angle', '?')}] {br['query']}")
if "results" in br: if "results" in br:
parts.append(format_urls(br["results"])) parts.append(format_urls(br["results"]))
else: else:
parts.append(f"# [ERROR: {br['error']}]") parts.append(f"# [ERROR: {br['error']}]")
# 再输出合并去重后的总览
if merged_count > 0:
parts.append("")
parts.append(f"# [MERGED] {topic} ({merged_count} unique results)")
parts.append(format_urls(merged))
output = "\n".join(parts) output = "\n".join(parts)
else: # brief else: # brief
parts = [] parts = []
@@ -2294,6 +2419,12 @@ Use --config FILE to load a non-default config file (overrides the auto-discover
else: else:
parts.append(f"[ERROR: {br['error']}]") parts.append(f"[ERROR: {br['error']}]")
parts.append("") parts.append("")
# 合并去重后的总览
if merged_count > 0:
parts.append("=" * 60)
parts.append(f"[MERGED] {topic} ({merged_count} unique results)")
parts.append("=" * 60)
parts.append(format_brief(merged))
output = "\n".join(parts) output = "\n".join(parts)
if args.output: if args.output:
+3 -2
View File
@@ -44,7 +44,7 @@ class TestExpandResearchQueries:
def test_angles_match_definition(self): def test_angles_match_definition(self):
queries = expand_research_queries("test") queries = expand_research_queries("test")
angles = [a for a, _ in queries] angles = [a for a, _ in queries]
assert angles == [a for a, _ in _RESEARCH_ANGLES] assert angles == [a for a, _, _ in _RESEARCH_ANGLES]
def test_empty_topic_returns_empty(self): def test_empty_topic_returns_empty(self):
assert expand_research_queries("") == [] assert expand_research_queries("") == []
@@ -65,7 +65,8 @@ class TestExpandResearchQueries:
def test_multi_word_topic(self): def test_multi_word_topic(self):
queries = expand_research_queries("Python asyncio tutorial") queries = expand_research_queries("Python asyncio tutorial")
assert queries[0] == ("overview", "Python asyncio tutorial") assert queries[0] == ("overview", "Python asyncio tutorial")
assert "简介" in queries[1][1] # v2.1.0:英文主题用英文后缀(避免跨语言组合匹配度低)
assert "profile" in queries[1][1]
# ===== --research CLI mutex checks ===== # ===== --research CLI mutex checks =====
+43
View File
@@ -14,6 +14,7 @@ from search import (
_merge_headers, _merge_headers,
_normalize_csv, _normalize_csv,
_read_queries_file, _read_queries_file,
_warn_unresponsive_engines,
deduplicate_results, deduplicate_results,
filter_results_by_domain, filter_results_by_domain,
load_config, load_config,
@@ -224,6 +225,14 @@ def test_build_params_time_range_none_excluded():
assert "time_range" not in p assert "time_range" not in p
def test_build_params_time_range_week_included():
"""v2.1.1: 'week' is a valid SearXNG API time_range tier and must be
passed through to params. Previously argparse choices omitted 'week',
forcing users to use config-file workaround."""
p = _build_params("x", _Args(time_range="week"))
assert p["time_range"] == "week"
def test_build_params_pageno_as_string(): def test_build_params_pageno_as_string():
p = _build_params("x", _Args(pageno=3)) p = _build_params("x", _Args(pageno=3))
assert p["pageno"] == "3" assert p["pageno"] == "3"
@@ -252,6 +261,40 @@ def test_merge_headers_all_none():
assert _merge_headers(None, None) == {} assert _merge_headers(None, None) == {}
# ----- _warn_unresponsive_engines (v2.1.1) -----
def test_warn_unresponsive_no_field():
"""No unresponsive_engines field → no warning, no exception."""
_warn_unresponsive_engines({}, "test")
def test_warn_unresponsive_empty_list():
"""Empty unresponsive_engines list → no warning."""
_warn_unresponsive_engines({"unresponsive_engines": []}, "test")
def test_warn_unresponsive_with_reasons():
"""[engine, reason] format should be formatted as 'engine (reason)'."""
results = {"unresponsive_engines": [
["brave", "Suspended: too many requests"],
["duckduckgo", "CAPTCHA"],
]}
# Should not raise; logger.warning is called internally
_warn_unresponsive_engines(results, "test", result_count=1)
def test_warn_unresponsive_engine_only():
"""[engine] single-element format should be handled."""
results = {"unresponsive_engines": [["brave"]]}
_warn_unresponsive_engines(results, "test", result_count=5)
def test_warn_unresponsive_string_format():
"""Plain string entries (non-list) should be handled gracefully."""
results = {"unresponsive_engines": ["brave", "duckduckgo"]}
_warn_unresponsive_engines(results, "test", result_count=0)
# ----- deduplicate_results ----- # ----- deduplicate_results -----
def test_dedup_removes_exact_duplicate_url(): def test_dedup_removes_exact_duplicate_url():
+13 -4
View File
@@ -113,12 +113,21 @@ class TestIsHardBlockedDomain:
assert is_hard_blocked_domain("https://github.com/python/cpython") is False assert is_hard_blocked_domain("https://github.com/python/cpython") is False
def test_baidu_search_not_blocked(self): def test_baidu_search_not_blocked(self):
"""baidu.com search page is NOT in the hard-blocked list — only subdomains """www.baidu.com search page is in the hard-blocked list.
like baike.baidu.com, zhidao.baidu.com are.""" v2.1.0: moved from _SUBDOMAIN_BLOCKED (baidu.com) to HARD_BLOCKED_DOMAINS
# Actually, baidu.com is in _SUBDOMAIN_BLOCKED, so www.baidu.com matches. (www.baidu.com) to avoid blocking pan.baidu.com / cloud.baidu.com."""
# This is intentional — Baidu's main search also has strong anti-bot.
assert is_hard_blocked_domain("https://www.baidu.com/s?wd=test") is True assert is_hard_blocked_domain("https://www.baidu.com/s?wd=test") is True
def test_baidu_pan_not_blocked(self):
"""pan.baidu.com (百度网盘) should NOT be blocked after v2.1.0 fix.
Previously matched by _SUBDOMAIN_BLOCKED 'baidu.com' — now only
www/baike/zhidao/tieba/wenku are in the precise list."""
assert is_hard_blocked_domain("https://pan.baidu.com/s/abc123") is False
def test_baidu_cloud_not_blocked(self):
"""cloud.baidu.com (百度智能云) should NOT be blocked after v2.1.0 fix."""
assert is_hard_blocked_domain("https://cloud.baidu.com/product/abc") is False
def test_empty_url(self): def test_empty_url(self):
assert is_hard_blocked_domain("") is False assert is_hard_blocked_domain("") is False