fix(v2.0.1): 修复反爬误判 + Wayback 兜底未触发两个 bug
Bug 1: Wayback 兜底未触发 (search.py fetch_page) - 根因: Cloudflare JS 质询页常返回 HTTP 200 (非 403), 主抓取 result 非 None - _should_try_fallback 在 result 非 None 时直接返回 False, 跳过兜底 - 反爬检测在兜底判断之后执行, 错过兜底入口 - 修复: 反爬检测提前到兜底判断之前, anti_bot_detected=True 也触发 Wayback - Wayback 结果重新做反爬检测 (防御性) Bug 2: WAF 指纹库误判正常内容 (search.py WAF_FINGERPRINTS) - 根因: 裸公司名 (cloudflare/akamai) 和宽泛词 (captcha/challenge/dd-) 做全文匹配 - DataCamp 文章引用 cloudflare.com 文档链接 -> 误判为 cloudflare WAF - 'coding challenges' 正常内容 -> 误判为 generic 反爬 - Wayback 归档正文被误判, 兜底返回的有效内容被丢弃 - 修复: 移除裸公司名和宽泛词, 改用技术标识符 (cf-ray/incap_ses/bm_sz 等) - 通用文案用完整短语 (please complete the captcha) 替代单词 - 增加 <title> 标签精准检测 (反爬页 title 是特征文案, 误判率极低) - 新增 Anubis 反爬系统检测 (anubis_challenge/miserere) 真实测试验证 (search.metona.cn 实例): - v2.0.0: --fetch 3 全部失败 (3 ERR: cloudflare/generic, Wayback 未触发) - v2.0.1: --fetch 3 全部成功 (3 OK: 38100/93442/1945 chars, UA 轮换绕过 Cloudflare) 测试: 458 个全部通过 (新增 7 个测试覆盖修复行为)
This commit is contained in:
@@ -46,9 +46,12 @@ def test_blocked_page_checks_first_2000_chars():
|
||||
v2.0.0 changed this to full-document scanning because large anti-bot
|
||||
pages (e.g. Cloudflare challenges with big JS blobs) may place the
|
||||
telltale keyword beyond the 2000-char boundary.
|
||||
|
||||
v2.0.1: 用完整短语 'please complete the captcha' 替代裸 'captcha',
|
||||
避免正常内容误判。
|
||||
"""
|
||||
padding = "x" * 2500
|
||||
html = f"<html>{padding}captcha</html>"
|
||||
html = f"<html>{padding}please complete the captcha</html>"
|
||||
# v2.0.0: now detected (was: not detected)
|
||||
assert _is_blocked_page(html)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user