fix(v2.0.1): 修复反爬误判 + Wayback 兜底未触发两个 bug

Bug 1: Wayback 兜底未触发 (search.py fetch_page)

- 根因: Cloudflare JS 质询页常返回 HTTP 200 (非 403), 主抓取 result 非 None

- _should_try_fallback 在 result 非 None 时直接返回 False, 跳过兜底

- 反爬检测在兜底判断之后执行, 错过兜底入口

- 修复: 反爬检测提前到兜底判断之前, anti_bot_detected=True 也触发 Wayback

- Wayback 结果重新做反爬检测 (防御性)

Bug 2: WAF 指纹库误判正常内容 (search.py WAF_FINGERPRINTS)

- 根因: 裸公司名 (cloudflare/akamai) 和宽泛词 (captcha/challenge/dd-) 做全文匹配

- DataCamp 文章引用 cloudflare.com 文档链接 -> 误判为 cloudflare WAF

- 'coding challenges' 正常内容 -> 误判为 generic 反爬

- Wayback 归档正文被误判, 兜底返回的有效内容被丢弃

- 修复: 移除裸公司名和宽泛词, 改用技术标识符 (cf-ray/incap_ses/bm_sz 等)

- 通用文案用完整短语 (please complete the captcha) 替代单词

- 增加 <title> 标签精准检测 (反爬页 title 是特征文案, 误判率极低)

- 新增 Anubis 反爬系统检测 (anubis_challenge/miserere)

真实测试验证 (search.metona.cn 实例):

- v2.0.0: --fetch 3 全部失败 (3 ERR: cloudflare/generic, Wayback 未触发)

- v2.0.1: --fetch 3 全部成功 (3 OK: 38100/93442/1945 chars, UA 轮换绕过 Cloudflare)

测试: 458 个全部通过 (新增 7 个测试覆盖修复行为)
This commit is contained in:
2026-08-01 21:44:58 +08:00
parent 28ff7c0a48
commit d9a08716bc
7 changed files with 215 additions and 66 deletions
+39 -8
View File
@@ -85,6 +85,7 @@ def test_detect_akamai_reference():
# ----- Generic detection -----
def test_detect_generic_captcha():
"""v2.0.1:收窄为完整短语 'please complete the captcha',避免正文误判。"""
html = "<html><body>Please complete the CAPTCHA</body></html>"
assert _detect_anti_bot(html) == "generic"
@@ -95,7 +96,15 @@ def test_detect_generic_verify_human():
def test_detect_generic_access_denied():
html = "<html><body>Access Denied</body></html>"
"""v2.0.1:裸 'access denied' 从全文档扫描移除(避免权限文章误判),
改由 <title> 标签检测覆盖。反爬页 title 常为 'Access Denied'"""
html = "<html><head><title>Access Denied</title></head><body>Access Denied</body></html>"
assert _detect_anti_bot(html) == "generic"
def test_detect_generic_access_denied_sucuri():
"""v2.0.1Sucuri WAF 的完整文案在全文档扫描中保留。"""
html = "<html><body>Access Denied - Sucuri Website Firewall</body></html>"
assert _detect_anti_bot(html) == "generic"
@@ -130,7 +139,8 @@ def test_detect_anti_bot_beyond_2000_chars():
def test_detect_anti_bot_large_page_end():
"""反爬关键词在文档末尾也能检测到。"""
padding = "y" * 5000
html = f"<html><body>{padding}captcha</body></html>"
# v2.0.1:用完整短语而非裸 captcha,避免误判
html = f"<html><body>{padding}please complete the captcha</body></html>"
assert _detect_anti_bot(html) == "generic"
@@ -149,14 +159,34 @@ def test_detect_anti_bot_empty_content():
def test_detect_anti_bot_article_mentions_captcha_in_context():
"""文章讨论 captcha 但不是反爬页(上下文判断的局限——接受误报)
"""v2.0.1 修复:文章讨论 captcha 但不是反爬页,不再误判
注意:当前实现是关键词匹配,无法区分"讨论 captcha 的文章"
"captcha 拦截页"。这是已知局限,测试记录此行为。
v2.0.0 用裸 "captcha" 做全文匹配,导致"讨论 captcha 工作原理的文章"
被误判为反爬页。v2.0.1 收窄为 "please complete the captcha" 等完整短语,
正常讨论 captcha 的文章不再触发。
"""
html = "<html><body><p>This article explains how CAPTCHA works.</p></body></html>"
# 关键词匹配会误报为 generic
assert _detect_anti_bot(html) == "generic"
# v2.0.1:收窄后不再误判
assert _detect_anti_bot(html) is None
def test_detect_anti_bot_article_mentions_cloudflare_in_context():
"""v2.0.1 修复:文章引用 Cloudflare 文档链接,不再误判为 cloudflare WAF。
v2.0.0 用裸 "cloudflare" 做全文匹配,Wayback 归档的 DataCamp 文章
正文里有 cloudflare.com 链接,被误判为反爬页导致兜底失败。
"""
html = ("<html><body><article>"
"<p>Learn more at <a href='https://www.cloudflare.com/learn/'>Cloudflare</a></p>"
"<p>Join our daily coding challenges!</p>"
"</article></body></html>")
assert _detect_anti_bot(html) is None
def test_detect_anti_bot_article_mentions_challenge_in_context():
"""v2.0.1 修复:文章含 'coding challenge' 等正常内容,不再误判。"""
html = "<html><body><p>Daily 5-minute coding challenges.</p></body></html>"
assert _detect_anti_bot(html) is None
# ----- Priority: specialized WAF before generic -----
@@ -197,7 +227,8 @@ def test_waf_fingerprints_all_have_indicators():
def test_is_blocked_page_delegates_to_detect():
"""_is_blocked_page 应委托给 _detect_anti_bot。"""
assert _is_blocked_page("<html>captcha</html>") is True
# v2.0.1:用完整反爬短语而非裸 captcha
assert _is_blocked_page("<html><body>please complete the captcha</body></html>") is True
assert _is_blocked_page("<html>normal content</html>") is False
+4 -1
View File
@@ -46,9 +46,12 @@ def test_blocked_page_checks_first_2000_chars():
v2.0.0 changed this to full-document scanning because large anti-bot
pages (e.g. Cloudflare challenges with big JS blobs) may place the
telltale keyword beyond the 2000-char boundary.
v2.0.1: 用完整短语 'please complete the captcha' 替代裸 'captcha'
避免正常内容误判。
"""
padding = "x" * 2500
html = f"<html>{padding}captcha</html>"
html = f"<html>{padding}please complete the captcha</html>"
# v2.0.0: now detected (was: not detected)
assert _is_blocked_page(html)
+51 -6
View File
@@ -208,22 +208,67 @@ def test_fetch_page_normal_page_no_anti_bot():
def test_fetch_page_anti_bot_triggers_wayback():
"""被反爬拦截后应尝试 Wayback_should_try_fallback 防御性检查)。
"""被反爬拦截后应尝试 Wayback 兜底(v2.0.1 修复)。
当前实现:主抓取成功返回反爬页内容 → result 非 None → 不触发兜底。
此测试记录此行为:反爬页被当作"成功抓取"返回,在 fetch_page 内部检测
v2.0.0 bug:主抓取"成功"返回反爬页(HTTP 200 + Cloudflare 质询)时,
result 非 None 导致 _should_try_fallback 返回 FalseWayback 永不触发
v2.0.1 修复:反爬检测提前到兜底判断之前,反爬阳性也触发 Wayback。
本测试 mock fetch_url 两次调用:
1. 主抓取 → 返回 Cloudflare 质询页
2. Wayback 兜底 → 返回正常归档页
"""
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
wb_html = "<html><body><p>Archived article content.</p></body></html>"
cf_result = _make_result(content=cf_html, final_url="https://example.com")
wb_result = _make_result(content=wb_html,
final_url="https://web.archive.org/web/2024/https://example.com")
with patch("search.fetch_url",
return_value=_make_result(content=cf_html)):
side_effect=[cf_result, wb_result]) as mock_fetch:
result = fetch_page("https://example.com", fallback_enabled=True)
# 反爬页被检测到,标记为 error + anti_bot_detected
# Wayback 兜底成功,反爬标记清除,status=ok
assert result["status"] == "ok"
assert result["anti_bot_detected"] is False
assert result["waf_type"] is None
assert result["fallback_used"] == "wayback"
assert "Archived article content." in result["text"]
# 确认 fetch_url 被调用两次:主抓取 + Wayback
assert mock_fetch.call_count == 2
# 第二次调用应该是 Wayback URL
assert "web.archive.org/web/2/" in mock_fetch.call_args_list[1][0][0]
def test_fetch_page_anti_bot_wayback_also_blocked():
"""主抓取反爬 + Wayback 也反爬/失败 → 最终返回 error。
Wayback 兜底返回 None(失败)时,保留主抓取的反爬检测结果。
"""
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
cf_result = _make_result(content=cf_html, final_url="https://example.com")
with patch("search.fetch_url",
side_effect=[cf_result, RuntimeError("wayback timeout")]):
result = fetch_page("https://example.com", fallback_enabled=True)
# Wayback 失败,保留反爬 error
assert result["status"] == "error"
assert result["anti_bot_detected"] is True
# 但不会触发 Wayback(因为主抓取"成功"了,只是内容是反爬页)
assert result["waf_type"] == "cloudflare"
assert result["fallback_used"] is None
def test_fetch_page_anti_bot_fallback_disabled():
"""--no-fallback 时反爬页直接返回 error,不尝试 Wayback。"""
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
with patch("search.fetch_url",
return_value=_make_result(content=cf_html)) as mock_fetch:
result = fetch_page("https://example.com", fallback_enabled=False)
assert result["status"] == "error"
assert result["anti_bot_detected"] is True
assert result["waf_type"] == "cloudflare"
assert result["fallback_used"] is None
# 只调用一次(主抓取),没有 Wayback
assert mock_fetch.call_count == 1
def test_fetch_page_passes_referer():
"""referer 透传给 fetch_url。"""
with patch("search.fetch_url", return_value=_make_result()) as mock_fetch: