fix(v2.0.1): 修复反爬误判 + Wayback 兜底未触发两个 bug
Bug 1: Wayback 兜底未触发 (search.py fetch_page) - 根因: Cloudflare JS 质询页常返回 HTTP 200 (非 403), 主抓取 result 非 None - _should_try_fallback 在 result 非 None 时直接返回 False, 跳过兜底 - 反爬检测在兜底判断之后执行, 错过兜底入口 - 修复: 反爬检测提前到兜底判断之前, anti_bot_detected=True 也触发 Wayback - Wayback 结果重新做反爬检测 (防御性) Bug 2: WAF 指纹库误判正常内容 (search.py WAF_FINGERPRINTS) - 根因: 裸公司名 (cloudflare/akamai) 和宽泛词 (captcha/challenge/dd-) 做全文匹配 - DataCamp 文章引用 cloudflare.com 文档链接 -> 误判为 cloudflare WAF - 'coding challenges' 正常内容 -> 误判为 generic 反爬 - Wayback 归档正文被误判, 兜底返回的有效内容被丢弃 - 修复: 移除裸公司名和宽泛词, 改用技术标识符 (cf-ray/incap_ses/bm_sz 等) - 通用文案用完整短语 (please complete the captcha) 替代单词 - 增加 <title> 标签精准检测 (反爬页 title 是特征文案, 误判率极低) - 新增 Anubis 反爬系统检测 (anubis_challenge/miserere) 真实测试验证 (search.metona.cn 实例): - v2.0.0: --fetch 3 全部失败 (3 ERR: cloudflare/generic, Wayback 未触发) - v2.0.1: --fetch 3 全部成功 (3 OK: 38100/93442/1945 chars, UA 轮换绕过 Cloudflare) 测试: 458 个全部通过 (新增 7 个测试覆盖修复行为)
This commit is contained in:
+39
-8
@@ -85,6 +85,7 @@ def test_detect_akamai_reference():
|
||||
# ----- Generic detection -----
|
||||
|
||||
def test_detect_generic_captcha():
|
||||
"""v2.0.1:收窄为完整短语 'please complete the captcha',避免正文误判。"""
|
||||
html = "<html><body>Please complete the CAPTCHA</body></html>"
|
||||
assert _detect_anti_bot(html) == "generic"
|
||||
|
||||
@@ -95,7 +96,15 @@ def test_detect_generic_verify_human():
|
||||
|
||||
|
||||
def test_detect_generic_access_denied():
|
||||
html = "<html><body>Access Denied</body></html>"
|
||||
"""v2.0.1:裸 'access denied' 从全文档扫描移除(避免权限文章误判),
|
||||
改由 <title> 标签检测覆盖。反爬页 title 常为 'Access Denied'。"""
|
||||
html = "<html><head><title>Access Denied</title></head><body>Access Denied</body></html>"
|
||||
assert _detect_anti_bot(html) == "generic"
|
||||
|
||||
|
||||
def test_detect_generic_access_denied_sucuri():
|
||||
"""v2.0.1:Sucuri WAF 的完整文案在全文档扫描中保留。"""
|
||||
html = "<html><body>Access Denied - Sucuri Website Firewall</body></html>"
|
||||
assert _detect_anti_bot(html) == "generic"
|
||||
|
||||
|
||||
@@ -130,7 +139,8 @@ def test_detect_anti_bot_beyond_2000_chars():
|
||||
def test_detect_anti_bot_large_page_end():
|
||||
"""反爬关键词在文档末尾也能检测到。"""
|
||||
padding = "y" * 5000
|
||||
html = f"<html><body>{padding}captcha</body></html>"
|
||||
# v2.0.1:用完整短语而非裸 captcha,避免误判
|
||||
html = f"<html><body>{padding}please complete the captcha</body></html>"
|
||||
assert _detect_anti_bot(html) == "generic"
|
||||
|
||||
|
||||
@@ -149,14 +159,34 @@ def test_detect_anti_bot_empty_content():
|
||||
|
||||
|
||||
def test_detect_anti_bot_article_mentions_captcha_in_context():
|
||||
"""文章讨论 captcha 但不是反爬页(上下文判断的局限——接受误报)。
|
||||
"""v2.0.1 修复:文章讨论 captcha 但不是反爬页,不再误判。
|
||||
|
||||
注意:当前实现是关键词匹配,无法区分"讨论 captcha 的文章"和
|
||||
"captcha 拦截页"。这是已知局限,测试记录此行为。
|
||||
v2.0.0 用裸 "captcha" 做全文匹配,导致"讨论 captcha 工作原理的文章"
|
||||
被误判为反爬页。v2.0.1 收窄为 "please complete the captcha" 等完整短语,
|
||||
正常讨论 captcha 的文章不再触发。
|
||||
"""
|
||||
html = "<html><body><p>This article explains how CAPTCHA works.</p></body></html>"
|
||||
# 关键词匹配会误报为 generic
|
||||
assert _detect_anti_bot(html) == "generic"
|
||||
# v2.0.1:收窄后不再误判
|
||||
assert _detect_anti_bot(html) is None
|
||||
|
||||
|
||||
def test_detect_anti_bot_article_mentions_cloudflare_in_context():
|
||||
"""v2.0.1 修复:文章引用 Cloudflare 文档链接,不再误判为 cloudflare WAF。
|
||||
|
||||
v2.0.0 用裸 "cloudflare" 做全文匹配,Wayback 归档的 DataCamp 文章
|
||||
正文里有 cloudflare.com 链接,被误判为反爬页导致兜底失败。
|
||||
"""
|
||||
html = ("<html><body><article>"
|
||||
"<p>Learn more at <a href='https://www.cloudflare.com/learn/'>Cloudflare</a></p>"
|
||||
"<p>Join our daily coding challenges!</p>"
|
||||
"</article></body></html>")
|
||||
assert _detect_anti_bot(html) is None
|
||||
|
||||
|
||||
def test_detect_anti_bot_article_mentions_challenge_in_context():
|
||||
"""v2.0.1 修复:文章含 'coding challenge' 等正常内容,不再误判。"""
|
||||
html = "<html><body><p>Daily 5-minute coding challenges.</p></body></html>"
|
||||
assert _detect_anti_bot(html) is None
|
||||
|
||||
|
||||
# ----- Priority: specialized WAF before generic -----
|
||||
@@ -197,7 +227,8 @@ def test_waf_fingerprints_all_have_indicators():
|
||||
|
||||
def test_is_blocked_page_delegates_to_detect():
|
||||
"""_is_blocked_page 应委托给 _detect_anti_bot。"""
|
||||
assert _is_blocked_page("<html>captcha</html>") is True
|
||||
# v2.0.1:用完整反爬短语而非裸 captcha
|
||||
assert _is_blocked_page("<html><body>please complete the captcha</body></html>") is True
|
||||
assert _is_blocked_page("<html>normal content</html>") is False
|
||||
|
||||
|
||||
|
||||
@@ -46,9 +46,12 @@ def test_blocked_page_checks_first_2000_chars():
|
||||
v2.0.0 changed this to full-document scanning because large anti-bot
|
||||
pages (e.g. Cloudflare challenges with big JS blobs) may place the
|
||||
telltale keyword beyond the 2000-char boundary.
|
||||
|
||||
v2.0.1: 用完整短语 'please complete the captcha' 替代裸 'captcha',
|
||||
避免正常内容误判。
|
||||
"""
|
||||
padding = "x" * 2500
|
||||
html = f"<html>{padding}captcha</html>"
|
||||
html = f"<html>{padding}please complete the captcha</html>"
|
||||
# v2.0.0: now detected (was: not detected)
|
||||
assert _is_blocked_page(html)
|
||||
|
||||
|
||||
@@ -208,22 +208,67 @@ def test_fetch_page_normal_page_no_anti_bot():
|
||||
|
||||
|
||||
def test_fetch_page_anti_bot_triggers_wayback():
|
||||
"""被反爬拦截后也应尝试 Wayback(_should_try_fallback 防御性检查)。
|
||||
"""被反爬拦截后应尝试 Wayback 兜底(v2.0.1 修复)。
|
||||
|
||||
当前实现:主抓取成功返回反爬页内容 → result 非 None → 不触发兜底。
|
||||
此测试记录此行为:反爬页被当作"成功抓取"返回,在 fetch_page 内部检测。
|
||||
v2.0.0 bug:主抓取"成功"返回反爬页(HTTP 200 + Cloudflare 质询)时,
|
||||
result 非 None 导致 _should_try_fallback 返回 False,Wayback 永不触发。
|
||||
v2.0.1 修复:反爬检测提前到兜底判断之前,反爬阳性也触发 Wayback。
|
||||
|
||||
本测试 mock fetch_url 两次调用:
|
||||
1. 主抓取 → 返回 Cloudflare 质询页
|
||||
2. Wayback 兜底 → 返回正常归档页
|
||||
"""
|
||||
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
|
||||
wb_html = "<html><body><p>Archived article content.</p></body></html>"
|
||||
cf_result = _make_result(content=cf_html, final_url="https://example.com")
|
||||
wb_result = _make_result(content=wb_html,
|
||||
final_url="https://web.archive.org/web/2024/https://example.com")
|
||||
with patch("search.fetch_url",
|
||||
return_value=_make_result(content=cf_html)):
|
||||
side_effect=[cf_result, wb_result]) as mock_fetch:
|
||||
result = fetch_page("https://example.com", fallback_enabled=True)
|
||||
# 反爬页被检测到,标记为 error + anti_bot_detected
|
||||
# Wayback 兜底成功,反爬标记清除,status=ok
|
||||
assert result["status"] == "ok"
|
||||
assert result["anti_bot_detected"] is False
|
||||
assert result["waf_type"] is None
|
||||
assert result["fallback_used"] == "wayback"
|
||||
assert "Archived article content." in result["text"]
|
||||
# 确认 fetch_url 被调用两次:主抓取 + Wayback
|
||||
assert mock_fetch.call_count == 2
|
||||
# 第二次调用应该是 Wayback URL
|
||||
assert "web.archive.org/web/2/" in mock_fetch.call_args_list[1][0][0]
|
||||
|
||||
|
||||
def test_fetch_page_anti_bot_wayback_also_blocked():
|
||||
"""主抓取反爬 + Wayback 也反爬/失败 → 最终返回 error。
|
||||
|
||||
Wayback 兜底返回 None(失败)时,保留主抓取的反爬检测结果。
|
||||
"""
|
||||
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
|
||||
cf_result = _make_result(content=cf_html, final_url="https://example.com")
|
||||
with patch("search.fetch_url",
|
||||
side_effect=[cf_result, RuntimeError("wayback timeout")]):
|
||||
result = fetch_page("https://example.com", fallback_enabled=True)
|
||||
# Wayback 失败,保留反爬 error
|
||||
assert result["status"] == "error"
|
||||
assert result["anti_bot_detected"] is True
|
||||
# 但不会触发 Wayback(因为主抓取"成功"了,只是内容是反爬页)
|
||||
assert result["waf_type"] == "cloudflare"
|
||||
assert result["fallback_used"] is None
|
||||
|
||||
|
||||
def test_fetch_page_anti_bot_fallback_disabled():
|
||||
"""--no-fallback 时反爬页直接返回 error,不尝试 Wayback。"""
|
||||
cf_html = "<html><body>Just a moment... cf-ray: 123</body></html>"
|
||||
with patch("search.fetch_url",
|
||||
return_value=_make_result(content=cf_html)) as mock_fetch:
|
||||
result = fetch_page("https://example.com", fallback_enabled=False)
|
||||
assert result["status"] == "error"
|
||||
assert result["anti_bot_detected"] is True
|
||||
assert result["waf_type"] == "cloudflare"
|
||||
assert result["fallback_used"] is None
|
||||
# 只调用一次(主抓取),没有 Wayback
|
||||
assert mock_fetch.call_count == 1
|
||||
|
||||
|
||||
def test_fetch_page_passes_referer():
|
||||
"""referer 透传给 fetch_url。"""
|
||||
with patch("search.fetch_url", return_value=_make_result()) as mock_fetch:
|
||||
|
||||
Reference in New Issue
Block a user