fix(v2.0.1): 修复反爬误判 + Wayback 兜底未触发两个 bug

Bug 1: Wayback 兜底未触发 (search.py fetch_page)

- 根因: Cloudflare JS 质询页常返回 HTTP 200 (非 403), 主抓取 result 非 None

- _should_try_fallback 在 result 非 None 时直接返回 False, 跳过兜底

- 反爬检测在兜底判断之后执行, 错过兜底入口

- 修复: 反爬检测提前到兜底判断之前, anti_bot_detected=True 也触发 Wayback

- Wayback 结果重新做反爬检测 (防御性)

Bug 2: WAF 指纹库误判正常内容 (search.py WAF_FINGERPRINTS)

- 根因: 裸公司名 (cloudflare/akamai) 和宽泛词 (captcha/challenge/dd-) 做全文匹配

- DataCamp 文章引用 cloudflare.com 文档链接 -> 误判为 cloudflare WAF

- 'coding challenges' 正常内容 -> 误判为 generic 反爬

- Wayback 归档正文被误判, 兜底返回的有效内容被丢弃

- 修复: 移除裸公司名和宽泛词, 改用技术标识符 (cf-ray/incap_ses/bm_sz 等)

- 通用文案用完整短语 (please complete the captcha) 替代单词

- 增加 <title> 标签精准检测 (反爬页 title 是特征文案, 误判率极低)

- 新增 Anubis 反爬系统检测 (anubis_challenge/miserere)

真实测试验证 (search.metona.cn 实例):

- v2.0.0: --fetch 3 全部失败 (3 ERR: cloudflare/generic, Wayback 未触发)

- v2.0.1: --fetch 3 全部成功 (3 OK: 38100/93442/1945 chars, UA 轮换绕过 Cloudflare)

测试: 458 个全部通过 (新增 7 个测试覆盖修复行为)
This commit is contained in:
2026-08-01 21:44:58 +08:00
parent 28ff7c0a48
commit d9a08716bc
7 changed files with 215 additions and 66 deletions
+115 -45
View File
@@ -13,6 +13,7 @@ import json
import logging
import os
import random
import re
import sys
import threading
import time
@@ -855,9 +856,35 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
except Exception as e:
error_msg = str(e) if str(e) else e.__class__.__name__
# Wayback 兜底:主抓取失败或被反爬拦截时尝试
# 反爬检测(v2.0.0 增强:全文档扫描 + WAF 指纹库)
# 必须在 Wayback 兜底判断之前执行:Cloudflare 质询页常返回 HTTP 200
# 此时 result 非 None 但内容是反爬页,必须识别出来才能触发兜底。
anti_bot_detected = False
waf_type = None
if result is not None:
content = result.content
content_type = result.content_type or ""
is_html = ("html" in content_type.lower() or
content.strip().startswith("<!") or
content.strip().startswith("<htm"))
if is_html:
waf_type = _detect_anti_bot(content)
if waf_type:
anti_bot_detected = True
# Wayback 兜底触发条件(v2.0.1 修复):
# 1. 主抓取抛异常且错误暗示 404/403/超时(_should_try_fallback
# 2. 主抓取"成功"但被反爬拦截(anti_bot_detected=True
# 原 v2.0.0 bug:仅条件 1 触发兜底,条件 2 漏网(Cloudflare 200 质询页)。
fallback_used = None
if fallback_enabled and _should_try_fallback(result, error_msg):
need_fallback = False
if fallback_enabled:
if anti_bot_detected:
need_fallback = True
elif result is None and _should_try_fallback(result, error_msg):
need_fallback = True
if need_fallback:
wb_result = _try_wayback_fallback(url, timeout=timeout,
auth_headers=auth_headers,
max_retries=max_retries,
@@ -866,6 +893,18 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
result = wb_result
error_msg = None
fallback_used = "wayback"
# Wayback 结果重新做反爬检测(防御性:Wayback 快照极少是反爬页)
anti_bot_detected = False
waf_type = None
content = result.content
content_type = result.content_type or ""
is_html = ("html" in content_type.lower() or
content.strip().startswith("<!") or
content.strip().startswith("<htm"))
if is_html:
waf_type = _detect_anti_bot(content)
if waf_type:
anti_bot_detected = True
# 仍然失败
if result is None:
@@ -877,25 +916,10 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
"fallback_used": None,
}
content = result.content
content_type = result.content_type
final_url = result.final_url
is_html = ("html" in content_type.lower() or
content.strip().startswith("<!") or
content.strip().startswith("<htm"))
# 反爬检测(v2.0.0 增强:全文档扫描 + WAF 指纹库)
anti_bot_detected = False
waf_type = None
if is_html:
waf_type = _detect_anti_bot(content)
if waf_type:
anti_bot_detected = True
# 反爬仍被检测到(Wayback 也无能为力或兜底被禁用)
if anti_bot_detected:
return {
"url": url, "final_url": final_url, "status": "error",
"url": url, "final_url": result.final_url, "status": "error",
"error": f"Bot protection detected ({waf_type})",
"text": "", "text_length": 0, "truncated": False,
"anti_bot_detected": True, "waf_type": waf_type,
@@ -905,7 +929,7 @@ def fetch_page(url: str, timeout: int = 10, auth_headers: dict = None,
text = extract_text(content) if is_html else content
return {
"url": url,
"final_url": final_url,
"final_url": result.final_url,
"status": "ok",
"content_type": content_type,
"text": text,
@@ -971,59 +995,105 @@ def _try_wayback_fallback(url: str, timeout: int = 10,
return None
# ----- 反爬检测(v2.0.0 增强版-----
# ----- 反爬检测(v2.0.1 收窄误判 + title 精准检测-----
# WAF 指纹库:每项 = (waf_type, [指示词])
# 指示词在页面 HTML/body/headers 中出现即判定为该 WAF。
# 顺序按检测优先级:专用指纹在前,通用指纹在后。
#
# v2.0.1 修复:v2.0.0 用裸公司名(cloudflare/akamai)和宽泛词(captcha/
# challenge/dd-)做全文匹配,导致正常文章(如引用 Cloudflare 文档、讨论
# "coding challenges" 的文章、Wayback 归档正文)被误判为反爬页,Wayback
# 兜底返回的有效内容也被丢弃。收窄原则:
# 1. 专用指纹只用 WAF 厂商的技术标识符(cookie 名/HTTP header 名/JS 变量名)
# ——这些不会出现在文章正文里
# 2. 通用文案用完整短语而非单词(如 "please complete the captcha" 而非
# "captcha"),避免正常内容误判
# 3. 增加 <title> 标签检测——反爬页 title 是特征文案,最精准
WAF_FINGERPRINTS = [
("cloudflare", [
# Cloudflare 技术标识符(cookie/header/JS 变量名,不会出现在正文)
"cf-ray", "cf-chl-bypass", "cf-mitigated",
"cloudflare", "cf-browser-verification",
"attention required! | cloudflare", "just a moment",
"checking your browser before accessing",
"cf-browser-verification", "cf-error-details", "cf-error-code",
# 质询页特征文案(足够具体,正常内容不会完整出现)
"just a moment", "checking your browser before accessing",
"attention required! | cloudflare",
"enable javascript and cookies to continue",
]),
("imperva", [
# Incapsula cookie/技术标识符
"incap_ses", "visid_incap", "incap_ses_",
"imperva", "incapsula",
"request unsuccessful. incapsula incident id",
"incapsula incident id", "request unsuccessful. incapsula",
"visit denied by incapsula",
]),
("perimeterx", [
"_px", "px-captcha", "pxhd", "pxcts", "pxcookie",
"perimeterx", "press & hold to confirm you are a human",
# PerimeterX 专有标识符(_px 太短会匹配 CSS 类名,已移除)
"px-captcha", "pxhd", "pxcts", "pxcookie",
"_pxff", "_pxhd",
"press & hold to confirm you are a human",
]),
("datadome", [
"datadome", "dd-", "data-dome",
"protected by datadome",
# DataDome 专有标识符(dd- 太宽泛会匹配 dd-class 等,已移除)
"datadome", "data-dome",
"protected by datadome", "datadome-bot-protect",
]),
("akamai", [
"akamai", "bm_sz", "_abck",
"reference #", "akamaighost",
# Akamai Bot Manager cookie/标识符(akamai 裸名会匹配正文引用,已移除)
"bm_sz", "_abck", "akamaighost", "akamai-bot-manager",
"ak_bmsc",
]),
# 通用反爬指示词(无明确 WAF 归属
# 通用反爬指示词(v2.0.1 收窄:完整短语而非单词,避免正文误判
("generic", [
"captcha", "challenge", "verify you are human",
"making sure you're not a bot",
"please enable javascript", "enable javascript to continue",
"ddos protection", "access denied",
"verify you are human", "verify that you are human",
"making sure you're not a bot", "are you a robot",
"robot or human", "human verification",
"please complete the captcha", "complete the security check",
"please enable javascript to continue",
"enable javascript to continue",
"ddos protection by", "access denied - sucuri",
"you have been blocked", "unusual traffic from your computer",
"robot or human", "are you a robot",
"pardon our interruption", "we'll be right back",
"bot protection", "anti-bot protection",
"anubis_challenge", "miserere", # Anubis 反爬系统(拦截 AI 爬虫)
]),
]
# <title> 标签检测:反爬页 title 通常是特征文案,比全文扫描更精准。
# key = title 中的特征子串(小写),value = 对应 WAF 类型
_TITLE_ANTI_BOT_SIGNATURES = {
"just a moment": "cloudflare",
"attention required": "cloudflare",
"access denied": "generic",
"are you a robot": "generic",
"robot check": "generic",
"human verification": "generic",
"please verify you are human": "generic",
"security check": "generic",
"verify you are human": "generic",
}
_TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.IGNORECASE | re.DOTALL)
def _detect_anti_bot(content: str) -> str:
"""检测反爬页面,返回 WAF 类型或 None。
v2.0.0 改进:
* 全文档扫描(去除 2000 字符限制——大页面反爬页可能在前 2000 字之外
* WAF 指纹库覆盖 Cloudflare/Imperva/PerimeterX/DataDome/Akamai/通用
* 返回具体 WAF 类型而非布尔值,让 AI Agent 可决策
v2.0.1 改进:
* 增加 <title> 标签精准检测(反爬页 title 是特征文案,误判率极低
* 收窄 WAF 指纹库关键词(移除裸公司名和宽泛词,改用技术标识符 + 完整短语)
* 保留全文档扫描(v2.0.0 改进,大页面反爬页可能在前 2000 字之外)
性能:全文档 lower() 一次,对 5MB 页面约 5ms,可接受
检测顺序:title 标签(最精准)→ 全文档指纹扫描(技术标识符 + 文案短语)
性能:全文档 lower() 一次 + title regex 一次,对 5MB 页面约 5ms,可接受。
"""
if not content:
return None
# 1. <title> 标签检测(最精准,误判率极低)
title_match = _TITLE_RE.search(content)
if title_match:
title_lower = title_match.group(1).strip().lower()
for sig, waf_type in _TITLE_ANTI_BOT_SIGNATURES.items():
if sig in title_lower:
return waf_type
# 2. 全文档指纹扫描(收窄后的技术标识符 + 完整文案)
lower = content.lower()
for waf_type, indicators in WAF_FINGERPRINTS:
for ind in indicators: