
Last30days Cn
- 753 installs
- 1.3k repo stars
- Updated July 20, 2026
- jesseovo/last30days-skill-cn
Researches the last 30 days of activity across Chinese platforms like Weibo, Xiaohongshu, Bilibili, Zhihu and Douyin.
About
Gathers and summarizes recent sentiment and topics from major Chinese social and search platforms, with Markdown, JSON and styled HTML report output. A developer or marketer uses it when researching Chinese-market trends on a topic.
- Covers Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, WeChat, Baidu and Toutiao
- Outputs Markdown, JSON, compact context and Swiss/IKB HTML reports
Last30days Cn by the numbers
- 753 all-time installs (skills.sh)
- +91 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #625 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jesseovo/last30days-skill-cn --skill last30days-cnAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 753 |
|---|---|
| repo stars | ★ 1.3k |
| Last updated | July 20, 2026 |
| Repository | jesseovo/last30days-skill-cn ↗ |
What it does
Researches the last 30 days of activity across Chinese platforms like Weibo, Xiaohongshu, Bilibili, Zhihu and Douyin.
Files
last30days-cn
You are a Chinese-platform research assistant. Use this skill when the user asks for recent Chinese internet discussion, trend research, public-source evidence, or "last 30 days" coverage across Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, WeChat public accounts, Baidu, and Toutiao.
Core Rule
Always ground claims in returned results. Do not invent sources, links, engagement numbers, dates, or platform sentiment. If coverage is sparse, say so clearly.
Run
Use the skill-local scripts directory:
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --emit compactUseful variants:
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --quick --emit compact
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --deep --emit md
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --emit html-path
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --search weibo,bilibili,zhihu --emit compact
python {{SKILL_DIR}}/scripts/last30days.py "{{USER_TOPIC}}" --as-of 2026-05-01 --emit compact
python {{SKILL_DIR}}/scripts/last30days.py --diagnose--as-of YYYY-MM-DD 以指定日期为终点回溯 N 天(历史回溯);未指定 --search 时回退到环境变量 LAST30DAYS_DEFAULT_SEARCH,EXCLUDE_SOURCES 可排除指定源。输出中若多个平台讨论同一事件,会先给出「跨平台聚合热点」。
Output Modes
compact: concise Markdown evidence for the agent to synthesize.md: full Markdown report.html: complete standalone HTML report.html-path: path to the generatedreport.html.json: structured report data.context: reusable context snippet.path: path tolast30days.context.md.
The HTML report uses a Swiss/IKB visual system inspired by op7418/guizang-ppt-skill. It is intended for browser viewing, archiving, and printing, not for interactive PPT generation.
Configuration
Most sources can be tried with no configuration. Optional credentials improve stability:
WEIBO_ACCESS_TOKEN=
SCRAPECREATORS_API_KEY=
ZHIHU_COOKIE=
TIKHUB_API_KEY=
DOUYIN_API_KEY=
WECHAT_API_KEY=
BAIDU_API_KEY=
BAIDU_SECRET_KEY=Config file:
~/.config/last30days-cn/.envOptional crawler mode:
python -m pip install playwright
python -m playwright install chromiumSynthesis Guidance
When presenting the final answer:
1. State the date range and the active sources. 2. Separate confirmed findings from weak or sparse signals. 3. Cite platform and URL for important claims. 4. Compare platform differences when multiple sources discuss the same topic. 5. Mention unavailable or failed sources if that affects confidence. 6. Keep the final answer in Chinese unless the user requests otherwise.
Compliance
This skill is for learning, research, and personal knowledge work. Use low frequency, respect platform terms and robots.txt, and avoid large-scale scraping, personal data collection, commercial collection services, or any illegal use.
{
"$schema": "https://anthropic.com/claude-code/marketplace.schema.json",
"name": "last30days-skill-cn",
"description": "中国平台深度研究引擎 - 搜索微博、小红书、B站、知乎等8大平台最近30天的热门内容,生成有据可查的研究简报。",
"owner": {
"name": "Jesseovo",
"url": "https://github.com/Jesseovo"
},
"plugins": [
{
"name": "last30days-cn",
"description": "中国互联网平台深度研究技能,搜索微博、小红书、B站、知乎、抖音、微信公众号、百度、今日头条等平台近30天热门内容,生成有据可查的研究报告。",
"version": "2.1.0",
"author": {
"name": "Jesse",
"url": "https://github.com/Jesseovo"
},
"source": "./",
"category": "productivity",
"homepage": "https://github.com/Jesseovo/last30days-skill-cn"
}
]
}
{
"name": "last30days-cn",
"version": "2.1.0",
"description": "中国平台深度研究引擎 - 覆盖微博、小红书、B站、知乎、抖音、微信公众号、百度搜索、今日头条等8大平台。v2.1 修复反爬问题,XHR拦截+Bing兜底。",
"author": {
"name": "Jesse",
"url": "https://github.com/Jesseovo"
},
"license": "MIT",
"keywords": ["research", "weibo", "xiaohongshu", "bilibili", "zhihu", "douyin", "wechat", "baidu", "toutiao", "chinese-platforms"],
"skills": ["./"],
"hooks": {}
}
# 排除开发/测试文件,仅保留 ClawHub 发布所需
fixtures/
tests/
agents/
release-notes.md
SPEC.md
TASKS.md
*.jsonl
*.mp3
*.jpeg
*.jpg
*.png
*.gif
# OS / tool files
.DS_Store
.claude/
.entire/
__pycache__/
*.pyc
.pytest_cache/
mise.toml
*.egg-info/
dist/
build/
.env
interface:
display_name: "近30天研究"
short_description: "搜索微博、小红书、B站、知乎、抖音、微信公众号、百度、今日头条等8大中国平台最近30天的热门内容,生成综合研究报告。"
default_prompt: "研究这个话题在中国互联网平台上最近30天的讨论。综合各平台观点,分析趋势,找出核心发现。"
brand_color: "#E53935"
policy:
allow_implicit_invocation: true
last30days-cn
Agent-facing notes for local development.
Purpose
last30days-cn researches recent Chinese-platform discussion across Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, WeChat public accounts, Baidu, and Toutiao. The default window is the last 30 days.
Structure
skills/last30days/SKILL.md: installable Agent Skill entrypoint.skills/last30days/scripts/: self-contained runtime payload for Agent Skills installs.scripts/: root development copy kept for compatibility.scripts/lib/render.py: Markdown, JSON context, and HTML report rendering.tests/: focused regression tests.
Commands
python scripts/last30days.py "你的主题" --emit compact
python scripts/last30days.py "你的主题" --emit html-path
python scripts/last30days.py --diagnoseWhen validating on this Windows workspace, prefer:
py -m pytestThe python command may resolve to the Windows Store shim on some machines.
Release Notes
For v3.0.0, the skill payload is self-contained under skills/last30days, and the new HTML renderer emits a Guizang-inspired Swiss/IKB report.
{
"data": {
"result": [
{
"type": "video",
"bvid": "BV1xx411c7XY",
"title": "2026年最值得学的AI工具盘点",
"description": "本期视频盘点了2026年最火的AI工具...",
"author": "科技视频号",
"mid": 12345678,
"play": 150000,
"danmaku": 800,
"like": 5000,
"coin": 2000,
"favorites": 3000,
"review": 600,
"pubdate": 1710892800,
"duration": "15:30",
"arcurl": "https://www.bilibili.com/video/BV1xx411c7XY"
},
{
"type": "video",
"bvid": "BV1yy411c8ZW",
"title": "AI编程助手横评:谁是最强?",
"description": "测试了多款AI编程助手的实际表现...",
"author": "编程小课堂",
"mid": 23456789,
"play": 80000,
"danmaku": 400,
"like": 3000,
"coin": 1000,
"favorites": 2000,
"review": 350,
"pubdate": 1711324800,
"duration": "20:45",
"arcurl": "https://www.bilibili.com/video/BV1yy411c8ZW"
}
]
}
}
{
"statuses": [
{
"id": 5120001234567890,
"mid": "5120001234567890",
"text": "最新AI工具测试,效果惊人!#人工智能# #AI工具# 大家觉得怎么样?",
"user": {
"id": 1234567890,
"screen_name": "科技圈小编",
"followers_count": 50000
},
"created_at": "2026-03-15 14:30:00",
"reposts_count": 128,
"comments_count": 256,
"attitudes_count": 1024,
"source": "微博 weibo.com"
},
{
"id": 5120001234567891,
"mid": "5120001234567891",
"text": "深度体验了一周的AI编程助手,写一下使用心得和踩坑记录",
"user": {
"id": 1234567891,
"screen_name": "程序猿日记",
"followers_count": 30000
},
"created_at": "2026-03-20 10:15:00",
"reposts_count": 64,
"comments_count": 128,
"attitudes_count": 512,
"source": "微博 weibo.com"
}
],
"total_number": 2
}
{
"data": [
{
"type": "search_result",
"object": {
"type": "answer",
"id": 3456789012,
"question": {
"id": 678901234,
"title": "2026年有哪些值得关注的AI工具?"
},
"author": {
"name": "AI研究员",
"url_token": "ai-researcher",
"headline": "人工智能从业者"
},
"excerpt": "作为一个每天都在使用各种AI工具的从业者,我来分享一下2026年最值得关注的几个工具...",
"voteup_count": 2500,
"comment_count": 180,
"created_time": 1711065600,
"updated_time": 1711152000,
"url": "https://www.zhihu.com/question/678901234/answer/3456789012"
}
},
{
"type": "search_result",
"object": {
"type": "article",
"id": 789012345,
"title": "深入理解大语言模型的工作原理",
"author": {
"name": "机器学习专栏",
"url_token": "ml-column"
},
"excerpt": "本文从技术角度深入分析大语言模型的核心原理...",
"voteup_count": 1800,
"comment_count": 95,
"created": 1710979200,
"url": "https://zhuanlan.zhihu.com/p/789012345"
}
}
]
}
{
"name": "last30days-cn",
"version": "3.0.0-cn",
"description": "Chinese-platform last-30-days research skill covering Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, WeChat, Baidu, and Toutiao. Includes compact Markdown, full Markdown, JSON, and Guizang-inspired Swiss/IKB HTML report output.",
"settings": [
{
"name": "Extension Directory",
"description": "Gemini CLI extension install directory.",
"envVar": "GEMINI_EXTENSION_DIR",
"sensitive": false
},
{
"name": "Weibo Access Token",
"description": "Optional Weibo API access token.",
"envVar": "WEIBO_ACCESS_TOKEN",
"sensitive": true
},
{
"name": "ScrapeCreators API Key",
"description": "Optional API key kept for compatibility with existing configuration.",
"envVar": "SCRAPECREATORS_API_KEY",
"sensitive": true
},
{
"name": "Zhihu Cookie",
"description": "Optional Zhihu login cookie for higher-quality search.",
"envVar": "ZHIHU_COOKIE",
"sensitive": true
},
{
"name": "TikHub API Key",
"description": "Optional TikHub key for Douyin search.",
"envVar": "TIKHUB_API_KEY",
"sensitive": true
},
{
"name": "WeChat API Key",
"description": "Optional WeChat public-account search API key.",
"envVar": "WECHAT_API_KEY",
"sensitive": true
},
{
"name": "Baidu API Key",
"description": "Optional Baidu search API key.",
"envVar": "BAIDU_API_KEY",
"sensitive": true
},
{
"name": "Baidu Secret Key",
"description": "Optional Baidu secret key paired with BAIDU_API_KEY.",
"envVar": "BAIDU_SECRET_KEY",
"sensitive": true
}
]
}
{
"hooks": {
"SessionStart": [
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "bash ${CLAUDE_PLUGIN_ROOT}/hooks/scripts/check-config.sh",
"timeout": 5
}
]
}
]
}
}
#!/bin/bash
# Author: Jesse (https://github.com/Jesseovo)
set -euo pipefail
# 检查 last30days-cn 配置状态并显示欢迎/就绪信息(中国平台)。
# 优先级:.claude/last30days-cn.env > ~/.config/last30days-cn/.env > 环境变量
PROJECT_ENV=".claude/last30days-cn.env"
GLOBAL_ENV="$HOME/.config/last30days-cn/.env"
# 可选:配置文件权限过宽时警告
check_perms() {
local file="$1"
if [[ ! -f "$file" ]]; then return; fi
local perms
perms=$(stat -f '%Lp' "$file" 2>/dev/null || stat -c '%a' "$file" 2>/dev/null || echo "")
if [[ -n "$perms" && "$perms" != "600" && "$perms" != "400" ]]; then
echo "last30days-cn:警告 — $file 权限为 $perms(建议 600)。"
echo " 修复:chmod 600 $file"
fi
}
# 将 env 文件读入变量供检查(不 export)
load_env_vars() {
local file="$1"
if [[ -f "$file" ]]; then
while IFS='=' read -r key value; do
[[ "$key" =~ ^[[:space:]]*# ]] && continue
[[ -z "$key" ]] && continue
key=$(echo "$key" | xargs)
value=$(echo "$value" | xargs | sed 's/^["'\''"]//;s/["'\''"]$//')
if [[ -n "$key" && -n "$value" ]]; then
eval "ENV_${key}=\"${value}\""
fi
done < "$file"
fi
}
CONFIG_FILE=""
if [[ -f "$PROJECT_ENV" ]]; then
CONFIG_FILE="$PROJECT_ENV"
check_perms "$PROJECT_ENV"
elif [[ -f "$GLOBAL_ENV" ]]; then
CONFIG_FILE="$GLOBAL_ENV"
check_perms "$GLOBAL_ENV"
fi
if [[ -n "$CONFIG_FILE" ]]; then
load_env_vars "$CONFIG_FILE"
fi
SETUP_COMPLETE="${ENV_SETUP_COMPLETE:-${SETUP_COMPLETE:-}}"
# 中国平台可选密钥(文件或环境)
HAS_WEIBO="${ENV_WEIBO_ACCESS_TOKEN:-${WEIBO_ACCESS_TOKEN:-}}"
HAS_SCRAPE="${ENV_SCRAPECREATORS_API_KEY:-${SCRAPECREATORS_API_KEY:-}}"
HAS_ZHIHU_COOKIE="${ENV_ZHIHU_COOKIE:-${ZHIHU_COOKIE:-}}"
HAS_TIKHUB="${ENV_TIKHUB_API_KEY:-${TIKHUB_API_KEY:-}}"
HAS_WECHAT="${ENV_WECHAT_API_KEY:-${WECHAT_API_KEY:-}}"
HAS_BAIDU="${ENV_BAIDU_API_KEY:-${BAIDU_API_KEY:-}}"
any_cn_key_set() {
[[ -n "$HAS_WEIBO" || -n "$HAS_SCRAPE" || -n "$HAS_ZHIHU_COOKIE" || -n "$HAS_TIKHUB" || -n "$HAS_WECHAT" || -n "$HAS_BAIDU" ]]
}
# 从未配置且无密钥:中文欢迎
if [[ -z "$SETUP_COMPLETE" && -z "$CONFIG_FILE" ]] && ! any_cn_key_set; then
cat <<'EOF'
last30days-cn:已就绪。运行研究流程即可开始 — 配置可选密钥可解锁更多平台。
B站、知乎、百度(基础)、今日头条 可免费直接使用,无需 API Key。
配置 WEIBO_ACCESS_TOKEN、SCRAPECREATORS_API_KEY、ZHIHU_COOKIE、TIKHUB_API_KEY、WECHAT_API_KEY、BAIDU_API_KEY 可分别启用或增强 微博、小红书、知乎、抖音、微信公众号、百度搜索 等能力。
EOF
exit 0
fi
# 基础可用源:四大免费平台
SOURCE_COUNT=4
[[ -n "$HAS_WEIBO" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
[[ -n "$HAS_SCRAPE" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
[[ -n "$HAS_ZHIHU_COOKIE" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
[[ -n "$HAS_TIKHUB" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
[[ -n "$HAS_WECHAT" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
[[ -n "$HAS_BAIDU" ]] && SOURCE_COUNT=$((SOURCE_COUNT + 1))
echo "last30days-cn:就绪 — 当前约 ${SOURCE_COUNT} 路数据源可用(含 B站/知乎/百度基础/头条 等免费源)。"
if ! any_cn_key_set; then
echo " 提示:在 ~/.config/last30days-cn/.env 或 .claude/last30days-cn.env 中配置可选密钥,可启用微博、小红书、抖音等更多平台。"
fi
MIT License
Copyright (c) 2026 Matt Van Horn (Original last30days-skill)
Copyright (c) 2026 Jesse (Chinese localization fork: last30days-skill-cn)
This project is a Chinese-localized fork of https://github.com/mvanhorn/last30days-skill
Original author: Matt Van Horn (mvanhorn@gmail.com)
Fork author: Jesse (https://github.com/Jesseovo)
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
<p align="center"> <img src="assets/banner.jpg" alt="last30days-cn — Chinese Platform Deep Research Engine" width="380"> </p>
<p align="center"> <a href="README.md">简体中文</a> · <b>English</b> </p>
📰 last30days-cn — Chinese Platform Deep Research Engine
🚀 30 days of research, 30 seconds of work. 8 platforms. Zero stale info.
last30days-cn is an AI Agent skill that automatically searches the 8 major Chinese internet platforms for the last 30 days of content and generates well-cited research reports.
🔗 This project is a deeply localized fork of mvanhorn/last30days-skill, fully adapted for Chinese users and Chinese internet platforms.
🕷️ v2.0 integrates the MediaCrawler crawler-engine approach to greatly reduce API-key dependency. v2.1 fixed Baidu/Xiaohongshu anti-bot issues (XHR interception over DOM parsing, Bing search fallback) and removed the ineffective ScrapeCreators Xiaohongshu integration.
Current version: v3.0.0
👤 Author: Jesse (@Jesseovo)
---
✨ What's new in v3.0.0
- Matches upstream v3's Agent Skills package layout:
skills/last30daysis now an independently installable runtime payload. - The Chinese CLI uses a single entry point,
last30days.py, with the same structure in both the repo root and the skill payload. - Added
--emit htmland--emit html-pathto generate an offline-openablereport.html. - The HTML report adopts the Swiss/IKB visual language from op7418/guizang-ppt-skill — good for browsing, archiving, and printing.
- Xiaohongshu and Zhihu search now report empty-result fallbacks, listing which paths were tried and likely causes on failure.
- Douyin and Toutiao now fall back to a public search engine when their native APIs are rate-limited, so they no longer silently return 0 results (see issue #8).
- Fixed
npxinstall failure on macOS/Linux caused by a corrupted symlink atskills/last30days/SKILL.md(see issue #10). - Synced several platform-agnostic capabilities from upstream mvanhorn/last30days-skill:
--as-ofhistorical lookback, cross-source cluster merging,LAST30DAYS_DEFAULT_SEARCH/EXCLUDE_SOURCESconfig knobs, honest--diagnose(live probes), and HTML report XSS hardening. - The repo-root
scripts/is kept for local development and legacy paths; Agent Skills installs use the self-contained payload underskills/last30days/scripts.
✅ Quality checks
Done before this release:
- The published remote tag is unified as
v3.0.0, with no extra v3-derived tags. - Both the repo root and the skill payload use the single
last30days.pyentry point — no extra entry files. - Full test suite passes:
py -m pytest,176 passed. - Both entry points verified:
py scripts/last30days.py --diagnoseandpy skills/last30days/scripts/last30days.py --diagnoseboth print platform availability diagnostics.
---
⚠️ Disclaimer
Please read carefully. By using this project you agree to all of the terms below.
1. This project is for learning and research only. All crawler features are intended solely for technical study and exchange; commercial use is strictly prohibited. 2. Users must comply with all applicable laws and regulations (data security, personal-information protection, anti-unfair-competition, etc.). 3. Users must respect each platform's Terms of Service (ToS) and robots.txt. 4. Do NOT use this project for: large-scale / high-frequency scraping; collecting, storing, or disseminating others' personal data; disrupting platform operations; reselling data or any commercial gain; or providing automated data-collection services to third parties. 5. The developer assumes no liability for any legal consequences arising from use of this project. Users bear all legal risk themselves. 6. For any infringement concerns, contact the author and it will be addressed promptly.
Technical notes
- The crawler features rely on Playwright browser automation to mimic normal browsing; they do not reverse-engineer encryption or break security mechanisms.
- Platform interfaces can change at any time; this project does not guarantee that every feature always works.
- Keep request frequency reasonable (e.g. ≥ 5s between searches) to avoid being blocked.
💡 Illegal scraping cases are common — use this lawfully and compliantly.
Reference: Chinese crawler legal cases
---
📋 Platform support
| Platform | Module | Data sources | Configuration |
|---|---|---|---|
weibo.py | API / 🕷️ crawler / public API | ✅ No config in crawler mode | |
| 📕 Xiaohongshu | xiaohongshu.py | API / 🕷️ crawler / public API / search fallback | ✅ No config in crawler mode |
| 📺 Bilibili | bilibili.py | Public API / 🕷️ crawler backup | ✅ No config |
| 💬 Zhihu | zhihu.py | Public search / 🕷️ crawler backup / search fallback | ✅ No config |
| 🎵 Douyin | douyin.py | API / 🕷️ crawler / public API / search fallback | ✅ No config in crawler mode |
wechat.py | API / Sogou search | WECHAT_API_KEY (optional) | |
| 🔵 Baidu | baidu.py | Public search / API | ✅ No config for basic search |
| 📰 Toutiao | toutiao.py | Public API / search fallback | ✅ No config |
🕷️ = requires Playwright (pip install playwright && playwright install chromium)---
🤖 Install
| Surface | Install |
|---|---|
| Agent Skills (recommended) | npx skills add Jesseovo/last30days-skill-cn -g |
| Claude Code | npx skills add Jesseovo/last30days-skill-cn -g |
| Cursor | Clone the repo and add SKILL.md as a project skill |
| OpenClaw / ClawHub | git clone https://github.com/Jesseovo/last30days-skill-cn.git ~/.agents/skills/last30days-cn |
| Gemini CLI | Clone and load as a Gemini extension |
| Any agent | Any AI agent with Bash / Read / Write tools |
Agent Skills (recommended)
npx skills add Jesseovo/last30days-skill-cn -gManual (developer)
git clone https://github.com/Jesseovo/last30days-skill-cn.git ~/.claude/skills/last30days-cn---
⚙️ Configuration
Step 1 — Install dependencies
pip install jiebaStep 2 — Install the crawler engine (recommended; enables 7/8 platforms without API keys)
pip install playwright
playwright install chromiumWith Playwright installed, Weibo, Xiaohongshu, Douyin, Bilibili (backup), and Zhihu (backup) all work without API keys.
Step 3 — Create a config file (optional)
For more stable API-mode data, or to use WeChat public-account search:
mkdir -p ~/.config/last30days-cn
touch ~/.config/last30days-cn/.env
chmod 600 ~/.config/last30days-cn/.envWindows (PowerShell): create the config directory with New-Item -ItemType Directory -Force -Path "$env:USERPROFILE\.config\last30days-cn", then create the file with New-Item -ItemType File -Path "$env:USERPROFILE\.config\last30days-cn\.env" -Force (or edit %USERPROFILE%\.config\last30days-cn\.env). To restrict .env to the current user (similar intent to chmod 600; the ACL model differs), run icacls "$env:USERPROFILE\.config\last30days-cn\.env" /inheritance:r /grant:r "$($env:USERNAME):(R,W)".
Edit ~/.config/last30days-cn/.env and fill in API keys as needed (all optional):
# All API keys are OPTIONAL. With Playwright installed, most platforms
# already work via crawler mode; API keys just give more stable data.
WEIBO_ACCESS_TOKEN= # Weibo Open Platform (optional; crawler covers it)
SCRAPECREATORS_API_KEY= # Xiaohongshu (optional; crawler covers it)
ZHIHU_COOKIE= # Zhihu cookie (optional; improves search quality)
TIKHUB_API_KEY= # Douyin (optional; crawler covers it)
WECHAT_API_KEY= # WeChat public accounts (no crawler alternative; required for WeChat)
BAIDU_API_KEY= # Baidu Search API (optional; public search already works)
BAIDU_SECRET_KEY=Step 4 — Verify
python scripts/last30days.py --diagnosePrints each platform's availability plus crawler-engine status:
{
"weibo": true,
"xiaohongshu": false,
"bilibili": true,
"zhihu": true,
"douyin": true,
"wechat": false,
"baidu_api": false,
"toutiao": true,
"crawler_engine": {
"playwright_available": true,
"cached_logins": [],
"note": "With Playwright installed, Weibo/Xiaohongshu/Douyin/Bilibili/Zhihu work without API keys"
},
"note_douyin_toutiao": "Douyin/Toutiao native APIs require signature params and are often rate-limited; on failure they fall back to a public search engine, which only yields public links (no real engagement data or precise dates)."
}---
🚀 Usage
Basic
python scripts/last30days.py "AI coding assistant" --emit compact
python scripts/last30days.py "AI coding assistant" --emit html-pathCLI flags
| Flag | Description | Example |
|---|---|---|
--emit | Output mode | compact / json / md / context / path / html / html-path |
--quick | Quick search | Fewer sources, faster |
--deep | Deep search | More sources, more thorough |
--days N | Look-back days | --days 7 (last week) |
--as-of | Historical end date | --as-of 2026-05-01 (look back N days from that date) |
--search | Specific sources | --search weibo,bilibili,zhihu |
--diagnose | Diagnose config | Show per-platform availability |
--timeout SECS | Global timeout | Override the default global timeout |
--save-dir DIR | Auto-save raw output | Write raw output to the given dir |
--debug | Debug mode | Verbose logging |
🔧 Env vars: when--searchis omitted, it falls back toLAST30DAYS_DEFAULT_SEARCH(a comma-separated default source set);EXCLUDE_SOURCESremoves sources from the active set.
Examples
# 🔍 Search an AI topic
python scripts/last30days.py "latest AI tools" --emit compact
# ⚡ Quick search, Bilibili + Zhihu only
python scripts/last30days.py "Python tutorial" --quick --search bilibili,zhihu
# 📊 Deep search and save results
python scripts/last30days.py "EV cars" --deep --save-dir ~/Documents/research
# 📋 JSON output (for programmatic use)
python scripts/last30days.py "ChatGPT alternatives" --emit json
# 🗓️ Last 7 days only
python scripts/last30days.py "trending topics" --days 7
# 🧾 Generate an offline-openable HTML report
python scripts/last30days.py "embodied AI" --deep --emit html-path---
🔧 Data-fetching strategy (tiered fallback)
Tier 1: API mode (if an API key is configured)
↓ failed or not configured
Tier 2: Crawler mode (MediaCrawler, requires Playwright)
↓ failed or not installed
Tier 3: Public API (direct HTTP, no config needed)
↓ still no results (Douyin / Toutiao / Xiaohongshu / Zhihu)
Fallback: Public search engine (Bing `site:` search, yields public links)⚠️ Douyin/Toutiao native web APIs now require signature params (a_bogus/_signature) and are often rate-limited. On failure they fall back to a public search engine, which only yields public links — no real engagement data or precise dates.
---
🏗️ Project layout
last30days-skill-cn/
├── SKILL.md # Agent skill definition
├── README.md # Docs (Chinese)
├── README.en.md # Docs (English, this file)
├── LICENSE # MIT license
├── requirements.txt # Python deps
├── assets/ # README images
├── scripts/ # Dev copy: last30days.py + lib/ (per-platform modules)
├── skills/last30days/ # Self-contained Agent Skills payload (SKILL.md + scripts/)
├── fixtures/ # Sample data
├── tests/ # Test cases
└── hooks/ # Agent hooks---
📊 Scoring
Each result gets a 0–100 composite score:
| Dimension | Weight | Notes |
|---|---|---|
| 🎯 Relevance | 45% | Text match against the query |
| 🕐 Recency | 25% | Freshness of the publish time |
| 🔥 Engagement | 30% | Per-platform interaction metrics |
---
🙏 Acknowledgements
- mvanhorn/last30days-skill — the original English project
- NanmiCoder/MediaCrawler — crawler-engine inspiration
---
📜 License
Released under the MIT License.
- 🔗 Original: mvanhorn/last30days-skill by Matt Van Horn
- 🇨🇳 Chinese fork: Jesse (@Jesseovo)
<p align="center"> <img src="assets/banner.jpg" alt="last30days-cn — 中国平台深度研究引擎" width="380"> </p>
<p align="center"> <b>简体中文</b> · <a href="README.en.md">English</a> </p>
📰 last30days-cn — 中国平台深度研究引擎
🚀 30 天的研究,30 秒的结果。8 大平台。零过时信息。
last30days-cn 是一个 AI Agent 技能(Skill),能够自动搜索中国互联网 8 大主流平台最近 30 天的内容,综合分析后生成有据可查的研究报告。
🔗 本项目基于 mvanhorn/last30days-skill 进行深度本土化改造,完全面向中国用户和中文互联网平台。
🕷️ v2.0 集成 MediaCrawler 爬虫引擎思路,大幅减少 API Key 依赖。v2.1 修复百度/小红书反爬问题,XHR 拦截替代 DOM 解析,Bing 兜底搜索,已移除无效的 ScrapeCreators 小红书集成。
当前版本:v3.0.0
👤 作者 / Author: Jesse (@Jesseovo)
---
✨ v3.0.0 升级内容
- 追平原版 v3 的 Agent Skills 包结构:
skills/last30days现在是可独立安装的运行载荷。 - 中文 CLI 统一使用单入口
last30days.py,根目录和 Skill 载荷保持同名结构。 - 新增
--emit html和--emit html-path,可生成离线可打开的report.html。 - HTML 报告融入 op7418/guizang-ppt-skill 的 Swiss/IKB 视觉语言,适合浏览、归档、打印。
- 小红书和知乎搜索增加空结果兜底说明,失败时会标注已尝试路径与可能原因。
- 抖音、头条在原生接口被风控时新增公开搜索引擎兜底,不再静默返回 0 条(见 issue #8)。
- 修复 macOS/Linux 下
skills/last30days/SKILL.md损坏 symlink 导致的npx安装失败(见 issue #10)。 - 对照上游 mvanhorn/last30days-skill 同步若干平台无关能力:
--as-of历史回溯、跨平台聚合热点、LAST30DAYS_DEFAULT_SEARCH/EXCLUDE_SOURCES配置开关、诚实--diagnose(实时探测)、HTML 报告 XSS 加固。 - 根目录
scripts/继续保留,方便本地开发和旧路径调用;Agent Skills 安装使用skills/last30days/scripts下的自包含载荷。
✅ 质量验证
本次发布前已完成以下质检:
- 远端发布 tag 统一为
v3.0.0,没有额外 v3 派生 tag。 - 根目录与 Skill 载荷均统一使用
last30days.py单入口,没有额外入口文件。 - 全量测试通过:
py -m pytest,共176 passed。 - 根目录入口和 Skill 载荷入口均已验证:
py scripts/last30days.py --diagnose与py skills/last30days/scripts/last30days.py --diagnose均可正常输出平台可用性诊断。
---
⚠️ 免责声明 / Disclaimer
请务必仔细阅读以下内容。使用本项目即表示您同意以下所有条款。
法律合规声明
1. 本项目仅供学习和研究目的。所有爬虫功能仅用于技术学习与研究交流,严禁用于商业用途。 2. 使用者必须严格遵守中华人民共和国相关法律法规,包括但不限于:
- 《中华人民共和国网络安全法》
- 《中华人民共和国数据安全法》
- 《中华人民共和国个人信息保护法》
- 《中华人民共和国反不正当竞争法》
3. 使用者必须遵守各平台的服务条款(ToS)和 robots.txt 规定。 4. 禁止将本项目用于以下行为:
- 大规模、高频率地抓取平台数据
- 收集、存储或传播他人个人隐私信息
- 破坏或干扰平台正常运营
- 任何形式的非法数据倒卖或商业牟利
- 对外提供自动化数据采集服务
5. 本项目开发者不承担因使用本项目而产生的任何法律责任。用户应自行承担使用本项目的全部法律风险。 6. 如有侵权,请联系作者,将在第一时间处理。
技术免责
- 爬虫功能依赖 Playwright 浏览器自动化,模拟正常用户浏览行为,不涉及逆向加密算法或破解安全机制。
- 各平台接口随时可能变更,本项目不保证所有功能始终可用。
- 建议将请求频率控制在合理范围内(如每次搜索间隔 ≥ 5 秒),避免被平台封禁。
💡 爬虫违法违规的案例频发,请务必合法合规使用。
参考:中国爬虫相关法律案例汇总
---
✨ v2.0 新特性
🆕 v2.0 vs v1.0 对比
| 特性 | v1.0 | v2.0 |
|---|---|---|
| 免费可用平台数 | 4 个 | 7 个(安装 Playwright 后) |
| 需要 API Key 的平台 | 微博、小红书、抖音、微信 | 仅微信(其余可用爬虫替代) |
| 数据获取方式 | 仅 API + 公开接口 | API + 爬虫引擎 + 公开接口 |
| 安装难度 | 需配置多个 API Key | pip install playwright 即可 |
| marketplace.json | 缺少 owner 字段(Bug) | ✅ 已修复 |
核心升级
1. 集成 MediaCrawler 爬虫引擎 — 基于 Playwright 浏览器自动化,无需逆向加密算法,大幅降低使用门槛 2. 7/8 平台零配置可用 — 除微信外,所有平台均可无需 API Key 使用 3. 智能降级策略 — API 优先 → 爬虫模式 → 公开接口,三级自动降级 4. 修复 Issue #1 — 修复 marketplace.json 缺少 owner 字段导致 Claude Code 安装失败的 bug 5. 登录态缓存 — 爬虫模式支持 Cookie 持久化,减少重复登录
---
📋 平台支持
| 平台 | 模块 | 数据获取方式 | 需要配置 |
|---|---|---|---|
| 🔴 微博 | weibo.py | API / 🕷️爬虫 / 公开接口 | ✅ 爬虫模式无需配置 |
| 📕 小红书 | xiaohongshu.py | API / 🕷️爬虫 / 公开接口 | ✅ 爬虫模式无需配置 |
| 📺 B站 | bilibili.py | 公开 API / 🕷️爬虫备用 | ✅ 无需配置 |
| 💬 知乎 | zhihu.py | 公开搜索 / 🕷️爬虫备用 | ✅ 无需配置 |
| 🎵 抖音 | douyin.py | API / 🕷️爬虫 / 公开接口 / 搜索兜底 | ✅ 爬虫模式无需配置 |
| 💚 微信 | wechat.py | API / 搜狗搜索 | WECHAT_API_KEY(可选) |
| 🔵 百度 | baidu.py | 公开搜索 / API | ✅ 基础搜索无需配置 |
| 📰 头条 | toutiao.py | 公开接口 / 搜索兜底 | ✅ 无需配置 |
🕷️ = 需要安装 Playwright(pip install playwright && playwright install chromium)---
🤖 Agent 平台安装
Agent Skills(推荐)
npx skills add Jesseovo/last30days-skill-cn -gCursor(推荐)
将项目克隆到 Cursor 技能目录:
git clone https://github.com/Jesseovo/last30days-skill-cn.git然后在 Cursor 中将 SKILL.md 添加为项目技能。
Claude Code
# 方式一:通过 Agent Skills 安装(推荐)
npx skills add Jesseovo/last30days-skill-cn -g
# 方式二:手动安装
git clone https://github.com/Jesseovo/last30days-skill-cn.git ~/.claude/skills/last30days-cnOpenClaw / ClawHub
git clone https://github.com/Jesseovo/last30days-skill-cn.git ~/.agents/skills/last30days-cnGemini CLI
git clone https://github.com/Jesseovo/last30days-skill-cn.git
# 在 Gemini CLI 中作为扩展加载通用 Agent
任何支持 Bash / Read / Write 工具的 AI Agent 都可以使用本技能。
---
⚙️ 配置指南
📍 第一步:安装依赖
pip install jieba📍 第二步:安装爬虫引擎(推荐,可获取 7/8 平台数据)
pip install playwright
playwright install chromium安装 Playwright 后,微博、小红书、抖音、B站(备用)、知乎(备用)均可无需 API Key 使用。
📍 第三步:创建配置文件(可选)
如果您希望使用 API 模式获取更稳定的数据,或需要使用微信公众号搜索:
mkdir -p ~/.config/last30days-cn
touch ~/.config/last30days-cn/.env
chmod 600 ~/.config/last30days-cn/.envWindows(PowerShell)等价: 创建目录与空配置文件可用 New-Item -ItemType Directory -Force -Path "$env:USERPROFILE\.config\last30days-cn" 与 New-Item -ItemType File -Path "$env:USERPROFILE\.config\last30days-cn\.env" -Force。限制 .env 仅当前用户可读写可近似使用 icacls "$env:USERPROFILE\.config\last30days-cn\.env" /inheritance:r /grant:r "$($env:USERNAME):(R,W)"(与 Unix chmod 600 意图相近,权限模型不同)。
编辑 ~/.config/last30days-cn/.env,按需填入 API Key:
# ============================================
# last30days-cn v2.0 配置文件
# ============================================
# 📌 说明:所有 API Key 均为可选
# 安装 Playwright 后,大部分平台已可通过爬虫模式使用
# API Key 提供更稳定的数据获取方式
# ============================================
# 🔴 微博开放平台(可选,已有爬虫模式替代)
# 获取方式: https://open.weibo.com → 创建应用 → 获取 Access Token
WEIBO_ACCESS_TOKEN=
# 📕 小红书(可选,已有爬虫模式替代)
# 获取方式: https://scrapecreators.com → 注册 → 获取 API Key
SCRAPECREATORS_API_KEY=
# 💬 知乎 Cookie(可选,增强搜索质量)
# 获取方式: 浏览器登录知乎 → F12 → Network → 复制 Cookie 值
ZHIHU_COOKIE=
# 🎵 抖音(可选,已有爬虫模式替代)
# 获取方式: https://tikhub.io → 注册 → 获取 API Key
TIKHUB_API_KEY=
# 💚 微信公众号搜索(目前无爬虫替代,需 API Key 才能使用)
# 获取方式: 使用第三方微信搜索 API 服务商
WECHAT_API_KEY=
# 🔵 百度搜索 API(可选,公开搜索已可用)
# 获取方式: https://cloud.baidu.com → 搜索服务 → 创建应用
BAIDU_API_KEY=
BAIDU_SECRET_KEY=📍 第四步:验证配置
python scripts/last30days.py --diagnose将输出各平台的可用状态和爬虫引擎状态:
{
"weibo": true,
"xiaohongshu": false,
"bilibili": true,
"zhihu": true,
"douyin": true,
"wechat": false,
"baidu_api": false,
"toutiao": true,
"crawler_engine": {
"playwright_available": true,
"cached_logins": [],
"note": "安装 Playwright 后,微博/小红书/抖音/B站/知乎可无需 API Key 使用爬虫模式"
},
"note_douyin_toutiao": "抖音/头条原生接口需签名参数,常被风控;接口失败时改用公开搜索引擎兜底,仅能拿到公开链接,无真实互动数据与精确日期。"
}---
🚀 使用方式
基本用法
python scripts/last30days.py "AI编程助手" --emit compact
python scripts/last30days.py "AI编程助手" --emit html-path命令行参数
| 参数 | 说明 | 示例 |
|---|---|---|
--emit | 输出模式 | compact / json / md / context / path / html / html-path |
--quick | 快速搜索 | 更少数据源,更快速度 |
--deep | 深度搜索 | 更多数据源,更全面 |
--days N | 回溯天数 | --days 7(最近一周) |
--as-of | 历史回溯终点日期 | --as-of 2026-05-01(以该日为终点回溯 N 天) |
--search | 指定搜索源 | --search weibo,bilibili,zhihu |
--diagnose | 诊断配置 | 显示各平台可用状态 |
--timeout SECS | 全局超时秒数 | 覆盖默认全局超时 |
--save-dir DIR | 自动保存原始输出目录 | 将原始输出写入指定目录 |
--debug | 调试模式 | 输出详细日志 |
🔧 环境变量:未指定--search时,回退到LAST30DAYS_DEFAULT_SEARCH(逗号分隔的默认源集);EXCLUDE_SOURCES可从启用集合中排除指定源。
使用示例
# 🔍 搜索 AI 相关话题
python scripts/last30days.py "最新AI工具" --emit compact
# ⚡ 快速搜索,仅 B站和知乎
python scripts/last30days.py "Python教程" --quick --search bilibili,zhihu
# 📊 深度搜索并保存结果
python scripts/last30days.py "新能源汽车" --deep --save-dir ~/Documents/research
# 📋 输出 JSON 格式(适合程序处理)
python scripts/last30days.py "ChatGPT替代品" --emit json
# 🗓️ 仅搜索最近 7 天
python scripts/last30days.py "热门话题" --days 7
# 🧾 生成可离线打开的 HTML 报告
python scripts/last30days.py "具身智能" --deep --emit html-path---
🔧 数据获取策略(三级降级)
v2.0 采用三级自动降级策略,确保最大可用性:
优先级 1: API 模式(如配置了 API Key)
↓ 失败或未配置
优先级 2: 爬虫模式(MediaCrawler,需要 Playwright)
↓ 失败或未安装
优先级 3: 公开接口(HTTP 直接请求,无需任何配置)
↓ 仍无结果(抖音/头条/小红书/知乎)
兜底: 公开搜索引擎(Bing site: 搜索,获取公开链接)各平台数据获取方式对比
| 平台 | API 模式 | 爬虫模式 | 公开接口 | 搜索兜底 |
|---|---|---|---|---|
| 微博 | WEIBO_ACCESS_TOKEN | ✅ Playwright | ✅ m.weibo.cn | - |
| 小红书 | MCP HTTP API(可选) | ✅ Playwright (XHR拦截) | ⚠️ 命中率低 | ✅ Bing |
| B站 | - | ✅ Playwright(备用) | ✅ 公开 API | - |
| 知乎 | ZHIHU_COOKIE(增强) | ✅ Playwright(备用) | ✅ 公开搜索 | ✅ Bing |
| 抖音 | TIKHUB_API_KEY | ✅ Playwright | ⚠️ 需签名 | ✅ Bing |
| 微信 | WECHAT_API_KEY | - | ✅ 搜狗搜索 | - |
| 百度 | BAIDU_API_KEY | - | ⚠️ 公开搜索可能被拦截 | ✅ Bing |
| 头条 | - | - | ⚠️ 需签名 | ✅ Bing |
⚠️ 抖音/头条原生 web 接口现在强制要求签名参数(a_bogus/_signature),常被风控。接口失败时改用公开搜索引擎兜底,只能拿到公开链接,无真实互动数据与精确日期。
---
🏗️ 项目架构
last30days-skill-cn/
├── 📄 SKILL.md # Agent 技能定义文件
├── 📄 README.md # 项目说明(中文,本文件)
├── 📄 README.en.md # 项目说明(English)
├── 📄 LICENSE # MIT 许可证
├── 📄 requirements.txt # Python 依赖
├── 📁 assets/ # README 配图
├── 📁 scripts/
│ ├── 🐍 last30days.py # 中文主入口 CLI
│ └── 📁 lib/
│ ├── crawler_bridge.py # 🆕 MediaCrawler 爬虫桥接模块
│ ├── weibo.py # 微博搜索模块
│ ├── xiaohongshu.py # 小红书搜索模块
│ ├── bilibili.py # B站搜索模块
│ ├── zhihu.py # 知乎搜索模块
│ ├── douyin.py # 抖音搜索模块
│ ├── wechat.py # 微信公众号模块
│ ├── baidu.py # 百度搜索模块
│ ├── toutiao.py # 今日头条模块
│ ├── schema.py # 数据结构定义
│ ├── score.py # 评分系统
│ ├── normalize.py # 数据标准化
│ ├── dedupe.py # 去重
│ ├── render.py # 输出渲染
│ ├── relevance.py # 相关性计算
│ ├── query.py # 查询预处理
│ ├── query_type.py # 查询类型检测
│ ├── entity_extract.py # 实体抽取
│ ├── env.py # 环境配置管理
│ ├── cache.py # 缓存管理
│ ├── dates.py # 日期工具
│ ├── http.py # HTTP 客户端
│ ├── ui.py # 终端 UI
│ └── setup_wizard.py # 配置向导
├── 📁 skills/
│ └── 📁 last30days/ # Agent Skills 自包含运行载荷
│ ├── 📄 SKILL.md # 安装后供 Agent 读取的技能说明
│ └── 📁 scripts/
│ ├── 🐍 last30days.py
│ └── 📁 lib/ # 与根目录脚本同步的运行依赖
├── 📁 fixtures/ # 示例数据
├── 📁 tests/ # 测试用例
└── 📁 hooks/ # Agent 钩子---
📊 评分系统
每条搜索结果的综合评分(0-100)基于:
| 维度 | 权重 | 说明 |
|---|---|---|
| 🎯 相关性 | 45% | 与查询主题的文本匹配度 |
| 🕐 时效性 | 25% | 内容发布时间的新鲜程度 |
| 🔥 互动度 | 30% | 各平台互动指标(见下表) |
各平台互动指标
| 平台 | 互动指标 |
|---|---|
| 微博 | 转发 + 评论 + 点赞 |
| 小红书 | 点赞 + 收藏 + 评论 + 分享 |
| B站 | 播放 + 弹幕 + 评论 + 投币 + 收藏 |
| 知乎 | 赞同 + 评论 + 收藏 |
| 抖音 | 点赞 + 评论 + 分享 + 播放 |
| 头条 | 评论 + 阅读 + 点赞 |
---
🙏 致谢
- mvanhorn/last30days-skill — 原始英文版项目
- NanmiCoder/MediaCrawler — 爬虫引擎技术灵感来源
---
📜 许可证
本项目基于 MIT License 发布。
- 🔗 原始项目: mvanhorn/last30days-skill by Matt Van Horn
- 🇨🇳 中文本土化: Jesse (@Jesseovo)
Release Notes
v3.0.0
中文分支 v3 维护升级。
重点更新
- 新增
skills/last30days自包含 Agent Skills 运行载荷。 - 中文 CLI 统一使用单入口
last30days.py,根目录和 Skill 载荷保持同名结构。 - 新增
--emit html和--emit html-path,支持生成可离线打开的report.html。 - 引入受
op7418/guizang-ppt-skill启发的 Swiss/IKB HTML 报告样式。 - README、Skill 说明、SPEC 与同步脚本统一改为推荐
last30days.py。 - README 在原先长文档基础上合入 v3 内容,保留免责声明、平台支持、配置、评分系统和中英文说明。
修复与改进
- 小红书和知乎在 API 与 Playwright 路径返回空结果时,增加公开站内搜索兜底。
quick模式下小红书和知乎跳过较慢的 Playwright 路径,避免长时间等待后超时。- 小红书和知乎空结果时会明确说明已尝试路径和可能原因。
--diagnose在可尝试兜底路径时会把小红书标记为可用。- 增加小红书/知乎兜底解析与可用性诊断回归测试。
- 增加 HTML 渲染回归测试。
兼容性
根目录 scripts/ 仍保留用于本地开发和旧用法。通过 Agent Skills 安装时,推荐使用 {{SKILL_DIR}}/scripts/last30days.py。
已知限制
小红书和知乎仍可能因登录态失效、验证码/反爬、平台 API 变化、搜索引擎未收录公开链接而返回 0 条。当前版本会明确暴露原因,不再静默失败。
验证
py -m pytest tests/test_html_render.py tests/test_render_wechat.pyv2.1.0
- 修复微信公众号渲染回归。
- 改进百度与小红书兜底行为。
- 增加爬虫和反爬相关回归覆盖。
v2.0.0
- 增加受 MediaCrawler 启发的 Playwright 兜底路径。
- 降低强制 API Key 依赖。
- 增加多 Agent 兼容文档。
v1.0.0
- 首次完成
mvanhorn/last30days-skill的中文平台本土化。
jieba>=0.42.1
playwright>=1.40.0
#!/usr/bin/env python3
"""
last30days-cn - 研究过去30天内中国平台上的热门话题。
Author: Jesse (https://github.com/Jesseovo)
Usage:
python3 last30days.py <topic> [options]
Options:
--emit=MODE 输出模式: compact|json|md|html|context|path|html-path (default: compact)
--quick 快速搜索,减少数据源
--deep 深度搜索,更多数据源
--debug 启用调试日志
--days N 回溯天数 (1-30, default: 30)
--as-of DATE 历史回溯:以 YYYY-MM-DD 为终点回溯 N 天 (default: 今天)
--search SOURCES 指定搜索源 (逗号分隔): weibo,xiaohongshu,bilibili,zhihu,douyin,wechat,baidu,toutiao
未指定时回退到环境变量 LAST30DAYS_DEFAULT_SEARCH;EXCLUDE_SOURCES 可排除源
--diagnose 显示数据源可用性诊断
"""
import argparse
import atexit
import json
import os
import signal
import sys
import threading
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime, timezone
from pathlib import Path
SCRIPT_DIR = Path(__file__).parent.resolve()
sys.path.insert(0, str(SCRIPT_DIR))
_child_pids: set = set()
_child_pids_lock = threading.Lock()
TIMEOUT_PROFILES = {
"quick": {"global": 90, "future": 30, "weibo_future": 30, "bilibili_future": 30, "zhihu_future": 30, "douyin_future": 60, "xiaohongshu_future": 30, "wechat_future": 30, "baidu_future": 30, "toutiao_future": 15, "http": 15},
"default": {"global": 180, "future": 60, "weibo_future": 60, "bilibili_future": 60, "zhihu_future": 60, "douyin_future": 90, "xiaohongshu_future": 60, "wechat_future": 60, "baidu_future": 60, "toutiao_future": 30, "http": 30},
"deep": {"global": 300, "future": 90, "weibo_future": 90, "bilibili_future": 90, "zhihu_future": 90, "douyin_future": 120, "xiaohongshu_future": 90, "wechat_future": 90, "baidu_future": 90, "toutiao_future": 45, "http": 30},
}
VALID_SEARCH_SOURCES = {
"weibo", "xiaohongshu", "xhs", "bilibili", "zhihu",
"douyin", "wechat", "baidu", "toutiao",
}
def parse_search_flag(search_str: str) -> set:
sources = set()
for s in search_str.split(","):
s = s.strip().lower()
if not s:
continue
if s == "xhs":
s = "xiaohongshu"
if s not in VALID_SEARCH_SOURCES:
print(f"错误: 未知搜索源 '{s}'。可用: {', '.join(sorted(VALID_SEARCH_SOURCES))}", file=sys.stderr)
sys.exit(1)
sources.add(s)
if not sources:
print("错误: --search 需要至少一个搜索源。", file=sys.stderr)
sys.exit(1)
return sources
# 8 个规范源(不含 xhs 别名),用于默认/排除集合运算
ALL_SOURCE_IDS = {
"weibo", "xiaohongshu", "bilibili", "zhihu",
"douyin", "wechat", "baidu", "toutiao",
}
def resolve_search_sources(cli_search):
"""解析最终启用的搜索源。
优先级: --search > 环境变量 LAST30DAYS_DEFAULT_SEARCH > 全部源;
随后减去环境变量 EXCLUDE_SOURCES(逗号分隔)。
返回 None 表示"按查询类型自动启用全部源"(沿用 run_research 既有行为)。
"""
sources = None
if cli_search:
sources = parse_search_flag(cli_search)
else:
default_env = os.environ.get("LAST30DAYS_DEFAULT_SEARCH", "").strip()
if default_env:
sources = parse_search_flag(default_env)
exclude_env = os.environ.get("EXCLUDE_SOURCES", "").strip()
if exclude_env:
excluded = {
("xiaohongshu" if s.strip().lower() == "xhs" else s.strip().lower())
for s in exclude_env.split(",")
if s.strip()
}
base = sources if sources is not None else set(ALL_SOURCE_IDS)
sources = base - excluded
if not sources:
print("错误: EXCLUDE_SOURCES 排除后没有可用搜索源。", file=sys.stderr)
sys.exit(1)
return sources
def _cleanup_children():
with _child_pids_lock:
pids = list(_child_pids)
for pid in pids:
try:
os.kill(pid, signal.SIGTERM)
except (ProcessLookupError, PermissionError, OSError):
pass
atexit.register(_cleanup_children)
def _install_global_timeout(timeout_seconds: int):
if hasattr(signal, 'SIGALRM'):
def _handler(signum, frame):
sys.stderr.write(f"\n[超时] 全局超时 ({timeout_seconds}s) 已超过。正在清理。\n")
sys.stderr.flush()
_cleanup_children()
sys.exit(1)
signal.signal(signal.SIGALRM, _handler)
signal.alarm(timeout_seconds)
else:
def _watchdog():
sys.stderr.write(f"\n[超时] 全局超时 ({timeout_seconds}s) 已超过。正在清理。\n")
sys.stderr.flush()
_cleanup_children()
os._exit(1)
timer = threading.Timer(timeout_seconds, _watchdog)
timer.daemon = True
timer.start()
from lib import (
weibo,
xiaohongshu,
bilibili,
zhihu,
douyin,
wechat,
baidu,
toutiao,
dates,
dedupe,
cluster,
env,
normalize,
query,
render,
schema,
score,
setup_wizard,
query_type as qt,
crawler_bridge,
)
def _search_weibo(topic, config, from_date, to_date, depth):
try:
token = config.get("WEIBO_ACCESS_TOKEN")
items = weibo.search_weibo(topic, from_date, to_date, depth=depth, token=token)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_xiaohongshu(topic, config, from_date, to_date, depth):
try:
token = config.get("SCRAPECREATORS_API_KEY")
api_base = env.get_xiaohongshu_api_base(config)
items = xiaohongshu.search_xiaohongshu(topic, from_date, to_date, depth=depth, token=token, api_base=api_base)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_bilibili(topic, from_date, to_date, depth):
try:
items = bilibili.search_bilibili(topic, from_date, to_date, depth=depth)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_zhihu(topic, config, from_date, to_date, depth):
try:
cookie = config.get("ZHIHU_COOKIE")
items = zhihu.search_zhihu(topic, from_date, to_date, depth=depth, cookie=cookie)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_douyin(topic, config, from_date, to_date, depth):
try:
token = config.get("TIKHUB_API_KEY") or config.get("DOUYIN_API_KEY")
items = douyin.search_douyin(topic, from_date, to_date, depth=depth, token=token)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_wechat(topic, config, from_date, to_date, depth):
try:
api_key = config.get("WECHAT_API_KEY")
items = wechat.search_wechat(topic, from_date, to_date, depth=depth, api_key=api_key)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_baidu(topic, config, from_date, to_date, depth):
try:
api_key = config.get("BAIDU_API_KEY")
secret_key = config.get("BAIDU_SECRET_KEY")
items = baidu.search_baidu(topic, from_date, to_date, depth=depth, api_key=api_key, secret_key=secret_key)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def _search_toutiao(topic, from_date, to_date, depth):
try:
items = toutiao.search_toutiao(topic, from_date, to_date, depth=depth)
return items, None
except Exception as e:
return [], f"{type(e).__name__}: {e}"
def run_research(
topic: str,
config: dict,
from_date: str,
to_date: str,
depth: str = "default",
timeouts: dict = None,
search_sources: set = None,
query_type: str = "breaking_news",
) -> dict:
if timeouts is None:
timeouts = TIMEOUT_PROFILES[depth]
future_timeout = timeouts["future"]
all_sources = {"weibo", "xiaohongshu", "bilibili", "zhihu", "douyin", "wechat", "baidu", "toutiao"}
if search_sources:
active = search_sources & all_sources
else:
active = {s for s in all_sources if qt.is_source_enabled(s, query_type)}
results = {src: {"items": [], "error": None} for src in all_sources}
futures = {}
max_workers = len(active)
with ThreadPoolExecutor(max_workers=max(max_workers, 1)) as executor:
if "weibo" in active:
sys.stderr.write("[微博] 搜索中...\n")
futures["weibo"] = executor.submit(_search_weibo, topic, config, from_date, to_date, depth)
if "xiaohongshu" in active:
sys.stderr.write("[小红书] 搜索中...\n")
futures["xiaohongshu"] = executor.submit(_search_xiaohongshu, topic, config, from_date, to_date, depth)
if "bilibili" in active:
sys.stderr.write("[B站] 搜索中...\n")
futures["bilibili"] = executor.submit(_search_bilibili, topic, from_date, to_date, depth)
if "zhihu" in active:
sys.stderr.write("[知乎] 搜索中...\n")
futures["zhihu"] = executor.submit(_search_zhihu, topic, config, from_date, to_date, depth)
if "douyin" in active:
sys.stderr.write("[抖音] 搜索中...\n")
futures["douyin"] = executor.submit(_search_douyin, topic, config, from_date, to_date, depth)
if "wechat" in active:
sys.stderr.write("[微信] 搜索中...\n")
futures["wechat"] = executor.submit(_search_wechat, topic, config, from_date, to_date, depth)
if "baidu" in active:
sys.stderr.write("[百度] 搜索中...\n")
futures["baidu"] = executor.submit(_search_baidu, topic, config, from_date, to_date, depth)
if "toutiao" in active:
sys.stderr.write("[头条] 搜索中...\n")
futures["toutiao"] = executor.submit(_search_toutiao, topic, from_date, to_date, depth)
for source, future in futures.items():
timeout = timeouts.get(f"{source}_future", future_timeout)
try:
items, error = future.result(timeout=timeout)
results[source]["items"] = items
results[source]["error"] = error
if error:
sys.stderr.write(f"[{source}] 错误: {error}\n")
else:
sys.stderr.write(f"[{source}] {len(items)} 条结果\n")
except TimeoutError:
results[source]["error"] = f"{source} 搜索超时 ({timeout}s)"
sys.stderr.write(f"[{source}] 超时 ({timeout}s)\n")
except Exception as e:
results[source]["error"] = f"{type(e).__name__}: {e}"
sys.stderr.write(f"[{source}] 错误: {e}\n")
sys.stderr.flush()
return results
def main():
if sys.platform == "win32":
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
sys.stderr.reconfigure(encoding="utf-8", errors="replace")
parser = argparse.ArgumentParser(description="研究过去N天内中国平台上的热门话题")
parser.add_argument("topic", nargs="*", help="研究主题")
parser.add_argument("--emit", choices=["compact", "json", "md", "html", "context", "path", "html-path"], default="compact", help="输出模式")
parser.add_argument("--quick", action="store_true", help="快速搜索")
parser.add_argument("--deep", action="store_true", help="深度搜索")
parser.add_argument("--debug", action="store_true", help="启用调试日志")
parser.add_argument("--days", type=int, default=30, choices=range(1, 31), metavar="N", help="回溯天数 (1-30)")
parser.add_argument("--as-of", dest="as_of", type=str, default=None, metavar="YYYY-MM-DD", help="历史回溯:以指定日期为终点回溯 N 天")
parser.add_argument("--diagnose", action="store_true", help="显示数据源诊断")
parser.add_argument("--timeout", type=int, default=None, metavar="SECS", help="全局超时秒数")
parser.add_argument("--search", type=str, default=None, metavar="SOURCES", help="逗号分隔的搜索源列表")
parser.add_argument("--save-dir", type=str, default=None, metavar="DIR", help="自动保存原始输出")
args = parser.parse_args()
args.topic = " ".join(args.topic) if args.topic else None
if args.debug:
os.environ["LAST30DAYS_DEBUG"] = "1"
if args.quick and args.deep:
print("错误: 不能同时使用 --quick 和 --deep", file=sys.stderr)
sys.exit(1)
elif args.quick:
depth = "quick"
elif args.deep:
depth = "deep"
else:
depth = "default"
timeouts = TIMEOUT_PROFILES[depth]
global_timeout = args.timeout or timeouts["global"]
_install_global_timeout(global_timeout)
config = env.get_config()
if args.diagnose:
crawler_status = crawler_bridge.get_crawler_status()
diag = {
"weibo": env.is_weibo_available(config),
"xiaohongshu": env.is_xiaohongshu_available(config),
"bilibili": env.probe_bilibili(),
"zhihu": env.probe_zhihu(),
"douyin": env.is_douyin_available(config),
"wechat": env.is_wechat_available(config),
"baidu_api": env.is_baidu_api_available(config),
"toutiao": env.probe_toutiao(),
"xiaohongshu_api_base": env.get_xiaohongshu_api_base(config),
"crawler_engine": {
"playwright_available": crawler_status["playwright_available"],
"cached_logins": crawler_status["cached_logins"],
"note": "安装 Playwright 后,微博/小红书/抖音/B站/知乎可无需 API Key 使用爬虫模式",
},
"note_douyin_toutiao": "抖音/头条原生接口需签名参数,常被风控;接口失败时改用公开搜索引擎兜底,仅能拿到公开链接,无真实互动数据与精确日期。",
}
print(json.dumps(diag, indent=2, ensure_ascii=False))
sys.exit(0)
if args.topic and args.topic.strip().lower() == "setup":
results = setup_wizard.run_auto_setup(config)
env_path = env.CONFIG_FILE
if env_path:
written = setup_wizard.write_setup_config(env_path)
results["env_written"] = written
else:
results["env_written"] = False
print(setup_wizard.get_setup_status_text(results))
sys.exit(0)
if not args.topic:
print("错误: 请提供研究主题。", file=sys.stderr)
print("用法: python3 last30days.py <topic> [options]", file=sys.stderr)
sys.exit(1)
try:
from_date, to_date = dates.get_date_range(args.days, as_of=args.as_of)
except ValueError as e:
print(f"错误: {e}", file=sys.stderr)
sys.exit(1)
search_sources = resolve_search_sources(args.search)
query_type = qt.detect_query_type(args.topic)
search_topic = query.extract_core_subject(args.topic)
sys.stderr.write(f"正在搜索: {args.topic}\n")
if search_topic != args.topic:
sys.stderr.write(f"提纯关键词: {search_topic}\n")
sys.stderr.write(f"查询类型: {query_type} | 日期范围: {from_date} 至 {to_date}\n")
sys.stderr.flush()
raw_results = run_research(
search_topic, config, from_date, to_date, depth,
timeouts=timeouts, search_sources=search_sources,
query_type=query_type,
)
for source in ("xiaohongshu", "zhihu"):
if source in (search_sources or set()) and not raw_results[source]["items"] and not raw_results[source]["error"]:
if source == "xiaohongshu":
raw_results[source]["error"] = (
"未获取到结果;已尝试 MCP/公开接口/站内搜索兜底。"
"默认或 deep 模式还会尝试 Playwright。若仍为空,通常是登录态失效、验证码、平台反爬或搜索引擎未收录。"
)
else:
raw_results[source]["error"] = (
"未获取到结果;已尝试知乎 API/热榜/站内搜索兜底。"
"默认或 deep 模式还会尝试 Playwright。若仍为空,通常是 API 限制、登录态失效、反爬验证或搜索引擎未收录。"
)
sys.stderr.write("正在处理结果...\n")
sys.stderr.flush()
norm_weibo = normalize.normalize_weibo_items(raw_results["weibo"]["items"], from_date, to_date)
norm_xhs = normalize.normalize_xiaohongshu_items(raw_results["xiaohongshu"]["items"], from_date, to_date)
norm_bili = normalize.normalize_bilibili_items(raw_results["bilibili"]["items"], from_date, to_date)
norm_zhihu = normalize.normalize_zhihu_items(raw_results["zhihu"]["items"], from_date, to_date)
norm_douyin = normalize.normalize_douyin_items(raw_results["douyin"]["items"], from_date, to_date)
norm_wechat = normalize.normalize_wechat_items(raw_results["wechat"]["items"], from_date, to_date)
norm_baidu = normalize.normalize_baidu_items(raw_results["baidu"]["items"], from_date, to_date)
norm_toutiao = normalize.normalize_toutiao_items(raw_results["toutiao"]["items"], from_date, to_date)
filt_weibo = normalize.filter_by_date_range(norm_weibo, from_date, to_date)
filt_xhs = normalize.filter_by_date_range(norm_xhs, from_date, to_date)
filt_bili = normalize.filter_by_date_range(norm_bili, from_date, to_date)
filt_zhihu = normalize.filter_by_date_range(norm_zhihu, from_date, to_date)
filt_douyin = normalize.filter_by_date_range(norm_douyin, from_date, to_date)
filt_wechat = normalize.filter_by_date_range(norm_wechat, from_date, to_date)
filt_baidu = normalize.filter_by_date_range(norm_baidu, from_date, to_date)
filt_toutiao = normalize.filter_by_date_range(norm_toutiao, from_date, to_date)
scored_weibo = score.score_weibo_items(filt_weibo)
scored_xhs = score.score_xiaohongshu_items(filt_xhs)
scored_bili = score.score_bilibili_items(filt_bili)
scored_zhihu = score.score_zhihu_items(filt_zhihu)
scored_douyin = score.score_douyin_items(filt_douyin)
scored_wechat = score.score_wechat_items(filt_wechat, query_type=query_type)
scored_baidu = score.score_baidu_items(filt_baidu, query_type=query_type)
scored_toutiao = score.score_toutiao_items(filt_toutiao)
sorted_weibo = score.sort_items(scored_weibo, query_type=query_type)
sorted_xhs = score.sort_items(scored_xhs, query_type=query_type)
sorted_bili = score.sort_items(scored_bili, query_type=query_type)
sorted_zhihu = score.sort_items(scored_zhihu, query_type=query_type)
sorted_douyin = score.sort_items(scored_douyin, query_type=query_type)
sorted_wechat = score.sort_items(scored_wechat, query_type=query_type)
sorted_baidu = score.sort_items(scored_baidu, query_type=query_type)
sorted_toutiao = score.sort_items(scored_toutiao, query_type=query_type)
deduped_weibo = dedupe.dedupe_weibo(sorted_weibo)
deduped_xhs = dedupe.dedupe_xiaohongshu(sorted_xhs)
deduped_bili = dedupe.dedupe_bilibili(sorted_bili)
deduped_zhihu = dedupe.dedupe_zhihu(sorted_zhihu)
deduped_douyin = dedupe.dedupe_douyin(sorted_douyin)
deduped_wechat = dedupe.dedupe_wechat(sorted_wechat)
deduped_baidu = dedupe.dedupe_baidu(sorted_baidu)
deduped_toutiao = dedupe.dedupe_toutiao(sorted_toutiao)
deduped_weibo = score.relevance_filter(deduped_weibo, "WEIBO")
deduped_xhs = score.relevance_filter(deduped_xhs, "XIAOHONGSHU")
deduped_bili = score.relevance_filter(deduped_bili, "BILIBILI")
deduped_zhihu = score.relevance_filter(deduped_zhihu, "ZHIHU")
deduped_douyin = score.relevance_filter(deduped_douyin, "DOUYIN")
deduped_wechat = score.relevance_filter(deduped_wechat, "WECHAT")
deduped_baidu = score.relevance_filter(deduped_baidu, "BAIDU")
deduped_toutiao = score.relevance_filter(deduped_toutiao, "TOUTIAO")
dedupe.cross_source_link(
deduped_weibo, deduped_xhs, deduped_bili, deduped_zhihu,
deduped_douyin, deduped_wechat, deduped_baidu, deduped_toutiao,
)
clusters = cluster.build_clusters(
deduped_weibo, deduped_xhs, deduped_bili, deduped_zhihu,
deduped_douyin, deduped_wechat, deduped_baidu, deduped_toutiao,
)
report = schema.create_report(args.topic, from_date, to_date, "all")
report.clusters = clusters
report.weibo = deduped_weibo
report.xiaohongshu = deduped_xhs
report.bilibili = deduped_bili
report.zhihu = deduped_zhihu
report.douyin = deduped_douyin
report.wechat = deduped_wechat
report.baidu = deduped_baidu
report.toutiao = deduped_toutiao
report.weibo_error = raw_results["weibo"]["error"]
report.xiaohongshu_error = raw_results["xiaohongshu"]["error"]
report.bilibili_error = raw_results["bilibili"]["error"]
report.zhihu_error = raw_results["zhihu"]["error"]
report.douyin_error = raw_results["douyin"]["error"]
report.wechat_error = raw_results["wechat"]["error"]
report.baidu_error = raw_results["baidu"]["error"]
report.toutiao_error = raw_results["toutiao"]["error"]
report.context_snippet_md = render.render_context_snippet(report)
render.write_outputs(report)
total = sum(len(getattr(report, src, [])) for src in ["weibo", "xiaohongshu", "bilibili", "zhihu", "douyin", "wechat", "baidu", "toutiao"])
sys.stderr.write(f"\n完成! 共 {total} 条结果\n")
sys.stderr.flush()
if args.emit == "compact":
print(render.render_compact(report))
print(render.render_source_status(report))
elif args.emit == "json":
print(json.dumps(report.to_dict(), indent=2, ensure_ascii=False))
elif args.emit == "md":
print(render.render_full_report(report))
elif args.emit == "html":
print(render.render_html_report(report))
elif args.emit == "context":
print(report.context_snippet_md)
elif args.emit == "path":
print(render.get_context_path())
elif args.emit == "html-path":
print(render.get_html_path())
if args.save_dir:
import re as re_mod
save_dir = Path(args.save_dir).expanduser()
save_dir.mkdir(parents=True, exist_ok=True)
slug = re_mod.sub(r'[^a-z0-9\u4e00-\u9fff]+', '-', args.topic.lower()).strip('-')[:60]
save_path = save_dir / f"{slug}-raw.md"
if save_path.exists():
save_path = save_dir / f"{slug}-raw-{datetime.now().strftime('%Y-%m-%d')}.md"
content = render.render_compact(report)
content += "\n" + render.render_source_status(report)
save_path.write_text(content, encoding="utf-8")
print(f"已保存: {save_path}", file=sys.stderr)
if __name__ == "__main__":
main()
# last30days library modules
"""百度搜索模块 - 替代 Exa/Brave 等 Web 搜索后端。
Author: Jesse (https://github.com/Jesseovo)
v2.1 改动:
- 更真实的浏览器头(Accept-Language/Referer/Sec-Fetch-*)+ UA 轮换
- 若检测到「百度安全验证」拦截页,主动记录日志并降级
- 新增 Bing 国内版兜底(`https://cn.bing.com/search`)
- 同时兼容百度新版 HTML 的 `result c-container` 块级选择器
"""
import json
import random
import re
import sys
import time
import urllib.parse
import urllib.request
from typing import Any, Dict, List, Optional
from . import relevance
_UA_POOL = [
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36 Edg/125.0.0.0",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
]
_ANTIBOT_SIGNATURES = (
"wappass.baidu.com",
"百度安全验证",
"百度智能云验证",
"verify.baidu.com",
"/static/verify",
)
def search_baidu(
topic: str,
from_date: str,
to_date: str,
depth: str = "default",
api_key: Optional[str] = None,
secret_key: Optional[str] = None,
) -> List[Dict[str, Any]]:
"""搜索百度网页。
Args:
topic: 搜索关键词
from_date: 起始日期
to_date: 结束日期
depth: 搜索深度
api_key: 百度 API Key(可选)
secret_key: 百度 Secret Key(可选)
Returns:
网页搜索结果列表
"""
limit_map = {"quick": 10, "default": 20, "deep": 30}
limit = limit_map.get(depth, 20)
items: List[Dict[str, Any]] = []
if api_key and secret_key:
items = _search_via_api(topic, limit, api_key, secret_key)
if not items:
items = _search_via_public(topic, limit)
if not items:
items = _search_via_bing(topic, limit)
scored = []
for i, item in enumerate(items):
title = item.get("title", "")
snippet = item.get("snippet", "")
combined = f"{title} {snippet}"
rel = relevance.token_overlap_relevance(topic, combined)
item["id"] = f"BD{i+1}"
item["relevance"] = rel
item["why_relevant"] = f"{item.get('source_tag', '百度搜索')}:{title[:50]}"
scored.append(item)
scored.sort(key=lambda x: x.get("relevance", 0), reverse=True)
return scored[:limit]
def _build_headers(referer: str) -> Dict[str, str]:
ua = random.choice(_UA_POOL)
return {
"User-Agent": ua,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Accept-Encoding": "identity",
"Referer": referer,
"Sec-Ch-Ua": '"Chromium";v="124", "Not-A.Brand";v="99"',
"Sec-Ch-Ua-Mobile": "?0",
"Sec-Ch-Ua-Platform": '"Windows"',
"Sec-Fetch-Dest": "document",
"Sec-Fetch-Mode": "navigate",
"Sec-Fetch-Site": "same-origin",
"Sec-Fetch-User": "?1",
"Upgrade-Insecure-Requests": "1",
}
def _is_antibot_page(html: str) -> bool:
if not html:
return True
if len(html) < 2000:
return any(sig in html for sig in _ANTIBOT_SIGNATURES)
return any(sig in html for sig in _ANTIBOT_SIGNATURES)
def _fetch(url: str, headers: Dict[str, str], timeout: int = 15) -> str:
req = urllib.request.Request(url, headers=headers)
with urllib.request.urlopen(req, timeout=timeout) as response:
return response.read().decode("utf-8", errors="replace")
def _search_via_api(
topic: str, limit: int, api_key: str, secret_key: str
) -> List[Dict[str, Any]]:
"""通过百度搜索 API 进行搜索。"""
items = []
try:
encoded = urllib.parse.quote(topic)
url = f"https://api.baidu.com/search/v1?q={encoded}&rn={limit}&key={api_key}"
req = urllib.request.Request(url, headers={"User-Agent": _UA_POOL[0]})
with urllib.request.urlopen(req, timeout=15) as response:
data = json.loads(response.read().decode("utf-8"))
for r in data.get("result", []):
items.append({
"title": _clean_html(r.get("title", "")),
"snippet": _clean_html(r.get("abstract", "") or r.get("description", "")),
"url": r.get("url", ""),
"source_domain": _extract_domain(r.get("url", "")),
"date": r.get("date"),
"date_confidence": "med" if r.get("date") else "low",
"source_tag": "百度API",
})
except Exception as e:
sys.stderr.write(f"[百度] API 搜索失败: {e}\n")
return items
def _search_via_public(topic: str, limit: int) -> List[Dict[str, Any]]:
"""通过百度公开搜索(HTML 解析)。"""
items: List[Dict[str, Any]] = []
try:
encoded = urllib.parse.quote(topic)
now_ts = int(time.time())
ago_ts = now_ts - 30 * 86400
url = (
f"https://www.baidu.com/s?wd={encoded}&rn={min(limit, 20)}"
f"&gpc=stf%3D{ago_ts}%2C{now_ts}%7Cstftype%3D1"
)
headers = _build_headers("https://www.baidu.com/")
html = _fetch(url, headers)
if _is_antibot_page(html):
sys.stderr.write(
"[百度] 公开搜索被安全验证拦截,已自动降级到 Bing 兜底。"
"建议配置 BAIDU_API_KEY/BAIDU_SECRET_KEY 以获得稳定结果。\n"
)
return []
container_blocks = re.findall(
r'<div[^>]*class="[^"]*result[^"]*c-container[^"]*"[^>]*>([\s\S]*?)</div>\s*(?=<div[^>]*class="[^"]*result|<div[^>]*id="content_bottom)',
html,
)
for block in container_blocks[:limit]:
title_match = re.search(
r'<h3[^>]*>\s*<a[^>]*href="([^"]+)"[^>]*>(.*?)</a>',
block,
re.S,
)
if not title_match:
continue
href = title_match.group(1)
title = _clean_html(title_match.group(2))
snippet = ""
snip_match = re.search(
r'<span[^>]*class="[^"]*content-right_[^"]*"[^>]*>([\s\S]*?)</span>',
block,
)
if not snip_match:
snip_match = re.search(
r'<span[^>]*class="[^"]*c-font-normal[^"]*"[^>]*>([\s\S]*?)</span>',
block,
)
if snip_match:
snippet = _clean_html(snip_match.group(1))
items.append({
"title": title,
"snippet": snippet,
"url": href,
"source_domain": _extract_domain(href),
"date": None,
"date_confidence": "low",
"source_tag": "百度",
})
if not items:
results = re.findall(
r'<h3[^>]*class="[^"]*t[^"]*"[^>]*>\s*<a[^>]*href="([^"]*)"[^>]*>(.*?)</a>',
html,
re.S,
)
snippets = re.findall(
r'<span class="content-right_8Zs40">(.*?)</span>', html, re.S
)
for idx, (href, title_html) in enumerate(results[:limit]):
title = _clean_html(title_html)
snippet = _clean_html(snippets[idx]) if idx < len(snippets) else ""
items.append({
"title": title,
"snippet": snippet,
"url": href,
"source_domain": _extract_domain(href),
"date": None,
"date_confidence": "low",
"source_tag": "百度",
})
except Exception as e:
sys.stderr.write(f"[百度] 公开搜索失败: {e}\n")
return items
def _search_via_bing(topic: str, limit: int) -> List[Dict[str, Any]]:
"""Bing 国内版兜底搜索,避免百度完全失效时返回 0 条。"""
items: List[Dict[str, Any]] = []
try:
encoded = urllib.parse.quote(topic)
url = f"https://cn.bing.com/search?q={encoded}&setmkt=zh-CN&ensearch=0"
headers = _build_headers("https://cn.bing.com/")
html = _fetch(url, headers)
blocks = re.findall(
r'<li class="b_algo"[^>]*>([\s\S]*?)</li>',
html,
)
for block in blocks[:limit]:
title_match = re.search(
r'<h2[^>]*>\s*<a[^>]*href="([^"]+)"[^>]*>(.*?)</a>\s*</h2>',
block,
re.S,
)
if not title_match:
continue
href = title_match.group(1)
title = _clean_html(title_match.group(2))
snippet = ""
snip_match = re.search(
r'<p[^>]*class="b_lineclamp\d*\s*b_algoSlug"[^>]*>([\s\S]*?)</p>',
block,
) or re.search(r"<p[^>]*>([\s\S]*?)</p>", block)
if snip_match:
snippet = _clean_html(snip_match.group(1))
items.append({
"title": title,
"snippet": snippet,
"url": href,
"source_domain": _extract_domain(href),
"date": None,
"date_confidence": "low",
"source_tag": "Bing兜底",
})
except Exception as e:
sys.stderr.write(f"[百度] Bing 兜底搜索失败: {e}\n")
return items
def _clean_html(text: str) -> str:
"""清除 HTML 标签。"""
return re.sub(r"<[^>]+>", "", text or "").strip()
def _extract_domain(url: str) -> str:
"""从 URL 中提取域名。"""
try:
from urllib.parse import urlparse
return urlparse(url).netloc
except Exception:
return ""
"""B站搜索模块 - 搜索哔哩哔哩视频内容。
Author: Jesse (https://github.com/Jesseovo)
支持两种模式(自动切换):
1. B站公开搜索 API(无需 API Key)
2. MediaCrawler 浏览器爬虫(备用方案)
"""
import json
import re
import sys
import urllib.parse
import urllib.request
from typing import Any, Dict, List, Optional
from . import relevance
_UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
def search_bilibili(
topic: str,
from_date: str,
to_date: str,
depth: str = "default",
) -> List[Dict[str, Any]]:
"""搜索B站视频。
Args:
topic: 搜索关键词
from_date: 起始日期 YYYY-MM-DD
to_date: 结束日期 YYYY-MM-DD
depth: 搜索深度 quick/default/deep
Returns:
B站视频列表
"""
limit_map = {"quick": 10, "default": 20, "deep": 40}
limit = limit_map.get(depth, 20)
pages = 1 if depth == "quick" else (2 if depth == "default" else 3)
items: List[Dict[str, Any]] = []
for page_num in range(1, pages + 1):
try:
page_items = _search_page(topic, page_num)
items.extend(page_items)
except Exception as e:
sys.stderr.write(f"[B站] 搜索第 {page_num} 页失败: {e}\n")
break
if not items:
try:
from . import crawler_bridge
if crawler_bridge.is_playwright_available():
sys.stderr.write("[B站] API 无结果,尝试 MediaCrawler 爬虫模式...\n")
items = crawler_bridge.crawl_bilibili(topic, limit)
if items:
sys.stderr.write(f"[B站] 爬虫模式获取 {len(items)} 条结果\n")
except Exception as e:
sys.stderr.write(f"[B站] 爬虫模式失败: {e}\n")
scored = []
for i, item in enumerate(items):
title = _clean_html(item.get("title", ""))
rel = relevance.token_overlap_relevance(topic, title)
item["id"] = f"BL{i+1}"
item["title"] = title
item["relevance"] = rel
item["why_relevant"] = f"B站视频:{title[:50]}"
scored.append(item)
scored.sort(key=lambda x: x.get("relevance", 0), reverse=True)
return scored[:limit]
def _search_page(topic: str, page: int = 1) -> List[Dict[str, Any]]:
"""搜索B站单页结果。"""
encoded = urllib.parse.quote(topic)
url = (
f"https://api.bilibili.com/x/web-interface/search/type"
f"?search_type=video&keyword={encoded}&page={page}&page_size=20"
f"&order=totalrank"
)
headers = {
"User-Agent": _UA,
"Referer": "https://search.bilibili.com/",
}
req = urllib.request.Request(url, headers=headers)
with urllib.request.urlopen(req, timeout=15) as response:
data = json.loads(response.read().decode("utf-8"))
items = []
results = data.get("data", {}).get("result", [])
if not results:
return items
for r in results:
items.append(_parse_video(r))
return items
def _parse_video(v: dict) -> Dict[str, Any]:
"""解析B站视频搜索结果。"""
bvid = v.get("bvid", "")
pubdate = v.get("pubdate", 0)
date_str = None
if pubdate:
try:
from datetime import datetime
date_str = datetime.fromtimestamp(pubdate).strftime("%Y-%m-%d")
except Exception:
pass
return {
"title": v.get("title", ""),
"url": f"https://www.bilibili.com/video/{bvid}" if bvid else v.get("arcurl", ""),
"bvid": bvid,
"channel_name": v.get("author", ""),
"author_mid": v.get("mid", ""),
"date": date_str,
"duration": v.get("duration", ""),
"description": v.get("description", ""),
"engagement": {
"views": v.get("play", 0),
"danmaku": v.get("danmaku", 0),
"comments": v.get("review", 0) or v.get("comment", 0),
"favorites": v.get("favorites", 0),
"likes": v.get("like", 0),
},
}
def _clean_html(text: str) -> str:
"""清除搜索结果中的 HTML 高亮标签。"""
text = re.sub(r"<[^>]+>", "", text)
return text.strip()
"""Caching utilities for last30days skill.
Author: Jesse (https://github.com/Jesseovo)
"""
import hashlib
import json
import os
import tempfile
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Optional
CACHE_DIR = Path.home() / ".cache" / "last30days-cn"
DEFAULT_TTL_HOURS = 24
MODEL_CACHE_TTL_DAYS = 7
MODEL_CACHE_FILE = CACHE_DIR / "model_selection.json"
def ensure_cache_dir():
"""Ensure cache directory exists. Supports env override and sandbox fallback."""
global CACHE_DIR, MODEL_CACHE_FILE
env_dir = os.environ.get("LAST30DAYS_CACHE_DIR")
if env_dir:
CACHE_DIR = Path(env_dir)
MODEL_CACHE_FILE = CACHE_DIR / "model_selection.json"
try:
CACHE_DIR.mkdir(parents=True, exist_ok=True)
except PermissionError:
CACHE_DIR = Path(tempfile.gettempdir()) / "last30days-cn" / "cache"
MODEL_CACHE_FILE = CACHE_DIR / "model_selection.json"
CACHE_DIR.mkdir(parents=True, exist_ok=True)
def get_cache_key(topic: str, from_date: str, to_date: str, sources: str) -> str:
"""Generate a cache key from query parameters."""
key_data = f"{topic}|{from_date}|{to_date}|{sources}"
return hashlib.sha256(key_data.encode()).hexdigest()[:16]
def get_cache_path(cache_key: str) -> Path:
"""Get path to cache file."""
return CACHE_DIR / f"{cache_key}.json"
def is_cache_valid(cache_path: Path, ttl_hours: int = DEFAULT_TTL_HOURS) -> bool:
"""Check if cache file exists and is within TTL."""
if not cache_path.exists():
return False
try:
stat = cache_path.stat()
mtime = datetime.fromtimestamp(stat.st_mtime, tz=timezone.utc)
now = datetime.now(timezone.utc)
age_hours = (now - mtime).total_seconds() / 3600
return age_hours < ttl_hours
except OSError:
return False
def load_cache(cache_key: str, ttl_hours: int = DEFAULT_TTL_HOURS) -> Optional[dict]:
"""Load data from cache if valid."""
cache_path = get_cache_path(cache_key)
if not is_cache_valid(cache_path, ttl_hours):
return None
try:
with open(cache_path, 'r') as f:
return json.load(f)
except (json.JSONDecodeError, OSError):
return None
def get_cache_age_hours(cache_path: Path) -> Optional[float]:
"""Get age of cache file in hours."""
if not cache_path.exists():
return None
try:
stat = cache_path.stat()
mtime = datetime.fromtimestamp(stat.st_mtime, tz=timezone.utc)
now = datetime.now(timezone.utc)
return (now - mtime).total_seconds() / 3600
except OSError:
return None
def load_cache_with_age(cache_key: str, ttl_hours: int = DEFAULT_TTL_HOURS) -> tuple:
"""Load data from cache with age info.
Returns:
Tuple of (data, age_hours) or (None, None) if invalid
"""
cache_path = get_cache_path(cache_key)
if not is_cache_valid(cache_path, ttl_hours):
return None, None
age = get_cache_age_hours(cache_path)
try:
with open(cache_path, 'r') as f:
return json.load(f), age
except (json.JSONDecodeError, OSError):
return None, None
def save_cache(cache_key: str, data: dict):
"""Save data to cache."""
ensure_cache_dir()
cache_path = get_cache_path(cache_key)
try:
with open(cache_path, 'w') as f:
json.dump(data, f)
except OSError:
pass # Silently fail on cache write errors
def clear_cache():
"""Clear all cache files."""
if CACHE_DIR.exists():
for f in CACHE_DIR.glob("*.json"):
try:
f.unlink()
except OSError:
pass
# Model selection cache (longer TTL) — MODEL_CACHE_FILE is set at module level
# and updated by ensure_cache_dir() if env override or fallback is needed.
def load_model_cache() -> dict:
"""Load model selection cache."""
if not is_cache_valid(MODEL_CACHE_FILE, MODEL_CACHE_TTL_DAYS * 24):
return {}
try:
with open(MODEL_CACHE_FILE, 'r') as f:
return json.load(f)
except (json.JSONDecodeError, OSError):
return {}
def save_model_cache(data: dict):
"""Save model selection cache."""
ensure_cache_dir()
try:
with open(MODEL_CACHE_FILE, 'w') as f:
json.dump(data, f)
except OSError:
pass
def get_cached_model(provider: str) -> Optional[str]:
"""Get cached model selection for a provider."""
cache = load_model_cache()
return cache.get(provider)
def set_cached_model(provider: str, model: str):
"""Cache model selection for a provider."""
cache = load_model_cache()
cache[provider] = model
cache['updated_at'] = datetime.now(timezone.utc).isoformat()
save_model_cache(cache)
"""跨源聚类:把同一事件在多个平台上的条目聚合成一条热点。
复用 dedupe 的跨源相似度(``_get_cross_source_text`` + ``_hybrid_similarity``),
用并查集(union-find)求连通分量;仅保留覆盖 ≥2 个不同平台类型的簇,避免把单一
平台的近似条目当成"跨平台热点"。
Author: Jesse (https://github.com/Jesseovo)
"""
from typing import Any, Dict, List
from . import dedupe, schema
# 平台类型 → 中文标签
_SOURCE_LABEL = {
schema.WeiboItem: "微博",
schema.XiaohongshuItem: "小红书",
schema.BilibiliItem: "B站",
schema.ZhihuItem: "知乎",
schema.DouyinItem: "抖音",
schema.WechatItem: "微信",
schema.BaiduItem: "百度",
schema.ToutiaoItem: "头条",
}
def _title_of(item) -> str:
"""取代表项的展示文本。"""
for attr in ("title", "text"):
value = getattr(item, attr, None)
if value:
return value.strip()[:80]
return getattr(item, "url", "") or ""
def _rank_key(item):
"""簇代表项排序键:先看综合得分,再看相关性。"""
return (getattr(item, "score", 0) or 0, getattr(item, "relevance", 0) or 0)
def build_clusters(*source_lists: List[Any], threshold: float = 0.40) -> List[Dict[str, Any]]:
"""把多个平台的结果列表聚成跨平台簇。
Args:
*source_lists: 各平台已去重的条目列表。
threshold: 跨源相似度阈值(与 dedupe.cross_source_link 保持一致)。
Returns:
簇列表,每个簇为 dict:
``{representative_id, representative_title, representative_url,
member_ids, sources, size}``,按 size 降序。
"""
all_items: List[Any] = []
for source_list in source_lists:
all_items.extend(source_list)
n = len(all_items)
if n <= 1:
return []
texts = [dedupe._get_cross_source_text(item) for item in all_items]
parent = list(range(n))
def find(x: int) -> int:
while parent[x] != x:
parent[x] = parent[parent[x]] # 路径压缩
x = parent[x]
return x
def union(a: int, b: int) -> None:
ra, rb = find(a), find(b)
if ra != rb:
parent[rb] = ra
# 只在不同平台类型之间连边(同源近似条目已由各自 dedupe 处理)
for i in range(n):
for j in range(i + 1, n):
if type(all_items[i]) is type(all_items[j]):
continue
if dedupe._hybrid_similarity(texts[i], texts[j]) >= threshold:
union(i, j)
groups: Dict[int, List[int]] = {}
for idx in range(n):
groups.setdefault(find(idx), []).append(idx)
clusters: List[Dict[str, Any]] = []
for members in groups.values():
if len(members) < 2:
continue
member_items = [all_items[m] for m in members]
types = {type(it) for it in member_items}
if len(types) < 2:
continue # 必须跨 ≥2 个平台类型
rep = max(member_items, key=_rank_key)
sources = sorted({_SOURCE_LABEL.get(type(it), "?") for it in member_items})
clusters.append({
"representative_id": rep.id,
"representative_title": _title_of(rep),
"representative_url": getattr(rep, "url", "") or "",
"member_ids": [it.id for it in member_items],
"sources": sources,
"size": len(member_items),
})
clusters.sort(key=lambda c: c["size"], reverse=True)
return clusters
"""MediaCrawler 爬虫桥接模块 — 基于 Playwright 浏览器自动化的数据采集。
灵感来源: https://github.com/NanmiCoder/MediaCrawler
技术原理: 利用 Playwright 控制浏览器,通过保留登录态的上下文获取平台数据,
无需逆向复杂加密算法,无需付费 API Key。
v2.1 改动:
- 抽取 `_launch_browser_context` 公共函数,统一 locale/viewport/UA
- 小红书/抖音改用 `page.on("response")` XHR 拦截,避开 Virtual DOM 渲染
- `page.goto` 统一使用 `wait_until="domcontentloaded"`,配合条件轮询替代固定 sleep
⚠️ 免责声明:
本模块仅供学习和研究目的。使用者必须遵守相关法律法规及各平台的服务条款。
禁止用于商业用途、大规模数据采集或任何非法活动。
Author: Jesse (https://github.com/Jesseovo)
"""
import json
import os
import re
import sys
from contextlib import contextmanager
from pathlib import Path
from typing import Any, Dict, List, Optional
from datetime import datetime
COOKIE_DIR = Path.home() / ".config" / "last30days-cn" / "browser_cookies"
_playwright_available: Optional[bool] = None
_DESKTOP_UA = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
)
_MOBILE_UA = (
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) "
"AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1"
)
def is_playwright_available() -> bool:
"""检查 Playwright 是否已安装并可用。"""
global _playwright_available
if _playwright_available is not None:
return _playwright_available
try:
from playwright.sync_api import sync_playwright # noqa: F401
_playwright_available = True
except ImportError:
_playwright_available = False
return _playwright_available
def _ensure_cookie_dir():
COOKIE_DIR.mkdir(parents=True, exist_ok=True)
def _get_cookie_path(platform: str) -> Path:
_ensure_cookie_dir()
return COOKIE_DIR / f"{platform}_cookies.json"
def save_cookies(platform: str, cookies: list):
path = _get_cookie_path(platform)
path.write_text(json.dumps(cookies, ensure_ascii=False, indent=2), encoding="utf-8")
def load_cookies(platform: str) -> Optional[list]:
path = _get_cookie_path(platform)
if path.exists():
try:
return json.loads(path.read_text(encoding="utf-8"))
except Exception:
return None
return None
def _clean_html(text: str) -> str:
if not text:
return ""
text = re.sub(r"<[^>]+>", "", str(text))
return re.sub(r"\s+", " ", text).strip()
@contextmanager
def _launch_browser_context(platform: str, mobile: bool = False, headless: bool = True):
"""统一构造 Playwright 浏览器上下文,自动加载并回写 cookies。
Yields:
(browser, context, page) 三元组
"""
from playwright.sync_api import sync_playwright
ua = _MOBILE_UA if mobile else _DESKTOP_UA
viewport = {"width": 390, "height": 844} if mobile else {"width": 1280, "height": 800}
with sync_playwright() as p:
browser = p.chromium.launch(headless=headless)
try:
context = browser.new_context(
user_agent=ua,
locale="zh-CN",
timezone_id="Asia/Shanghai",
viewport=viewport,
)
cookies = load_cookies(platform)
if cookies:
try:
context.add_cookies(cookies)
except Exception as e:
sys.stderr.write(f"[爬虫-{platform}] 加载 Cookie 失败: {e}\n")
page = context.new_page()
try:
yield browser, context, page
finally:
try:
save_cookies(platform, context.cookies())
except Exception:
pass
finally:
try:
browser.close()
except Exception:
pass
def _wait_for(page, predicate, timeout_ms: int = 8000, interval_ms: int = 500) -> bool:
"""轮询等待条件满足,返回是否在超时前命中。"""
elapsed = 0
while elapsed < timeout_ms:
try:
if predicate():
return True
except Exception:
pass
page.wait_for_timeout(interval_ms)
elapsed += interval_ms
return False
def crawl_weibo(topic: str, limit: int = 20) -> List[Dict[str, Any]]:
"""通过浏览器自动化爬取微博搜索结果。"""
if not is_playwright_available():
return []
items: List[Dict[str, Any]] = []
try:
with _launch_browser_context("weibo", mobile=True) as (browser, context, page):
search_url = (
f"https://m.weibo.cn/api/container/getIndex?"
f"containerid=100103type%3D1%26q%3D{topic}&page_type=searchall"
)
page.goto(
f"https://m.weibo.cn/search?containerid=100103type%3D1%26q%3D{topic}",
wait_until="domcontentloaded",
timeout=20000,
)
page.wait_for_timeout(1500)
response = page.evaluate(
"""
async (url) => {
const resp = await fetch(url);
return await resp.json();
}
""",
search_url,
)
if isinstance(response, dict):
cards = response.get("data", {}).get("cards", [])
for card in cards[:limit]:
if card.get("card_type") == 9:
mblog = card.get("mblog", {})
if mblog:
items.append(_parse_weibo_mblog(mblog))
elif card.get("card_type") == 11:
for group in card.get("card_group", []):
mblog = group.get("mblog", {})
if mblog:
items.append(_parse_weibo_mblog(mblog))
except Exception as e:
sys.stderr.write(f"[爬虫-微博] 浏览器爬取失败: {e}\n")
return items
def _parse_weibo_mblog(mblog: dict) -> Dict[str, Any]:
text = _clean_html(mblog.get("text", ""))
user = mblog.get("user", {})
mid = mblog.get("mid", "") or mblog.get("id", "")
return {
"text": text,
"url": f"https://weibo.com/{user.get('id', '')}/{mid}",
"author_handle": user.get("screen_name", ""),
"author_id": str(user.get("id", "")),
"date": _parse_relative_date(mblog.get("created_at", "")),
"engagement": {
"reposts": mblog.get("reposts_count", 0),
"comments": mblog.get("comments_count", 0),
"likes": mblog.get("attitudes_count", 0),
},
"source": "crawler",
}
def crawl_xiaohongshu(topic: str, limit: int = 20) -> List[Dict[str, Any]]:
"""通过浏览器自动化爬取小红书搜索结果。
v2.1:参考 MediaCrawler 思路,优先拦截 XHR `/api/sns/web/v1/search/notes`
避开 Virtual DOM。失败时回退到 DOM 选择器(命中率较低,仅作兜底)。
"""
if not is_playwright_available():
return []
items: List[Dict[str, Any]] = []
try:
with _launch_browser_context("xiaohongshu") as (browser, context, page):
captured: Dict[str, Any] = {"payload": None}
def _on_response(resp):
try:
url = resp.url
if "/api/sns/web/v1/search/notes" in url and resp.status == 200:
captured["payload"] = resp.json()
except Exception:
pass
page.on("response", _on_response)
search_url = (
f"https://www.xiaohongshu.com/search_result?"
f"keyword={topic}&source=web_search_result_notes"
)
try:
page.goto(search_url, wait_until="domcontentloaded", timeout=30000)
except Exception as e:
sys.stderr.write(f"[爬虫-小红书] 页面加载失败: {e}\n")
for _ in range(5):
if captured["payload"]:
break
try:
page.mouse.wheel(0, 2000)
except Exception:
pass
page.wait_for_timeout(1500)
payload = captured["payload"]
if isinstance(payload, dict):
data = payload.get("data") or {}
raw_items = data.get("items") or []
for raw in raw_items[:limit]:
note_card = raw.get("note_card") or {}
if not note_card:
continue
note_id = raw.get("id") or note_card.get("note_id", "")
user = note_card.get("user") or {}
interact = note_card.get("interact_info") or {}
items.append({
"title": note_card.get("display_title") or note_card.get("title", ""),
"desc": note_card.get("desc", ""),
"url": f"https://www.xiaohongshu.com/explore/{note_id}" if note_id else "",
"author_name": user.get("nickname") or user.get("nick_name", ""),
"author_id": user.get("user_id") or user.get("userid", ""),
"date": None,
"engagement": {
"likes": _parse_count(str(interact.get("liked_count", "0"))),
"collects": _parse_count(str(interact.get("collected_count", "0"))),
"comments": _parse_count(str(interact.get("comment_count", "0"))),
"shares": _parse_count(str(interact.get("share_count", "0"))),
},
"hashtags": [],
"images": [
img.get("url_default") or img.get("url", "")
for img in (note_card.get("image_list") or [])
if isinstance(img, dict)
],
"source": "crawler-xhr",
})
elif not payload:
sys.stderr.write(
"[爬虫-小红书] 未捕获搜索 XHR,可能需要重新登录、验证码通过,或平台接口已调整;"
"将继续尝试 DOM/站内搜索兜底。\n"
)
if not items:
note_elements = page.query_selector_all(
"section.note-item, div[class*='note-item'], a[class*='cover'], a[href*='/explore/']"
)
for elem in note_elements[:limit]:
try:
title_el = elem.query_selector("span[class*='title'], div[class*='title']")
title = title_el.inner_text() if title_el else ""
link = elem.get_attribute("href") or ""
if link and not link.startswith("http"):
link = f"https://www.xiaohongshu.com{link}"
author_el = elem.query_selector("span[class*='name'], div[class*='author']")
author = author_el.inner_text() if author_el else ""
likes_el = elem.query_selector("span[class*='like'], span[class*='count']")
likes_text = likes_el.inner_text() if likes_el else "0"
likes = _parse_count(likes_text)
items.append({
"title": title,
"desc": "",
"url": link,
"author_name": author,
"author_id": "",
"date": None,
"engagement": {
"likes": likes,
"collects": 0,
"comments": 0,
"shares": 0,
},
"hashtags": [],
"images": [],
"source": "crawler-dom",
})
except Exception:
continue
if not items:
sys.stderr.write(
"[爬虫-小红书] Playwright 未解析到结果;这通常是登录态失效、反爬验证或页面结构变更导致。\n"
)
except Exception as e:
sys.stderr.write(f"[爬虫-小红书] 浏览器爬取失败: {e}\n")
return items
def crawl_douyin(topic: str, limit: int = 20) -> List[Dict[str, Any]]:
"""通过浏览器自动化爬取抖音搜索结果。
v2.1:优先拦截 XHR `/aweme/v1/web/search/item/`,DOM 解析作为兜底。
"""
if not is_playwright_available():
return []
items: List[Dict[str, Any]] = []
try:
with _launch_browser_context("douyin") as (browser, context, page):
captured: Dict[str, Any] = {"payload": None}
def _on_response(resp):
try:
url = resp.url
if "/aweme/v1/web/search/item" in url and resp.status == 200:
data = resp.json()
if isinstance(data, dict) and data.get("data"):
captured["payload"] = data
except Exception:
pass
page.on("response", _on_response)
search_url = f"https://www.douyin.com/search/{topic}?type=video"
try:
page.goto(search_url, wait_until="domcontentloaded", timeout=30000)
except Exception as e:
sys.stderr.write(f"[爬虫-抖音] 页面加载失败: {e}\n")
for _ in range(5):
if captured["payload"]:
break
try:
page.mouse.wheel(0, 2000)
except Exception:
pass
page.wait_for_timeout(1500)
payload = captured["payload"]
if isinstance(payload, dict):
for entry in (payload.get("data") or [])[:limit]:
aweme = entry.get("aweme_info") or entry
if not isinstance(aweme, dict):
continue
aweme_id = aweme.get("aweme_id", "")
author = aweme.get("author") or {}
stats = aweme.get("statistics") or {}
items.append({
"text": aweme.get("desc", ""),
"url": f"https://www.douyin.com/video/{aweme_id}" if aweme_id else "",
"author_name": author.get("nickname", ""),
"author_id": author.get("uid", ""),
"date": None,
"engagement": {
"views": stats.get("play_count", 0),
"likes": stats.get("digg_count", 0),
"comments": stats.get("comment_count", 0),
"shares": stats.get("share_count", 0),
},
"hashtags": [],
"duration": (aweme.get("duration") or 0) // 1000,
"source": "crawler-xhr",
})
if not items:
video_elements = page.query_selector_all(
"div[class*='video-card'], li[class*='search-result']"
)
for elem in video_elements[:limit]:
try:
title_el = elem.query_selector(
"a[class*='title'], span[class*='title'], p[class*='desc']"
)
title = title_el.inner_text() if title_el else ""
link_el = elem.query_selector("a[href*='/video/']")
link = ""
if link_el:
href = link_el.get_attribute("href") or ""
if href.startswith("/"):
link = f"https://www.douyin.com{href}"
else:
link = href
author_el = elem.query_selector(
"span[class*='author'], span[class*='nickname']"
)
author = author_el.inner_text() if author_el else ""
likes_el = elem.query_selector("span[class*='like'], span[class*='digg']")
likes = _parse_count(likes_el.inner_text() if likes_el else "0")
items.append({
"text": title,
"url": link,
"author_name": author,
"author_id": "",
"date": None,
"engagement": {
"views": 0,
"likes": likes,
"comments": 0,
"shares": 0,
},
"hashtags": [],
"duration": 0,
"source": "crawler-dom",
})
except Exception:
continue
except Exception as e:
sys.stderr.write(f"[爬虫-抖音] 浏览器爬取失败: {e}\n")
return items
def crawl_bilibili(topic: str, limit: int = 20) -> List[Dict[str, Any]]:
"""通过浏览器自动化爬取B站搜索结果。"""
if not is_playwright_available():
return []
items: List[Dict[str, Any]] = []
try:
with _launch_browser_context("bilibili") as (browser, context, page):
search_url = f"https://search.bilibili.com/all?keyword={topic}&order=totalrank"
try:
page.goto(search_url, wait_until="domcontentloaded", timeout=30000)
except Exception as e:
sys.stderr.write(f"[爬虫-B站] 页面加载失败: {e}\n")
_wait_for(
page,
lambda: bool(page.query_selector("div.bili-video-card, div[class*='video-list-item']")),
timeout_ms=8000,
)
video_elements = page.query_selector_all(
"div.bili-video-card, div[class*='video-list-item']"
)
for elem in video_elements[:limit]:
try:
title_el = elem.query_selector("h3[class*='title'], a[class*='title']")
title = _clean_html(title_el.inner_text()) if title_el else ""
link_el = elem.query_selector("a[href*='/video/']")
link = ""
if link_el:
href = link_el.get_attribute("href") or ""
if href.startswith("//"):
link = f"https:{href}"
elif href.startswith("/"):
link = f"https://www.bilibili.com{href}"
else:
link = href
author_el = elem.query_selector(
"span[class*='name'], span.bili-video-card__info--author"
)
author = author_el.inner_text() if author_el else ""
views_el = elem.query_selector("span[class*='play'], span[class*='view']")
views = _parse_count(views_el.inner_text() if views_el else "0")
items.append({
"title": title,
"url": link,
"bvid": "",
"channel_name": author,
"author_mid": "",
"date": None,
"duration": "",
"description": "",
"engagement": {
"views": views,
"danmaku": 0,
"comments": 0,
"favorites": 0,
"likes": 0,
},
"source": "crawler",
})
except Exception:
continue
except Exception as e:
sys.stderr.write(f"[爬虫-B站] 浏览器爬取失败: {e}\n")
return items
def crawl_zhihu(topic: str, limit: int = 20) -> List[Dict[str, Any]]:
"""通过浏览器自动化爬取知乎搜索结果。"""
if not is_playwright_available():
return []
items: List[Dict[str, Any]] = []
try:
with _launch_browser_context("zhihu") as (browser, context, page):
search_url = f"https://www.zhihu.com/search?type=content&q={topic}"
try:
page.goto(search_url, wait_until="domcontentloaded", timeout=30000)
except Exception as e:
sys.stderr.write(f"[爬虫-知乎] 页面加载失败: {e}\n")
_wait_for(
page,
lambda: bool(page.query_selector("div.SearchResult-Card, div[class*='List-item']")),
timeout_ms=8000,
)
result_elements = page.query_selector_all(
"div.SearchResult-Card, div[class*='List-item']"
)
for elem in result_elements[:limit]:
try:
title_el = elem.query_selector(
"h2[class*='ContentItem-title'], span[class*='Highlight']"
)
title = _clean_html(title_el.inner_text()) if title_el else ""
link_el = elem.query_selector("a[href*='/question/'], a[href*='/p/']")
link = ""
if link_el:
href = link_el.get_attribute("href") or ""
if href.startswith("/"):
link = f"https://www.zhihu.com{href}"
else:
link = href
excerpt_el = elem.query_selector(
"span[class*='RichText'], div[class*='RichText']"
)
excerpt = _clean_html(excerpt_el.inner_text()[:300]) if excerpt_el else ""
author_el = elem.query_selector(
"span[class*='AuthorInfo'] a, a[class*='UserLink']"
)
author = author_el.inner_text() if author_el else ""
voteup_el = elem.query_selector(
"button[class*='VoteButton'] span, span[class*='vote']"
)
voteups = _parse_count(voteup_el.inner_text() if voteup_el else "0")
items.append({
"title": title,
"excerpt": excerpt,
"url": link,
"author": author,
"date": None,
"content_type": "search_result",
"engagement": {
"voteups": voteups,
"comments": 0,
"collects": 0,
},
"source": "crawler",
})
except Exception:
continue
if not items:
sys.stderr.write(
"[爬虫-知乎] Playwright 未解析到结果;这通常是登录态失效、反爬验证或页面结构变更导致。\n"
)
except Exception as e:
sys.stderr.write(f"[爬虫-知乎] 浏览器爬取失败: {e}\n")
return items
def _parse_count(text: str) -> int:
"""解析数量文本,支持 '1.2万' / '1.2w' / '12k' 等格式。"""
if not text:
return 0
text = str(text).strip().replace(",", "")
try:
if "万" in text or text.lower().endswith("w"):
num = float(re.sub(r"[万wW]", "", text))
return int(num * 10000)
elif "亿" in text:
num = float(text.replace("亿", ""))
return int(num * 100000000)
elif text.lower().endswith("k"):
num = float(text[:-1])
return int(num * 1000)
return int(float(re.sub(r"[^\d.]", "", text or "0")))
except (ValueError, TypeError):
return 0
def _parse_relative_date(date_str: str) -> Optional[str]:
"""解析微博的相对时间格式。"""
if not date_str:
return None
from datetime import timedelta
now = datetime.now()
try:
dt = datetime.strptime(date_str, "%a %b %d %H:%M:%S %z %Y")
return dt.strftime("%Y-%m-%d")
except ValueError:
pass
if "分钟前" in date_str:
m = re.search(r"(\d+)", date_str)
if m:
return (now - timedelta(minutes=int(m.group(1)))).strftime("%Y-%m-%d")
if "小时前" in date_str:
m = re.search(r"(\d+)", date_str)
if m:
return (now - timedelta(hours=int(m.group(1)))).strftime("%Y-%m-%d")
if "昨天" in date_str:
return (now - timedelta(days=1)).strftime("%Y-%m-%d")
m = re.match(r"(\d{2})-(\d{2})", date_str)
if m:
return f"{now.year}-{m.group(1)}-{m.group(2)}"
return None
def get_crawler_status() -> Dict[str, Any]:
"""获取爬虫引擎的状态信息。"""
pw_available = is_playwright_available()
cached_platforms = []
if COOKIE_DIR.exists():
for f in COOKIE_DIR.glob("*_cookies.json"):
platform = f.stem.replace("_cookies", "")
cached_platforms.append(platform)
return {
"playwright_available": pw_available,
"cached_logins": cached_platforms,
"cookie_dir": str(COOKIE_DIR),
}
"""Date utilities for last30days skill.
Author: Jesse (https://github.com/Jesseovo)
"""
from datetime import datetime, timedelta, timezone
from typing import Optional, Tuple
CST = timezone(timedelta(hours=8))
def get_date_range(days: int = 30, as_of: Optional[str] = None) -> Tuple[str, str]:
"""Get the date range for the N days ending at ``as_of`` (or today).
Args:
days: 回溯天数。
as_of: 终点日期 YYYY-MM-DD(历史回溯);为空时以今天为终点。
Returns:
Tuple of (from_date, to_date) as YYYY-MM-DD strings
Raises:
ValueError: as_of 无法解析为有效日期时。
"""
if as_of:
parsed = parse_date(as_of)
if parsed is None:
raise ValueError(f"无效的 --as-of 日期: {as_of!r}(应为 YYYY-MM-DD)")
end_date = parsed.date()
else:
end_date = datetime.now(CST).date()
from_date = end_date - timedelta(days=days)
return from_date.isoformat(), end_date.isoformat()
def parse_date(date_str: Optional[str]) -> Optional[datetime]:
"""Parse a date string in various formats.
Supports: YYYY-MM-DD, ISO 8601, Unix timestamp
"""
if not date_str:
return None
# 尝试 Unix 时间戳
try:
ts = float(date_str)
return datetime.fromtimestamp(ts, tz=timezone.utc)
except (ValueError, TypeError):
pass
# Try ISO formats
formats = [
"%Y-%m-%d",
"%Y-%m-%dT%H:%M:%S",
"%Y-%m-%dT%H:%M:%SZ",
"%Y-%m-%dT%H:%M:%S%z",
"%Y-%m-%dT%H:%M:%S.%f%z",
]
for fmt in formats:
try:
return datetime.strptime(date_str, fmt).replace(tzinfo=timezone.utc)
except ValueError:
continue
return None
def timestamp_to_date(ts: Optional[float]) -> Optional[str]:
"""Convert Unix timestamp to YYYY-MM-DD string."""
if ts is None:
return None
try:
dt = datetime.fromtimestamp(ts, tz=timezone.utc)
return dt.date().isoformat()
except (ValueError, TypeError, OSError):
return None
def get_date_confidence(date_str: Optional[str], from_date: str, to_date: str) -> str:
"""Determine confidence level for a date.
Args:
date_str: The date to check (YYYY-MM-DD or None)
from_date: Start of valid range (YYYY-MM-DD)
to_date: End of valid range (YYYY-MM-DD)
Returns:
'high', 'med', or 'low'
"""
if not date_str:
return 'low'
try:
dt = datetime.strptime(date_str, "%Y-%m-%d").date()
start = datetime.strptime(from_date, "%Y-%m-%d").date()
end = datetime.strptime(to_date, "%Y-%m-%d").date()
if start <= dt <= end:
return 'high'
elif dt < start:
# Older than range
return 'low'
else:
# Future date (suspicious)
return 'low'
except ValueError:
return 'low'
def days_ago(date_str: Optional[str]) -> Optional[int]:
"""Calculate how many days ago a date is.
Returns None if date is invalid or missing.
"""
if not date_str:
return None
try:
dt = datetime.strptime(date_str, "%Y-%m-%d").date()
today = datetime.now(CST).date()
delta = today - dt
return delta.days
except ValueError:
return None
def recency_score(date_str: Optional[str], max_days: int = 30) -> int:
"""Calculate recency score (0-100).
0 days ago = 100, max_days ago = 0, clamped.
"""
age = days_ago(date_str)
if age is None:
return 0 # Unknown date gets worst score
if age < 0:
return 100 # Future date (treat as today)
if age >= max_days:
return 0
return int(100 * (1 - age / max_days))
"""Entity extraction from Phase 1 search results for supplemental searches.
Author: Jesse (https://github.com/Jesseovo)
"""
import re
from collections import Counter
from typing import Any, Dict, List, Optional
# Accounts that appear too frequently to be useful for targeted search.
GENERIC_HANDLES = frozenset(
{
"人民日报",
"央视新闻",
"新华社",
"人民网",
"环球网",
"中国日报",
"光明日报",
"经济日报",
"解放军报",
"共青团中央",
"中央广播电视总台",
"微博管理员",
"小红书官方",
}
)
def _account_key(name: str) -> str:
n = name.strip().lstrip("@")
return n.casefold() if n.isascii() else n
def _is_generic_account(name: str) -> bool:
if not name:
return True
key = _account_key(name)
return key in {_account_key(g) for g in GENERIC_HANDLES}
def extract_entities(
weibo_items: List[Dict[str, Any]],
xiaohongshu_items: List[Dict[str, Any]],
*,
zhihu_items: Optional[List[Dict[str, Any]]] = None,
max_weibo_users: int = 5,
max_xiaohongshu_topics: int = 3,
max_zhihu_questions: int = 5,
) -> Dict[str, List[str]]:
"""Extract key entities from Phase 1 results for supplemental searches.
Parses Weibo for @用户, 小红书 for #话题# (double-hash topics), Zhihu for 问题标题.
Args:
weibo_items: Raw Weibo item dicts from Phase 1
xiaohongshu_items: Raw 小红书 item dicts from Phase 1
zhihu_items: Raw 知乎 item dicts (optional; used for question titles)
max_weibo_users: Max Weibo users to return
max_xiaohongshu_topics: Max 小红书话题 to return
max_zhihu_questions: Max 知乎问题 strings to return
Returns:
Dict with keys: weibo_users, xiaohongshu_topics, zhihu_questions.
"""
if zhihu_items is None:
zhihu_items = []
users = _extract_weibo_users(weibo_items)
topics = _extract_xiaohongshu_topics(xiaohongshu_items)
questions = _extract_zhihu_questions(zhihu_items)
wu = users[:max_weibo_users]
xt = topics[:max_xiaohongshu_topics]
zq = questions[:max_zhihu_questions]
return {
"weibo_users": wu,
"xiaohongshu_topics": xt,
"zhihu_questions": zq,
}
def _extract_weibo_users(weibo_items: List[Dict[str, Any]]) -> List[str]:
"""Extract and rank @用户 from Weibo results (author + @mentions in text)."""
handle_counts = Counter()
canonical: Dict[str, str] = {}
mention_re = re.compile(r"@([\w\u4e00-\u9fff·]{1,40})")
def bump(raw: str) -> None:
raw = str(raw).strip().lstrip("@")
if not raw or _is_generic_account(raw):
return
nk = _account_key(raw)
if nk not in canonical:
canonical[nk] = raw
handle_counts[nk] += 1
for item in weibo_items:
bump(item.get("author_handle", "") or item.get("author", ""))
text = item.get("text", "") or item.get("title", "") or ""
for m in mention_re.findall(text):
bump(m)
return [canonical[k] for k, _ in handle_counts.most_common()]
def _extract_xiaohongshu_topics(xiaohongshu_items: List[Dict[str, Any]]) -> List[str]:
"""Extract and rank #话题# (double-hash) topics from 小红书 items."""
topic_counts = Counter()
topic_re = re.compile(r"#([^#\n][^#]{1,50}?)#")
for item in xiaohongshu_items:
chunks = []
for field in ("text", "title", "caption_snippet", "description"):
v = item.get(field)
if v:
chunks.append(str(v))
blob = "\n".join(chunks)
for raw in topic_re.findall(blob):
t = raw.strip()
if len(t) >= 2:
topic_counts[t] += 1
for tag in item.get("hashtags") or []:
t = str(tag).strip().lstrip("#")
if len(t) >= 2:
topic_counts[t] += 1
return [t for t, _ in topic_counts.most_common()]
def _extract_zhihu_questions(zhihu_items: List[Dict[str, Any]]) -> List[str]:
"""Extract and rank 知乎问题 titles from Zhihu items."""
q_counts = Counter()
for item in zhihu_items:
q = item.get("question") or item.get("title")
if not q:
continue
q = str(q).strip()
if len(q) >= 4:
q_counts[q] += 1
return [q for q, _ in q_counts.most_common()]
# last30days library modules
# last30days tests