
Rq Earnings Analysis
- 1 installs
- 43 repo stars
- Updated June 23, 2026
- ricequant/ricequant-skills
Generate structured Chinese earnings analysis reports from Ricequant JSON datasets via the earnings-analysis report template and generator script.
About
rq-earnings-analysis packages a Ricequant-oriented earnings report workflow for developers and small research stacks automating equity write-ups. The SKILL content defines a full Chinese-language report skeleton—from executive summary through market expectations, management wording, balance-sheet quality, updated investment thesis, valuation, and risk—filled by scripts such as earnings-analysis/scripts/generate_report.py reading a --data-dir of JSON files. Documented inputs include company_info, industry tiers, historical_financials with revenue and profit lines, ROE history, market cap and PE/PB/dividend yield series, and price_window bars for event-window reaction. Agents use it when turning standardized quantitative extracts into citable research memos during diligence or when validating finance-side assumptions before building trading tools, dashboards, or content products. It is template-plus-generator rather than a live market MCP; practitioners must supply correct Ricequant exports.
- Markdown report template with placeholders: executive summary, financial quality, thesis update, valuation, risks
- Data contract for generate_report.py: company_info, industry, historical_financials, roe_history, valuation factors, pri
- Derives YoY/QoQ metrics, margins, cash conversion, leverage, and post-earnings price reaction from JSON inputs
- Documents dividend_yield bps-to-percent conversion for display
- Bilingual finance workflow aimed at Ricequant / A-share style identifiers (order_book_id, quarters)
Rq Earnings Analysis by the numbers
- 1 all-time installs (skills.sh)
- Ranked #909 of 1,106 Finance & Trading skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 3, 2026 (Skillselion catalog sync)
npx skills add https://github.com/ricequant/ricequant-skills --skill rq-earnings-analysisAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 43 |
| Security audit | 1 / 3 scanners passed |
| Last updated | June 23, 2026 |
| Repository | ricequant/ricequant-skills ↗ |
What it does
Generate structured Chinese earnings analysis reports from Ricequant JSON datasets via the earnings-analysis report template and generator script.
Files
RQ 股票研究 - 财报分析
核心原则
- 所有内容必须遵循三阶段流程:数据采集 -> 报告生成 -> HTML 渲染
assets/template.md是唯一报告模板来源;Python 只做数据归一化、指标计算、占位符填充和结构校验- Python 只输出结构化 facts / tables / signal rules,不在代码里硬写观点性结论
RQData CLI是财务、估值、价格、公告和一致预期的主源;web_search只补充 CLI 无法直接提供的实时外部语境- 最新财报季度必须从真实
financial数据中按info_date <= report-date自动识别,不能硬写季度 - “超预期/低于预期”判断必须明确口径,优先使用财报前一致预期和高度相关的卖方点评
- 研报必须做相关性过滤;行业周报、策略报告不能冒充公司点评
- 缺失数据时必须明确写“无数据 / 未提供”,不能留空
数据源分工
RQData CLI 负责
- 公司信息、行业归属、历史财务、ROE、估值与股息率
- 财报前后价格反应、成交额变化、基准超额收益
- 一致预期、目标价、卖方研报
- 正式公告、业绩说明会、财报原文链接
web_search 负责
- 财报电话会安排、管理层最新外部表态
- 财报后行业或政策语境
RQData CLI未直接提供、但会影响财报解读的实时背景信息
web_search 禁止替代的内容
- 财务数字、估值指标、价格数据
- 正式公告、财报披露日、分红、交易所披露
- 一致预期和结构化卖方预测
web_search 使用规则
详细字段、来源等级、落盘示例和 fallback 规则见 references/web_search.md。
允许补充的内容:
- 管理层动态、财报电话会、IR 活动安排
- 行业景气、政策变化、监管动态
- 财报后几天内的重要公司新闻或权威媒体解读
落盘要求:
- 所有
web_search结果必须先写入web_search_findings.json - 只写结构化记录,不把自然语言笔记直接塞进报告
- 每条记录都必须保留来源、链接、发布时间、检索时间、相关性说明和置信度
硬性规则
以下任一条违反,视为输出失败:
[MUST-1]所有财务数字必须来自RQData CLI[MUST-2]web_search不得替代财报、估值、价格、公告和一致预期主源[MUST-3]金额类数据必须按可读口径展示;原始“元”金额在正文中应换算为“亿元”[MUST-4]必须保留正式公告原文链接;若原文提取失败,也必须保留失败状态和原文链接[MUST-5]“超预期 / 符合预期 / 低于预期”必须给出口径,不能空喊观点[MUST-6]每个关键数据点或关键结论都要标数据来源:XXX,置信度X[MUST-7]图表若未生成,必须由等价表格或趋势表完成降级,不能让关键章节失真[MUST-8]低置信度外部信息不得改写核心财务判断[MUST-9]最终输出必须严格来自模板,不得在脚本中自由拼写整篇报告
确信度评级
5:RQData CLI、交易所公告、上市公司官网、官方监管披露4:政府 / 监管 / 行业协会 / 官方机构、权威财经媒体3:一般媒体或二手整理,但来源清晰且与其他来源一致2:单一来源、细节不完整、时点未充分验证1:推断、估算窗口、未验证信息
使用规则:
- 混合结论的置信度取关键来源中的最低值
- 推断或估计信息不得标成高置信度
- 低置信度信息只能作为补充背景,不能成为财报结论唯一依据
图表 / 图片需求
当前脚本以趋势表降级交付,但本 skill 仍必须定义最终报告达标所需的图表需求。
- 图表名称:收入与净利润趋势图
- 图表目的:展示近 8 个季度累计与单季收入、净利润变化
- 使用的数据文件:
historical_financials.json - 关键字段:
quarter、revenue、net_profit - 建议图表类型:双轴折线或柱线组合
- 回答问题:增长趋势是否延续,单季拐点是否出现
- 放置位置:
## 财报概览 - 若图表缺失:保留累计趋势表和单季趋势表
- 图表名称:盈利能力与质量图
- 图表目的:展示毛利率、ROE、现金转化率和资产负债率变化
- 使用的数据文件:
historical_financials.json、roe_history.json - 关键字段:
gross_profit、revenue、return_on_equity_weighted_average、cash_from_operating_activities、total_assets、total_liabilities - 建议图表类型:折线图或分组柱状图
- 回答问题:财务质量是在改善还是恶化
- 放置位置:
## 财务质量与资产负债表 - 若图表缺失:保留财务质量表和 ROE 取值表
- 图表名称:预期修正与价格反应图
- 图表目的:展示财报前后预期变化与 1D / 3D / 5D 股价反馈
- 使用的数据文件:
consensus.json、price_window.json、benchmark_window.json - 关键字段:
comp_con_*、con_targ_price、close - 建议图表类型:对比柱图 + 收益曲线
- 回答问题:市场是否把这次财报解读为正面还是负面
- 放置位置:
## 市场预期、卖方反馈与价格反应 - 若图表缺失:保留预期对比表、价格反馈看板和超额收益表
目标产出
- 报告长度:8-12 页
- 输出文件:
- Markdown 报告
- HTML 报告(若本地已安装渲染器)
- 输出目录必须由
--data-dir/--output指定,不能写死固定路径
目录结构
earnings-analysis/
├── SKILL.md
├── scripts/
│ ├── extract_announcements.py
│ └── generate_report.py
├── assets/
│ └── template.md
└── references/
├── data_contract.md
└── web_search.md输入文件契约
原始数据目录由 --data-dir 指定,脚本会按下列文件名查找输入:
company_info.jsonindustry.jsonhistorical_financials.jsonroe_history.jsonmarket_cap.jsonpe_ratio.jsonpb_ratio.jsondividend_yield.jsonprice_window.jsonbenchmark_window.jsonconsensus.jsonresearch_reports.jsonannouncement_raw.jsonannouncement_extracts.json(可选)web_search_findings.json(可选)
完整字段说明见 references/data_contract.md。
工作流
步骤 1:准备参数
REPORT_DATE="${REPORT_DATE:-2026-04-07}"
ORDER_BOOK_ID="${ORDER_BOOK_ID:-600519.XSHG}"
PRICE_WINDOW_START="$(python3 - <<PY
from datetime import date, timedelta
report_date = date.fromisoformat("${REPORT_DATE}")
print((report_date - timedelta(days=45)).isoformat())
PY
)"
CONSENSUS_START="$(python3 - <<PY
from datetime import date, timedelta
report_date = date.fromisoformat("${REPORT_DATE}")
print((report_date - timedelta(days=90)).isoformat())
PY
)"
FACTOR_DATE="$(python3 - <<PY
import json
import subprocess
from datetime import date, timedelta
def has_non_empty_factor(day: str) -> bool:
payload = json.dumps({
"order_book_ids": ["${ORDER_BOOK_ID}"],
"factor": "market_cap",
"start_date": day,
"end_date": day,
}, ensure_ascii=False)
rows = json.loads(subprocess.check_output(
["rqdata", "stock", "cn", "financial-indicator", "--payload", payload, "--format", "json"],
text=True,
))
return any(isinstance(row, dict) and row.get("market_cap") not in (None, "", "null") for row in rows)
cursor = date.fromisoformat("${REPORT_DATE}")
for _ in range(15):
if has_non_empty_factor(cursor.isoformat()):
print(cursor.isoformat())
break
cursor -= timedelta(days=1)
else:
raise SystemExit("最近 15 个自然日都未找到非空 market_cap 因子日期")
PY
)"
HISTORY_START_QUARTER="$(python3 - <<PY
from datetime import date
report_date = date.fromisoformat("${REPORT_DATE}")
print(f"{report_date.year - 2}q1")
PY
)"
HISTORY_END_QUARTER="$(python3 - <<PY
from datetime import date
report_date = date.fromisoformat("${REPORT_DATE}")
print(f"{report_date.year}q4")
PY
)"
DATA_DIR="${DATA_DIR:-$HOME/rq_equities_reports/earnings_analysis}"
OUTPUT_MD="${OUTPUT_MD:-$DATA_DIR/earnings_analysis_${ORDER_BOOK_ID}_${REPORT_DATE}.md}"步骤 2:采集公司、财务、估值与价格数据
mkdir -p "$DATA_DIR"
rqdata stock cn instruments --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"]
}" --format json > "$DATA_DIR/company_info.json"
rqdata stock cn industry --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"date\": \"$REPORT_DATE\",
\"level\": 0,
\"source\": \"citics_2019\"
}" --format json > "$DATA_DIR/industry.json"
rqdata stock cn financial --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"fields\": [\"revenue\", \"net_profit\", \"gross_profit\", \"cash_from_operating_activities\", \"total_assets\", \"total_liabilities\"],
\"start_quarter\": \"$HISTORY_START_QUARTER\",
\"end_quarter\": \"$HISTORY_END_QUARTER\",
\"statements\": \"all\"
}" --format json > "$DATA_DIR/historical_financials.json"
rqdata stock cn financial-indicator --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"factor\": \"return_on_equity_weighted_average\",
\"start_date\": \"$CONSENSUS_START\",
\"end_date\": \"$FACTOR_DATE\"
}" --format json > "$DATA_DIR/roe_history.json"
rqdata stock cn financial-indicator --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"factor\": \"market_cap\",
\"start_date\": \"$FACTOR_DATE\",
\"end_date\": \"$FACTOR_DATE\"
}" --format json > "$DATA_DIR/market_cap.json"
rqdata stock cn financial-indicator --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"factor\": \"pe_ratio\",
\"start_date\": \"$FACTOR_DATE\",
\"end_date\": \"$FACTOR_DATE\"
}" --format json > "$DATA_DIR/pe_ratio.json"
rqdata stock cn financial-indicator --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"factor\": \"pb_ratio\",
\"start_date\": \"$FACTOR_DATE\",
\"end_date\": \"$FACTOR_DATE\"
}" --format json > "$DATA_DIR/pb_ratio.json"
rqdata stock cn financial-indicator --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"factor\": \"dividend_yield\",
\"start_date\": \"$FACTOR_DATE\",
\"end_date\": \"$FACTOR_DATE\"
}" --format json > "$DATA_DIR/dividend_yield.json"
rqdata stock cn price --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"start_date\": \"$PRICE_WINDOW_START\",
\"end_date\": \"$REPORT_DATE\",
\"fields\": [\"close\", \"volume\", \"total_turnover\"],
\"adjust_type\": \"post\"
}" --format json > "$DATA_DIR/price_window.json"
rqdata index price --payload "{
\"order_book_ids\": [\"000300.XSHG\"],
\"start_date\": \"$PRICE_WINDOW_START\",
\"end_date\": \"$REPORT_DATE\",
\"fields\": [\"close\"]
}" --format json > "$DATA_DIR/benchmark_window.json"步骤 3:采集市场预期、研报与公告
rqdata stock cn consensus --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"start_date\": \"$CONSENSUS_START\",
\"end_date\": \"$REPORT_DATE\",
\"report_range\": 3
}" --format json > "$DATA_DIR/consensus.json"
TARGET_FISCAL_YEAR="$(python3 - "$DATA_DIR/historical_financials.json" "$REPORT_DATE" <<'PY'
import json
import sys
from datetime import date, datetime
from pathlib import Path
def parse_iso_date(value):
if not value:
return None
for fmt in ("%Y-%m-%d", "%Y-%m-%d %H:%M:%S"):
try:
return datetime.strptime(str(value)[:19], fmt).date()
except ValueError:
continue
return None
rows = json.loads(Path(sys.argv[1]).read_text())
report_date = date.fromisoformat(sys.argv[2])
best = None
for row in rows:
info_date = parse_iso_date(row.get("info_date"))
quarter = row.get("quarter")
if not info_date or info_date > report_date or not quarter:
continue
if best is None or quarter > best:
best = quarter
print(best[:4] if best else report_date.year)
PY
)"
rqdata stock cn research-reports --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"fiscal_year\": \"$TARGET_FISCAL_YEAR\",
\"start_date\": \"$CONSENSUS_START\",
\"end_date\": \"$REPORT_DATE\",
\"date_rule\": \"create_tm\"
}" --format json > "$DATA_DIR/research_reports.json"
ANNOUNCEMENT_START="$(python3 - <<PY
from datetime import date, timedelta
report_date = date.fromisoformat("${REPORT_DATE}")
print((report_date - timedelta(days=45)).isoformat())
PY
)"
rqdata stock cn announcement --payload "{
\"order_book_ids\": [\"$ORDER_BOOK_ID\"],
\"start_date\": \"$ANNOUNCEMENT_START\",
\"end_date\": \"$REPORT_DATE\"
}" --format json > "$DATA_DIR/announcement_raw.json"步骤 4:提取公告原文片段
python3 earnings-analysis/scripts/extract_announcements.py \
--stock "$ORDER_BOOK_ID" \
--data-dir "$DATA_DIR" \
--report-date "$REPORT_DATE"要求:
- 若公告源站可读,会生成
announcement_extracts.json announcement_extracts.json保留raw_sections、summaries、失败状态和原文链接- 年报 / 中报正文优先抽取
company_intro、management_discussion、risk_warning、outlook - 若源站阻断或 PDF 不可读,也不能静默丢失
步骤 4.5:当前 LLM 回写 announcement_extracts.json summary
- 直接读取
records[].raw_sections.* - 回写
records[].summaries.* - 不额外创建
announcement_summaries.json summaries.*必须是客户可读表述,不得直接复制raw_sections.*原文、不允许粘贴 PDF 抽取碎片- 这一步必须由当前 LLM 完成;
generate_report.py不负责代写公告正文描述 - 若年报 / 中报已有可用
raw_sections.*,但summaries.*仍为空或只是原文截断,视为 workflow 未完成,不进入成稿阶段
步骤 5:补充 web_search 实时信息
仅当需要补充财报电话会、管理层外部表态、行业或政策语境时,才执行此步骤。
- 检索结果必须先写入
web_search_findings.json - 只能补充
RQData CLI直接缺失的信息 - 不得把新闻稿当作财务主数据
步骤 6:生成结构化 Markdown 草稿
python3 earnings-analysis/scripts/generate_report.py \
--stock "$ORDER_BOOK_ID" \
--data-dir "$DATA_DIR" \
--report-date "$REPORT_DATE" \
--output "$OUTPUT_MD"说明:
- 该脚本生成结构化事实草稿,包含表格、信号、研报摘录、公告原文链接和可选外部实时补充信息
- 不应在 Python 里写“因此看多 / 看空”这类分析句
- 若已产出
announcement_extracts.json,报告只消费其中summaries - 不允许在报告生成阶段根据
raw_sections现写客户可读公告摘要;summaries缺失时只能明确提示缺失
步骤 7:渲染 HTML
脚本会优先尝试本地安装的 rq-report-renderer,若不可用则回退到仓库内 report-renderer/scripts/render_report.py;两者都不可用时保留 Markdown 并打印警告。
阶段门控
阶段 1:数据采集完成标准
- 财务、价格、估值、一致预期、研报和公告原始文件齐全
- 最新财报季度已按
info_date <= report-date自动识别 - 若使用
web_search,web_search_findings.json已结构化落盘
阶段 2:研究核查完成标准
- 已核查财报季度、披露日、价格窗口、预期快照和相关公告
- 已核查
announcement_extracts.json -> summaries.*为当前 LLM 回写的客户可读摘要,不是原文截断 - 已确认卖方研报相关性,不包含行业周报 / 策略报告污染
- 公告原文提取成功或失败状态都已保留
阶段 3:成稿完成标准
- 模板占位符全部替换
- 关键章节完整
- 关键结论后有来源和置信度
- 公告章节中的文字描述仅来自
announcement_extracts.json -> summaries.* - 图表未生成时,表格降级仍能覆盖趋势、质量和价格反应
模板规则
- 报告必须严格基于 template.md 生成
- 占位符采用
[[TOKEN]]语法,不使用 Jinja - 当前模板仅允许以下占位符:
[[REPORT_DATE]][[COMPANY_NAME]][[STOCK_CODE]][[LATEST_QUARTER]][[EVENT_DATE]][[EXEC_SUMMARY]][[INFO_PANEL]][[EARNINGS_OVERVIEW]][[EXPECTATION_AND_REACTION]][[ANNOUNCEMENT_SECTION]][[FINANCIAL_QUALITY]][[THESIS_UPDATE]][[VALUATION_SECTION]][[RISK_SECTION]][[APPENDIX]]
报告质量要求
- 完整包含模板中的主章节
- 关键结论必须基于真实财报、预期和价格反应数据
- “超预期 / 符合预期 / 低于预期”必须给出口径
research_reports.json中的summary必须作为卖方文字解释层输出- 公告原文链接必须保留
- 公告正文描述必须先由当前 LLM 回写到
announcement_extracts.json -> summaries.* dividend_yield原始值单位为 bps,报告中必须换算为百分比- 若最新财报季度、关键财务指标或价格窗口缺失,生成器应直接失败
- 若财报前后一致预期字段在
*_t为空、但在更靠后的 forecast slot 非空,生成器必须继续核查并读取实际可用口径,不能机械把财报后预期写成“无数据” - 若使用
web_search,必须保留来源名称、链接、发布时间和检索时间
阶段验收清单
- [ ] 数据采集 -> 报告生成 -> HTML 渲染三阶段都按顺序执行
- [ ]
RQData CLI与web_search的边界没有混用 - [ ] 财报原文、预期对比、价格反应都有真实数据支撑
- [ ]
announcement_extracts.json -> summaries.*已由当前 LLM 回写成客户可读内容,而不是原文截断 - [ ] 卖方研报做过相关性过滤
- [ ] 公告原文链接和提取状态都保留
- [ ] 图表未生成时,表格降级仍覆盖核心问题
- [ ] Markdown / HTML 报告达到 8-12 页目标质量
常见错误
- 使用仓库级
utils/detect_latest_quarter.py - 把
financial-indicator误写成fields + start_quarter/end_quarter - 财报后分析却没有财报前预期口径
- 看到
consensus的*_t为空,就直接把财报后预期写成“无数据”,没有继续核查*_t1 / *_t2 - 把行业周报、策略报告直接当成公司财报点评写进正文
- 用
web_search替代财报、估值、价格、公告和一致预期主源
财报分析报告
- 报告日期:[[REPORT_DATE]]
- 公司:[[COMPANY_NAME]]([[STOCK_CODE]])
- 最新财报期:[[LATEST_QUARTER]]
- 披露日:[[EVENT_DATE]]
执行摘要
[[EXEC_SUMMARY]]
信息截面
[[INFO_PANEL]]
财报概览
[[EARNINGS_OVERVIEW]]
市场预期、卖方反馈与价格反应
[[EXPECTATION_AND_REACTION]]
公告原文与管理层表述
[[ANNOUNCEMENT_SECTION]]
财务质量与资产负债表
[[FINANCIAL_QUALITY]]
投资逻辑更新
[[THESIS_UPDATE]]
估值与定位
[[VALUATION_SECTION]]
风险提示
[[RISK_SECTION]]
附录:生成说明
[[APPENDIX]]
earnings-analysis 数据契约
earnings-analysis/scripts/generate_report.py 默认从 --data-dir 读取以下 JSON 文件。
1. company_info.json
典型字段:
order_book_idsymbolabbrev_symbolindustry_namelisted_date
用途:
- 公司名称、简称、上市信息
2. industry.json
典型字段:
first_industry_namesecond_industry_namethird_industry_name
用途:
- 行业分层描述
3. historical_financials.json
典型字段:
order_book_idquarterinfo_daterevenuenet_profitgross_profitcash_from_operating_activitiestotal_assetstotal_liabilities
用途:
- 自动识别最新财报季度
- 计算同比、环比、毛利率、现金转化率、资产负债率
4. roe_history.json
典型字段:
order_book_iddatereturn_on_equity_weighted_average
用途:
- 盈利质量趋势
5. market_cap.json / pe_ratio.json / pb_ratio.json / dividend_yield.json
典型字段:
order_book_iddate- 对应 factor 字段
说明:
dividend_yield原始值为 bps,生成报告时需要除以100后按百分比展示
用途:
- 当前估值与股东回报定位
6. price_window.json
典型字段:
order_book_iddatetimeclosevolumetotal_turnover
用途:
- 计算财报前后价格反应
- 计算成交额放大
7. benchmark_window.json
典型字段:
order_book_iddatetimeclose
用途:
- 计算相对沪深300的超额收益
8. consensus.json
典型字段:
datecreate_tmreport_year_tcomp_con_operating_revenue_t / t1 / t2 / t3comp_con_net_profit_t / t1 / t2 / t3comp_con_eps_t / t1 / t2 / t3con_targ_price
用途:
- 财报前一致预期
- 财报后预期变化
9. research_reports.json
典型字段:
create_tmdatereport_titlesummaryinstituteauthorfiscal_yearnet_profit_t / t1 / t2eps_t / t1 / t2targ_price
用途:
- 财报后卖方解读
- 目标价和年度利润口径的补充
说明:
summary是财报后“文字解释层”的首选字段,用于补充卖方对业绩、预期修正和核心关注点的描述report_title、summary、targ_price与net_profit_t / t1 / t2需要一起看,不能只保留数值预测
10. announcement_raw.json
典型字段:
info_datetitleinfo_typemediafile_typeannouncement_link
用途:
- 识别正式财报、主要经营数据公告和业绩说明会公告
- 在正文中保留公告原文链接,供后续 PDF / HTML 读取
11. announcement_extracts.json
该文件可选,可由具备 PDF / HTML 原文解析能力的流程生成。
典型字段:
records[].titlerecords[].info_daterecords[].announcement_linkrecords[].is_annual_or_interim_reportrecords[].fetch_statusrecords[].extract_statusrecords[].raw_sections.company_introrecords[].raw_sections.management_discussionrecords[].raw_sections.risk_warningrecords[].raw_sections.outlookrecords[].summaries.company_introrecords[].summaries.management_discussionrecords[].summaries.risk_warningrecords[].summaries.outlookrecords[].sections.company_introrecords[].sections.management_discussionrecords[].sections.risk_warningrecords[].sections.outlook
用途:
raw_sections保存较长原文段落,供当前 skill 内的 LLM 直接读取summaries保存基于raw_sections回写的总结性文本;报告生成时只消费该层summaries必须由当前 LLM 回写客户可读摘要,不能直接复制raw_sections原文或 PDF 抽取碎片sections为兼容旧结构保留,当前可视为raw_sections的兼容镜像company_intro/management_discussion/outlook主要面向年报、半年报正文;季度报告和临时公告默认不强制抽取这三类字段- 若源站拦截或 PDF 不可读,也必须保留失败状态和原文链接,不能静默丢失
- 不额外创建
announcement_summaries.json;LLM 应直接在announcement_extracts.json的records[].summaries.*中回写结果
12. web_search_findings.json
该文件可选,仅用于补充 RQData CLI 无法直接提供的实时外部语境。
每条记录至少包含:
querysource_namesource_typetitleurlpublished_atretrieved_atsummarywhy_relevantconfidencefinding_type
推荐附加字段:
subjectstance
允许的 finding_type:
company_newsmanagement_updateearnings_callindustry_contextpolicy_context
说明:
published_at是源内容发布时间,不是财报披露日retrieved_at是实际检索时间source_type/confidence需遵守references/web_search.md的来源等级约束web_search_findings.json不能替代财务、估值、价格、公告和一致预期主源
解析约定
- 所有文件都允许
{"data": [...]}、{"data": {...}}、[...]、{...}四种包装方式 - 同一季度多次披露时,脚本按
info_date <= report-date选择最新版本 - 财报分析使用的目标年度是最新已披露财报年度;读取
consensus时不能机械假设*_t一定非空,必须继续核查实际有值的 forecast slot - 研报必须做相关性过滤,至少要求标题或摘要命中公司名称/代码/英文名关键词
- 可直接运行
earnings-analysis/scripts/extract_announcements.py生成该文件;若源站阻断,也应保留失败状态 - 若存在
web_search_findings.json,脚本会校验必填字段、来源类别和置信度上限
Earnings Analysis Web Search Reference
Purpose
Use web_search only to supplement real-time information that RQData CLI does not directly provide for a post-earnings report.
Allowed Coverage
- Earnings call schedules or management public remarks
- Important post-earnings company news
- Industry or policy context relevant to the earnings interpretation
- External signals that help explain expectation revisions or market reaction
Prohibited Usage
- Do not replace financial statements, valuation multiples, prices, official announcements, or consensus data
- Do not use
web_searchto fabricate earnings dates or official disclosure details - Do not promote low-confidence media snippets into core earnings conclusions
Required Output File
All external findings must be written to web_search_findings.json.
Each record must contain:
querysource_namesource_typetitleurlpublished_atretrieved_atsummarywhy_relevantconfidencefinding_type
Recommended fields:
subjectstance
Allowed finding_type
company_newsmanagement_updateearnings_callindustry_contextpolicy_context
Source Types And Confidence Ceiling
official: max confidence5government: max confidence4association: max confidence4authoritative_media: max confidence4general_news: max confidence3inference: max confidence1
Search Workflow
1. Confirm the needed information is not directly available from RQData CLI. 2. Prefer official and primary sources first. 3. Save the finding into web_search_findings.json with structured metadata. 4. Keep a short summary and a concrete relevance note. 5. If confidence is low, keep it as context only and do not let it dominate the conclusion.
Fallback
1. Use the native web_search tool when available. 2. Otherwise use the configured network search tool in the current environment. 3. If neither is available:
- do not fabricate real-time information
- explicitly mark that context as unavailable or unverified
- lower confidence rather than guessing
Example
{
"data": [
{
"query": "贵州茅台 2026 业绩说明会 时间",
"source_name": "贵州茅台官网",
"source_type": "official",
"title": "2025年度业绩说明会召开公告",
"url": "https://www.example.com/ir-call",
"published_at": "2026-03-30",
"retrieved_at": "2026-04-07",
"summary": "公司披露 2025 年度业绩说明会将在 4 月中旬召开。",
"why_relevant": "有助于判断财报后管理层沟通节奏和市场关注焦点。",
"confidence": 5,
"finding_type": "earnings_call",
"subject": "业绩说明会",
"stance": "neutral"
}
]
}#!/usr/bin/env python3
"""Extract structured announcement snippets for earnings-analysis."""
from __future__ import annotations
import argparse
import json
import re
import zlib
from datetime import date
from pathlib import Path
from typing import Any, Dict, List, Optional, Sequence, Tuple
import requests
from generate_report import (
build_snapshot,
dedupe_financial_records,
extract_records,
parse_iso_date,
read_json_file,
select_relevant_announcements,
)
USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"
)
ACW_POS_LIST = [
0x0F,
0x23,
0x1D,
0x18,
0x21,
0x10,
0x01,
0x26,
0x0A,
0x09,
0x13,
0x1F,
0x28,
0x1B,
0x16,
0x17,
0x19,
0x0D,
0x06,
0x0B,
0x27,
0x12,
0x14,
0x08,
0x0E,
0x15,
0x20,
0x1A,
0x02,
0x1E,
0x07,
0x04,
0x11,
0x05,
0x03,
0x1C,
0x22,
0x25,
0x0C,
0x24,
]
ACW_MASK = "3000176000856006061501533003690027800375"
OBJ_RE = re.compile(rb"(\d+)\s+(\d+)\s+obj\b(.*?)endobj", re.S)
STREAM_RE = re.compile(rb"<<(.*?)>>\s*stream\r?\n(.*?)\r?\nendstream", re.S)
PAGE_RE = re.compile(rb"/Type\s*/Page\b")
TEXT_OP_RE = re.compile(
r"/([A-Za-z0-9]+)\s+[0-9.]+\s+Tf|"
r"<([0-9A-Fa-f\s]+)>\s*Tj|"
r"\[(.*?)\]\s*TJ|"
r"\(((?:\\.|[^\\)])*)\)\s*Tj|"
r"(-?[0-9.]+)\s+(-?[0-9.]+)\s+T[Dd]|"
r"T\*|BT|ET",
re.S,
)
TEXT_SECTION_STOP_MARKERS = [
"重要内容提示",
"一、主要财务数据",
"二、股东信息",
"三、其他提醒事项",
"四、季度财务报表",
"五、重要事项",
"六、其他事项",
"风险提示",
"重大风险提示",
"经营情况讨论与分析",
"管理层讨论与分析",
"投资者关系活动主要内容介绍",
"未来展望",
"经营计划",
"发展战略",
]
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="提取公告 PDF 正文片段并生成 announcement_extracts.json")
parser.add_argument("--stock", required=True, help="股票代码")
parser.add_argument("--data-dir", required=True, help="原始 JSON 数据目录")
parser.add_argument("--report-date", required=True, help="报告日期 (YYYY-MM-DD)")
parser.add_argument("--output", help="输出 JSON 路径,默认写到 data-dir/announcement_extracts.json")
parser.add_argument("--timeout", type=float, default=20.0, help="公告抓取超时时间,默认 20 秒")
return parser.parse_args()
def calc_sse_acw_cookie(arg1: str) -> str:
out = [""] * len(ACW_POS_LIST)
for idx, char in enumerate(arg1):
for out_idx, pos in enumerate(ACW_POS_LIST):
if pos == idx + 1:
out[out_idx] = char
break
arg2 = "".join(out)
pieces = []
for idx in range(0, min(len(arg2), len(ACW_MASK)), 2):
pieces.append(f"{int(arg2[idx:idx + 2], 16) ^ int(ACW_MASK[idx:idx + 2], 16):02x}")
return "".join(pieces)
def fetch_pdf_bytes(url: str, timeout: float) -> Tuple[Optional[bytes], str]:
headers = {"User-Agent": USER_AGENT, "Accept": "application/pdf,text/html,*/*", "Referer": url}
session = requests.Session()
try:
response = session.get(url, timeout=timeout, headers=headers, allow_redirects=True)
except requests.RequestException as exc:
return None, f"network_error:{type(exc).__name__}"
content_type = (response.headers.get("content-type") or "").lower()
if response.ok and (content_type.startswith("application/pdf") or response.content.startswith(b"%PDF-")):
return response.content, "ok"
if "static.sse.com.cn" in response.url and "text/html" in content_type:
match = re.search(r"arg1='([^']+)'", response.text)
if not match:
return None, "source_blocked:sse_html_without_arg1"
cookie = calc_sse_acw_cookie(match.group(1))
session.cookies.set("acw_sc__v2", cookie, domain="static.sse.com.cn", path="/")
try:
retry = session.get(
url,
timeout=timeout,
headers={"User-Agent": USER_AGENT, "Accept": "application/pdf,*/*", "Referer": "http://www.sse.com.cn/"},
allow_redirects=True,
)
except requests.RequestException as exc:
return None, f"network_error:{type(exc).__name__}"
retry_type = (retry.headers.get("content-type") or "").lower()
if retry.ok and (retry_type.startswith("application/pdf") or retry.content.startswith(b"%PDF-")):
return retry.content, "ok"
return None, f"source_blocked:sse_retry_{retry.status_code}"
if not response.ok:
return None, f"http_{response.status_code}"
return None, f"unsupported_content_type:{content_type or 'unknown'}"
def parse_pdf_objects(pdf_bytes: bytes) -> Dict[int, bytes]:
return {int(match.group(1)): match.group(3) for match in OBJ_RE.finditer(pdf_bytes)}
def parse_stream(raw_object: bytes) -> Tuple[Optional[bytes], Optional[bytes]]:
match = STREAM_RE.search(raw_object)
if not match:
return None, None
stream_dict = match.group(1)
stream_data = match.group(2)
if b"/FlateDecode" in stream_dict:
stream_data = zlib.decompress(stream_data)
return stream_dict, stream_data
def decode_utf16be_hex(value: str) -> str:
return bytes.fromhex(value).decode("utf-16-be", "ignore")
def build_cmap(stream_text: str) -> Dict[str, str]:
cmap: Dict[str, str] = {}
for block in re.findall(r"beginbfchar\s*(.*?)\s*endbfchar", stream_text, re.S):
for src, dst in re.findall(r"<([0-9A-Fa-f]+)>\s*<([0-9A-Fa-f]+)>", block):
cmap[src.upper()] = decode_utf16be_hex(dst)
for block in re.findall(r"beginbfrange\s*(.*?)\s*endbfrange", stream_text, re.S):
for start, end, dst in re.findall(r"<([0-9A-Fa-f]+)>\s*<([0-9A-Fa-f]+)>\s*<([0-9A-Fa-f]+)>", block):
start_int = int(start, 16)
end_int = int(end, 16)
dst_int = int(dst, 16)
width = len(start)
out_len = len(dst) // 2
for idx, code in enumerate(range(start_int, end_int + 1)):
cmap[f"{code:0{width}X}"] = (dst_int + idx).to_bytes(out_len, "big").decode("utf-16-be", "ignore")
for start, _end, arr in re.findall(r"<([0-9A-Fa-f]+)>\s*<([0-9A-Fa-f]+)>\s*\[(.*?)\]", block, re.S):
start_int = int(start, 16)
width = len(start)
for idx, dst in enumerate(re.findall(r"<([0-9A-Fa-f]+)>", arr)):
cmap[f"{start_int + idx:0{width}X}"] = decode_utf16be_hex(dst)
return cmap
def decode_pdf_hex(hex_text: str, cmap: Dict[str, str]) -> str:
hex_text = re.sub(r"\s+", "", hex_text)
if not hex_text:
return ""
key_lengths = sorted({len(key) for key in cmap}, reverse=True) if cmap else [2]
cursor = 0
output: List[str] = []
while cursor < len(hex_text):
matched = False
for width in key_lengths:
key = hex_text[cursor:cursor + width].upper()
if len(key) == width and key in cmap:
output.append(cmap[key])
cursor += width
matched = True
break
if matched:
continue
chunk = hex_text[cursor:cursor + 2]
if len(chunk) == 2:
try:
output.append(bytes.fromhex(chunk).decode("latin1"))
except ValueError:
pass
cursor += 2
return "".join(output)
def decode_pdf_literal(text: str) -> str:
return (
text.replace(r"\(", "(")
.replace(r"\)", ")")
.replace(r"\n", "\n")
.replace(r"\r", "")
.replace(r"\t", "\t")
.replace(r"\\", "\\")
)
def extract_pdf_text(pdf_bytes: bytes) -> str:
objects = parse_pdf_objects(pdf_bytes)
font_cmaps: Dict[int, Dict[str, str]] = {}
for obj_num, raw_object in objects.items():
match = re.search(rb"/ToUnicode\s+(\d+)\s+0\s+R", raw_object)
if not match:
continue
stream_ref = int(match.group(1))
if stream_ref not in objects:
continue
_stream_dict, stream_data = parse_stream(objects[stream_ref])
if not stream_data:
continue
font_cmaps[obj_num] = build_cmap(stream_data.decode("latin1", "ignore"))
pages: List[Tuple[int, List[int], Dict[str, int]]] = []
for obj_num, raw_object in objects.items():
if not PAGE_RE.search(raw_object):
continue
content_refs = [int(value) for value in re.findall(rb"/Contents\s+(\d+)\s+0\s+R", raw_object)]
if not content_refs:
array_match = re.search(rb"/Contents\s*\[(.*?)\]", raw_object, re.S)
if array_match:
content_refs = [int(value) for value in re.findall(rb"(\d+)\s+0\s+R", array_match.group(1))]
font_map: Dict[str, int] = {}
font_block = re.search(rb"/Font\s*<<(.+?)>>", raw_object, re.S)
if font_block:
for font_name, font_ref in re.findall(rb"/([A-Za-z0-9]+)\s+(\d+)\s+0\s+R", font_block.group(1)):
font_map[font_name.decode("ascii", "ignore")] = int(font_ref)
pages.append((obj_num, content_refs, font_map))
pages.sort(key=lambda item: item[0])
lines: List[str] = []
current_font: Optional[str] = None
for _page_num, content_refs, font_map in pages:
for content_ref in content_refs:
if content_ref not in objects:
continue
_stream_dict, stream_data = parse_stream(objects[content_ref])
if not stream_data:
continue
content_text = stream_data.decode("latin1", "ignore")
current_line: List[str] = []
for match in TEXT_OP_RE.finditer(content_text):
token = match.group(0)
if " Tf" in token:
current_font = match.group(1)
continue
if token == "BT":
current_line = []
continue
if token == "ET":
line = "".join(current_line).strip()
if line:
lines.append(line)
current_line = []
continue
if token == "T*" or token.endswith("TD") or token.endswith("Td"):
if match.group(6) and abs(float(match.group(6))) > 1e-6:
line = "".join(current_line).strip()
if line:
lines.append(line)
current_line = []
continue
if token.endswith("Tj") and token.startswith("<"):
font_ref = font_map.get(current_font or "")
cmap = font_cmaps.get(font_ref, {})
current_line.append(decode_pdf_hex(match.group(2), cmap))
continue
if token.endswith("TJ"):
font_ref = font_map.get(current_font or "")
cmap = font_cmaps.get(font_ref, {})
segment = match.group(3) or ""
for hex_group in re.findall(r"<([0-9A-Fa-f\s]+)>", segment):
current_line.append(decode_pdf_hex(hex_group, cmap))
for literal in re.findall(r"\(((?:\\.|[^\\)])*)\)", segment):
current_line.append(decode_pdf_literal(literal))
continue
current_line.append(decode_pdf_literal(match.group(4)))
text = "\n".join(line for line in lines if line.strip())
text = text.replace("\r", "\n").replace("\u3000", "")
text = re.sub(r"[ \t]+\n", "\n", text)
text = re.sub(r"\n{3,}", "\n\n", text)
return text.strip()
def squash_text(text: str) -> str:
return re.sub(r"\s+", "", text or "")
def clip_text(text: str, limit: int = 260) -> str:
text = str(text or "").strip()
if len(text) <= limit:
return text
return text[: limit - 1].rstrip(",、;: ") + "…"
def normalize_section_text(text: str, limit: Optional[int] = None) -> str:
text = str(text or "")
text = re.sub(r"[\x00-\x08\x0b\x0c\x0e-\x1f]", "", text)
text = re.sub(r"\s+", "", text)
if not text:
return ""
meaningful_chars = re.findall(r"[\u4e00-\u9fffA-Za-z0-9,。!?;:、“”‘’()()\-%./]", text)
if len(meaningful_chars) < max(20, int(len(text) * 0.6)):
return ""
if not re.search(r"[\u4e00-\u9fffA-Za-z]", text):
return ""
if limit is None:
return text
return clip_text(text, limit)
def is_annual_or_interim_report(title: str, info_type: str) -> bool:
title = str(title or "")
info_type = str(info_type or "")
if not re.search(r"(年度报告|年报|半年度报告|半年报|中报)", title):
return False
if re.search(r"(摘要|英文版|公告|业绩说明会|主要经营数据|信息披露公告)", title):
return False
return "定期报告" in info_type or bool(re.search(r"(年度报告|年报|半年度报告|半年报|中报)", title))
def find_marker_window(
text: str,
markers: Sequence[str],
stop_markers: Sequence[str],
max_chars: int,
forbidden_patterns: Sequence[str] = (),
) -> str:
candidates: List[Tuple[int, str]] = []
for marker in markers:
start = 0
while True:
idx = text.find(marker, start)
if idx < 0:
break
candidates.append((idx, marker))
start = idx + len(marker)
if not candidates:
return ""
candidates.sort(key=lambda item: item[0])
for best_start, matched_marker in candidates:
local_context = text[max(0, best_start - 80): min(len(text), best_start + 120)]
if re.search(r"[..。…]{12,}", local_context):
continue
search_start = best_start + len(matched_marker)
end_positions = [
text.find(stop_marker, search_start)
for stop_marker in stop_markers
if stop_marker not in markers and text.find(stop_marker, search_start) >= 0
]
end = min(end_positions) if end_positions else min(len(text), best_start + max_chars)
end = min(end, best_start + max_chars)
snippet = normalize_section_text(text[best_start:end], max_chars)
if snippet and forbidden_patterns and any(pattern in snippet for pattern in forbidden_patterns):
continue
if snippet:
return snippet
return ""
def find_sentence_by_keywords(text: str, keywords: Sequence[str], max_chars: int) -> str:
sentences = re.split(r"(?<=[。!?;])", text)
for sentence in sentences:
sentence = sentence.strip()
if sentence and any(keyword in sentence for keyword in keywords):
return normalize_section_text(sentence, max_chars)
collapsed = text
for keyword in keywords:
idx = collapsed.find(keyword)
if idx >= 0:
start = max(0, idx - 40)
end = min(len(collapsed), idx + max_chars)
return normalize_section_text(collapsed[start:end], max_chars)
return ""
def extract_sections(title: str, info_type: str, raw_text: str) -> Dict[str, str]:
squashed = squash_text(raw_text)
stop_markers = TEXT_SECTION_STOP_MARKERS
long_form_report = is_annual_or_interim_report(title, info_type)
company_intro = ""
management_discussion = ""
outlook = ""
if long_form_report:
intro_end = len(squashed)
for marker in ("重要内容提示", "一、主要财务数据"):
idx = squashed.find(marker)
if idx >= 0:
intro_end = min(intro_end, idx)
company_intro = normalize_section_text(squashed[:intro_end] or squashed[:220], 220)
company_intro_marked = find_marker_window(
squashed,
["公司简介", "公司基本情况", "发行人基本情况"],
stop_markers,
220,
)
if company_intro_marked:
company_intro = company_intro_marked
management_discussion = find_marker_window(
squashed,
[
"管理层讨论与分析",
"经营情况讨论与分析",
"董事会报告",
"经营回顾",
],
stop_markers,
280,
)
if not management_discussion:
management_discussion = find_sentence_by_keywords(
squashed,
["经营", "销量", "需求", "增长", "盈利能力", "毛利率", "渠道", "产能"],
240,
)
outlook = find_marker_window(
squashed,
["未来展望", "经营计划", "发展战略", "未来规划", "下半年展望", "后续规划"],
stop_markers,
220,
forbidden_patterns=("前瞻性陈述", "注意投资风险"),
)
if not outlook:
outlook = find_sentence_by_keywords(
squashed,
["未来", "展望", "预计", "计划", "规划", "目标", "将继续", "后续"],
220,
)
risk_warning = find_marker_window(
squashed,
["风险提示", "重大风险提示", "风险因素", "重大风险"],
stop_markers,
220,
)
return {
"company_intro": company_intro,
"management_discussion": management_discussion,
"risk_warning": risk_warning,
"outlook": outlook,
}
def build_raw_sections(title: str, info_type: str, raw_text: str) -> Dict[str, str]:
squashed = squash_text(raw_text)
stop_markers = TEXT_SECTION_STOP_MARKERS
long_form_report = is_annual_or_interim_report(title, info_type)
company_intro = ""
management_discussion = ""
outlook = ""
if long_form_report:
intro_end = len(squashed)
for marker in ("重要内容提示", "一、主要财务数据"):
idx = squashed.find(marker)
if idx >= 0:
intro_end = min(intro_end, idx)
company_intro = normalize_section_text(squashed[:intro_end] or squashed[:1200], 1200)
company_intro_marked = find_marker_window(
squashed,
["公司简介", "公司基本情况", "发行人基本情况"],
stop_markers,
1400,
)
if company_intro_marked:
company_intro = company_intro_marked
management_discussion = find_marker_window(
squashed,
["管理层讨论与分析", "经营情况讨论与分析", "董事会报告", "经营回顾"],
stop_markers,
2600,
)
if not management_discussion:
management_discussion = find_sentence_by_keywords(
squashed,
["经营", "销量", "需求", "增长", "盈利能力", "毛利率", "渠道", "产能"],
1600,
)
outlook = find_marker_window(
squashed,
["未来展望", "经营计划", "发展战略", "未来规划", "下半年展望", "后续规划"],
stop_markers,
1800,
forbidden_patterns=("前瞻性陈述", "注意投资风险"),
)
if not outlook:
outlook = find_sentence_by_keywords(
squashed,
["未来", "展望", "预计", "计划", "规划", "目标", "将继续", "后续"],
1200,
)
risk_warning = find_marker_window(
squashed,
["风险提示", "重大风险提示", "风险因素", "重大风险"],
stop_markers,
1400,
)
return {
"company_intro": company_intro,
"management_discussion": management_discussion,
"risk_warning": risk_warning,
"outlook": outlook,
}
def choose_extract_status(sections: Dict[str, str], title: str, info_type: str) -> str:
populated = sum(1 for value in sections.values() if value)
if not is_annual_or_interim_report(title, info_type) and populated == 0:
return "skipped_non_annual_interim"
if populated >= 4:
return "ok"
if populated >= 1:
return "partial"
return "no_sections"
def main() -> None:
args = parse_args()
report_date = date.fromisoformat(args.report_date)
data_dir = Path(args.data_dir).expanduser()
output_path = Path(args.output).expanduser() if args.output else data_dir / "announcement_extracts.json"
financial_records = extract_records(read_json_file(data_dir / "historical_financials.json"))
announcement_records = extract_records(read_json_file(data_dir / "announcement_raw.json"))
if not financial_records:
raise ValueError("缺少 historical_financials.json,无法定位财报事件日")
deduped_financials = dedupe_financial_records(financial_records, args.stock, report_date)
latest_snapshot = build_snapshot(deduped_financials[-1]) if deduped_financials else None
if latest_snapshot is None:
raise ValueError("未识别到 report-date 之前的最新财报季度")
selected_announcements = select_relevant_announcements(
announcement_records,
args.stock,
latest_snapshot.info_date,
report_date,
)
records: List[Dict[str, Any]] = []
for item in selected_announcements:
title = str(item.get("title") or "")
link = str(item.get("announcement_link") or "")
info_date = parse_iso_date(item.get("info_date") or item.get("date") or item.get("create_tm"))
empty_sections = {
"company_intro": "",
"management_discussion": "",
"risk_warning": "",
"outlook": "",
}
record: Dict[str, Any] = {
"title": title,
"info_date": info_date.isoformat() if info_date else str(item.get("info_date") or ""),
"announcement_link": link,
"media": item.get("media"),
"info_type": item.get("info_type"),
"is_annual_or_interim_report": is_annual_or_interim_report(title, str(item.get("info_type") or "")),
"fetch_status": "skipped",
"extract_status": "not_started",
"raw_sections": dict(empty_sections),
"summaries": dict(empty_sections),
"sections": dict(empty_sections),
}
if str(item.get("file_type") or "").upper() != "PDF":
record["fetch_status"] = "unsupported_file_type"
record["extract_status"] = "unsupported"
records.append(record)
continue
pdf_bytes, fetch_status = fetch_pdf_bytes(link, args.timeout)
record["fetch_status"] = fetch_status
if not pdf_bytes:
record["extract_status"] = "fetch_failed"
records.append(record)
continue
try:
extracted_text = extract_pdf_text(pdf_bytes)
except Exception as exc: # pragma: no cover - defensive branch for malformed PDFs
record["extract_status"] = f"pdf_parse_failed:{type(exc).__name__}"
records.append(record)
continue
raw_sections = build_raw_sections(title, str(item.get("info_type") or ""), extracted_text)
sections = extract_sections(title, str(item.get("info_type") or ""), extracted_text)
record["raw_sections"] = raw_sections
record["summaries"] = dict(empty_sections)
record["sections"] = raw_sections
record["extract_status"] = choose_extract_status(raw_sections, title, str(item.get("info_type") or ""))
records.append(record)
output_path.parent.mkdir(parents=True, exist_ok=True)
payload = {
"stock": args.stock,
"report_date": args.report_date,
"event_date": latest_snapshot.info_date.isoformat(),
"record_count": len(records),
"records": records,
}
output_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"✅ 公告提取结果已写入:{output_path}")
print(f"相关公告样本:{len(records)} 条")
if __name__ == "__main__":
main()
Related skills
FAQ
Is Rq Earnings Analysis safe to install?
skills.sh reports 1 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.