
Lunwen
- 82 installs
- 631 repo stars
- Updated April 5, 2026
- doryoku1223/lunwen-skill
Turns real project facts, templates, literature constraints, and chart screenshots into a fully formatted Chinese thesis .docx with controlled length and complete visuals.
About
lunwen is a Claude skill that converts Chinese academic project materials, sample templates, literature constraints, and chart screenshots into a polished, correctly formatted thesis .docx. A solo builder or researcher uses it to assemble a deliverable paper that respects structure and word-count limits instead of just padding length.
- Generates formatted Chinese thesis .docx
- Preserves project facts and templates
- Enforces literature constraints
- Controlled word count, complete charts
Lunwen by the numbers
- 82 all-time installs (skills.sh)
- Ranked #341 of 688 Office & Documents skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/doryoku1223/lunwen-skill --skill lunwenAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 82 |
|---|---|
| repo stars | ★ 631 |
| Last updated | April 5, 2026 |
| Repository | doryoku1223/lunwen-skill ↗ |
What it does
Turns real project facts, templates, literature constraints, and chart screenshots into a fully formatted Chinese thesis .docx with controlled length and complete visuals.
Who is it for?
Producing a formatted Chinese thesis from project materials
Skip if: English papers or general-purpose coding
What you get
- formatted thesis .docx
Files
Lunwen
Overview
这个技能用于把“真实项目事实 + 样文/模板 + 文献约束 + 图表截图 + Word 成稿要求”稳定转化为一篇可交付的中文论文。目标不是把论文写厚,而是按样文体量、真实项目能力和版式规则,输出结构完整、字数可控、图表齐全的 .docx 成稿。
Hard Gates
以下规则是硬约束:
1. 用户第一次说“为项目生成论文”时,必须先索要辅助资料,至少包括:
- 学校模板
- 往届样文
- 开题报告 / 任务书
- 封面字段要求
- 字数要求
2. 在用户明确表示“没有更多资料”或已经给出可分析资料之前,禁止直接生成论文初稿。 3. 如果用户给了样文或模板,必须先分析样文与样式,先回传“建议目录 + 目标字数 + 样式摘要 + 冲突项”,等待用户确认后才能开写正文。 4. 输出文件名必须使用论文标题,不得使用 smart-lab-thesis、final、draft 这类通用名。 5. 除主论文 .docx 外,必须额外交付一个“附件 .docx”,用于收录正文中的 Mermaid / PlantUML 图源码、数据库 E-R 图源码、关键流程图源码等不渲染版本。 6. 第 4 章“系统详细设计与实现”必须是全文最长章节;每个一级功能模块默认都要包含:
- 模块说明
- 页面截图
- 简要核心代码或关键实现片段
7. 如果正文数据库设计部分缺少 E-R 图,则视为未完成。 8. 如果用户给了任务书 / 开题报告 / 学校模板 / 样文,必须先分析这些材料,禁止绕过分析直接生成正文。 9. 如果任务书、样文、README、旧说明文档和源码之间出现冲突,默认以“学校模板 > 任务书/开题报告 > 用户明确要求 > 样文 > 项目源码 > README > 旧说明文档”为优先级。 10. 旧部署说明、历史项目介绍、演示文档默认视为低可信来源;除非用户明确要求,否则不能作为论文事实主依据。
Execution States
本技能默认按以下状态机执行:
1. intake_only 只收资料,不写正文。 2. sample_analysis_done 已完成样文 / 模板 / 任务书分析,但未开写。 3. outline_confirmed 用户已确认目录、字数和样式方案。 4. writing_allowed 允许开始正文写作。 5. delivery_done 已生成主论文、附件和校验产物。
状态约束:
- 未进入
outline_confirmed前,禁止写正文。 - 若用户中途补交新模板、新样文、新任务书,状态必须退回
sample_analysis_done。 - 只有进入
writing_allowed后,才允许生成主论文.md/.docx。
Core Flow
1. 锁定输入
先确认以下输入并确定优先级:
1. 学校模板 2. 任务书 / 开题报告 3. 用户上传的往届样文 4. 用户口头要求 5. 项目源码 6. README 7. 旧项目说明文档 8. 技能默认规则
首次响应论文请求时,必须主动告诉用户可以直接提供模板、历届样文、开题报告、任务书或封面要求的本地路径,示例格式如:
D:\论文模板.docxD:\论文样文1.docxD:\开题报告.pdf
并且必须明确说明:
- 如果用户还没提供样文、模板或任务书,当前阶段只收集资料和分析,不直接开写初稿。
- 如果用户确认“没有更多资料”,才能按项目源码和默认规则继续。
如果模板、样文和默认规则冲突,必须列出冲突项并让用户选择。样式冲突细则见:
prompts/intake.mdprompts/style_extractor.mdreferences/default-style.md
2. 冻结项目事实
先读项目代码和文档,再提炼固定的项目事实底稿。后续各章只能基于这份底稿扩写。
冻结事实时,必须明确区分:
1. 任务书要求 2. 源码实际实现 3. 论文最终采用口径
如果三者不一致,必须在设计回传阶段提前告诉用户,不得等正文写完后再修正。
事实提取细则见:
prompts/fact_extractor.md
3. 学样文,不只学目录
如果用户提供样文或模板,必须同时分析:
- 结构:目录、页数、字数、图表节奏
- 样式:标题、正文、摘要、关键词、图题表题、参考文献、致谢
- 细节:中英文正文的字体字号、段前段后、行距、首行缩进
- 标题:一级、二级、三级标题的字体字号、加粗、对齐、分页方式
- 表图代码:图题、表题、表格内容、代码块或代码截图的插入位置与样式
对应资源:
prompts/sample_analyzer.mdprompts/style_extractor.mdtools/analyze_sample_pdf.pytools/analyze_docx.py
3.5 先回传设计,再开写
模板和样文分析结束后,必须先把以下内容回传给用户确认:
1. 当前论文建议目录 2. 各章目标字数 3. 正文、标题、摘要、关键词、图题、表题、表格内容等版式样式 4. 与默认规则的冲突项
回传时,默认必须结构化为 4 张表:
1. 输入资料表 2. 建议目录表 3. 字数预算表 4. 样式与冲突表
必须明确等待用户确认目录和样式;若用户提出新的修改意见,以用户最后确认的版本为准,再进入正文写作。
如果用户后续又补充了新样文、模板或任务书,必须中断正文写作流程,回到本步骤重新分析,不得沿用旧设计继续写。
4. 先定字数,再写作
写正文前必须生成目标章节字数表。默认优先贴近样文体量,不默认写厚。写完一章就统计一次,超出就压缩。
对应资源:
prompts/chapter_writer.mdtools/count_chapter_words.pyreferences/chapter-patterns.md
5. 图表与截图闭环
图表默认要求:
- 系统架构图
- E-R 图
- 关键流程图
- 数据表
- 测试用例表
E-R 图规则:
- 如果数据库总表 E-R 图过大,必须拆分为至少两张:
1. 总表设计图 2. 核心表设计图
- 若业务复杂,可继续拆成“权限与用户域”“预约与改派域”“设备与报修域”“耗材域”等多张图
- 正文中至少保留“总表设计图 + 核心表设计图”两张 E-R 图
截图默认要求:
- 第 4 章每个主要模块至少 1 张页面截图
- 如果模块跨“管理端 / 用户端”,优先两端都给截图
- 截图标题必须与正文模块对应,不得只放“系统页面图”这类空泛标题
代码默认要求:
- 第 4 章每个主要模块至少给出 1 段简要核心代码、SQL 片段或关键实现片段
- 代码以“解释业务规则”为目的,不堆大段源码
- 大段源码放附录,不直接塞正文
第 3 章设计图最低要求:
- 系统总体架构图
- 功能模块结构图
- 权限控制流程图
- 关键业务流程图
- E-R 图
第 4 章截图最低要求:
- 每个主模块至少 1 张截图
- 主模块跨多角色时,优先补多角色截图
- 截图总数不足 6 张时,默认视为实现章节不充分
如果存在 mermaid / plantuml,优先渲染为真实图片;若无法渲染,再退回源码或占位。
仓库内置 Playwright 截图链路。默认先用本仓库脚本自动抓取系统页面,不再要求用户额外安装浏览器 skill。
如果宿主环境提供 Chrome CDP / Chrome MCP 连接地址,可以作为增强路径复用当前浏览器会话;否则默认由仓库自举 Playwright Chromium。
对应资源:
tools/render_mermaid.pytools/ensure_thesis_assets.pytools/extract_screenshot_placeholders.pytools/build_screenshot_plan.pytools/capture_thesis_screenshots.py
6. 参考文献先建池再回填
默认约束:
- 2020 年及以后
- 中文 10-12 篇
- 英文 3-5 篇
- 总数约 15 篇
必须优先真实可核验文献,不确定就不用。
文献工作流必须分 5 步执行:
1. 建候选池 2. 做真实性核验 3. 做相关性筛选 4. 格式化参考文献 5. 生成文献核验清单
默认文献质量规则:
- 中文优先使用 CNKI / 万方 / 维普中可核验的北大核心、CSCD 等来源
- 英文优先使用 IEEE、Elsevier、Springer、ACM、MDPI 等可核验 DOI 来源
- 优先近五年文献
- 优先被引较高文献
- 不确定真实性的条目直接丢弃,不写“猜测型引用”
文献交付产物默认至少包括:
- 正文参考文献列表
references-verified.json文献核验清单
文献核验清单默认字段:
- title
- authors
- year
- source
- doi_or_url
- citation_count_if_available
- relevance_note
- status
对应资源:
prompts/reference_selector.mdtools/build_reference_pool.py
7. DOCX 成稿交付
如果环境具备 doc / docx 能力,必须生成 .docx 成稿,而不是只停留在 Markdown。
交付物必须至少包括:
1. 主论文 .md 2. 主论文 .docx 3. 图像映射文件 4. 附件 .docx
文件名规则:
- 主论文文件名必须使用论文标题,如
智慧实验室管理系统的设计与实现.docx - 附件文件名必须使用论文标题加“附件”,如
智慧实验室管理系统的设计与实现-附件.docx - 禁止使用
smart-lab-thesis、paper-final、doc1等通用名
派生文件也应统一命名:
论文标题.md论文标题.docx论文标题-附件.docx论文标题-image-map.json论文标题-文献核验清单.json
附件 .docx 默认内容:
- 正文中的 Mermaid / PlantUML 源码
- 数据库 E-R 图源码
- 关键流程图源码
- 必要时补充核心 SQL / 接口结构说明
默认样式规则:
- 摘要、Abstract、参考文献、致谢标题居中
- 摘要与 Abstract 独立分页
- 一级章节分页开始
- 中文正文宋体
- 英文正文 Times New Roman
- 中文关键词单独成段,顶格,“关键词:”标签使用黑体小四加粗,内容使用宋体小四
- 中英文摘要正文除关键词行外,默认首行缩进 2 字符
- 英文摘要正文使用 Times New Roman 小四
- 英文关键词单独成段,顶格,不首行缩进,使用 Times New Roman 小四
- 参考文献悬挂缩进
对应资源:
prompts/docx_formatter.mdtools/generate_thesis_docx.pytools/analyze_docx.pyreferences/default-style.md
8. 最终检查
交付前必须检查:
- 章节完整
- 字数接近样文目标
- 参考文献比例正确
- 图表编号连续
- 是否残留占位符
.docx是否真实存在- 主论文文件名是否为论文标题
- 附件
.docx是否真实存在 - 第 4 章是否为全文最长章节或接近最长章节
- 第 4 章每个主要模块是否包含截图
- 第 4 章每个主要模块是否包含简要核心代码
- 第 3 章设计图数量是否达标
- E-R 图是否按需要拆分为总表图和核心表图
- 文献是否全部在目标时间范围内
- 是否还混入低可信旧说明文档中的事实
Style Guardrails
正文语言默认遵循以下风格约束:
1. 模仿样文的章节推进节奏,不机械模仿原句。 2. 优先写项目事实,不先写空泛结论。 3. 避免连续堆砌“具有重要意义”“实现了良好效果”“具有较高价值”这类空话。 4. 避免过密排比句、模板化总分总和明显 AI 套话。 5. 每章至少应有一部分直接对应源码中的真实模块、真实规则或真实数据结构。 6. 致谢可适度更自然、更有人情味,但正文不能过度口语化。
最终检查细则见:
prompts/final_checker.md
Resource Map
- 项目输入与冲突决策:
prompts/intake.md - 样文结构分析:
prompts/sample_analyzer.md - 样式提取:
prompts/style_extractor.md - 项目事实提取:
prompts/fact_extractor.md - 章节写作与控字:
prompts/chapter_writer.md - 参考文献筛选:
prompts/reference_selector.md - Word 格式化:
prompts/docx_formatter.md - 最终检查:
prompts/final_checker.md
- 默认版式:
references/default-style.md - 章节模式:
references/chapter-patterns.md
- 统计章节字数:
tools/count_chapter_words.py - 分析样文 PDF:
tools/analyze_sample_pdf.py - 分析样文 DOCX:
tools/analyze_docx.py - 检查参考文献池:
tools/build_reference_pool.py - 生成文献核验清单模板:
tools/write_reference_verification_template.py - 图表与截图补全检查:
tools/ensure_thesis_assets.py - 提取截图占位符:
tools/extract_screenshot_placeholders.py - 生成图源码附件 DOCX:
tools/generate_diagram_appendix_docx.py - 生成截图计划:
tools/build_screenshot_plan.py - 自动抓取页面截图:
tools/capture_thesis_screenshots.py - 渲染 Mermaid:
tools/render_mermaid.py - 生成 DOCX:
tools/generate_thesis_docx.py
你是一个专门负责中文论文写作与交付的子代理。
工作要求:
1. 先遵循 ../SKILL.md 2. 写作时优先读取:
../prompts/*.md../references/*.md../tools/*.py
3. 默认输出中文 4. 不要凭空补全项目事实 5. 不要写得明显超出样文体量 6. Word 成稿前必须检查图表、参考文献和截图是否闭环 7. 样文是 .docx 时必须先运行 ../tools/analyze_docx.py 8. 最终 .docx 不能残留 **、反引号或 Markdown 链接 9. 如果存在截图占位,优先调用仓库内置 Playwright 截图链路,而不是依赖外部 skill
使用仓库中的 lunwen 技能完成论文任务。
执行要求:
1. 先读取 SKILL.md 2. 按其中的主流程执行 3. 需要细则时按 prompts/、references/、tools/ 中的资源展开 4. 默认输出中文 5. 如果存在样文、模板、项目代码、参考文献约束、截图需求或 .docx 要求,全部纳入同一工作流 6. 如果样文是 .docx,先运行 tools/analyze_docx.py 生成样式配置,再开始正文 7. 写作期间必须用 tools/count_chapter_words.py 控字,并用 tools/ensure_thesis_assets.py 检查图表/截图闭环 8. 如果存在截图占位,优先用 tools/build_screenshot_plan.py + tools/capture_thesis_screenshots.py 生成 image-map.json
ARGUMENTS: $ARGUMENTS
__pycache__/
*.py[cod]
node_modules/
tools/browser/node_modules/
tools/browser/.cache/
你是一个专门负责中文论文写作与交付的子代理。
工作要求:
1. 先遵循 ../SKILL.md 2. 写作时优先读取:
../prompts/*.md../references/*.md../tools/*.py
3. 默认输出中文 4. 不要凭空补全项目事实 5. 不要写得明显超出样文体量 6. Word 成稿前必须检查图表、参考文献和截图是否闭环 7. 样文是 .docx 时必须先运行 ../tools/analyze_docx.py 8. 最终 .docx 不能残留 **、反引号或 Markdown 链接 9. 如果存在截图占位,优先调用仓库内置 Playwright 截图链路,而不是依赖外部 skill
使用仓库中的 lunwen 技能完成论文任务。
执行要求:
1. 先读取 SKILL.md 2. 按其中的主流程执行 3. 需要细则时按 prompts/、references/、tools/ 中的资源展开 4. 默认输出中文 5. 如果存在样文、模板、项目代码、参考文献约束、截图需求或 .docx 要求,全部纳入同一工作流 6. 如果样文是 .docx,先运行 tools/analyze_docx.py 生成样式配置,再开始正文 7. 写作期间必须用 tools/count_chapter_words.py 控字,并用 tools/ensure_thesis_assets.py 检查图表/截图闭环 8. 生成 .docx 时必须传入 --style-profile,且最终文档不能残留 **、反引号或 Markdown 链接 9. 如果存在截图占位,优先用 tools/build_screenshot_plan.py + tools/capture_thesis_screenshots.py 生成 image-map.json
ARGUMENTS: $ARGUMENTS
interface:
display_name: "Lunwen"
short_description: "Write Chinese theses, learn sample-paper structure, manage references, figures, screenshots, and DOCX delivery."
default_prompt: "Use this skill to write or revise a Chinese thesis or graduation project paper based on real project facts, sample papers, templates, references, figures, screenshots, and DOCX delivery. Before writing, ask for local file paths for template/sample/opening report. If a sample or template is DOCX, run `python tools/analyze_docx.py <sample.docx> --json-out output/style-profile.json`. If a sample is PDF, run `python tools/analyze_sample_pdf.py <sample.pdf> --json-out output/sample-analysis.json`. After analysis, summarize directory, target chapter lengths, and extracted styles, then explicitly wait for user confirmation before writing. During writing, run `python tools/count_chapter_words.py <thesis.md>` to control length, and run `python tools/ensure_thesis_assets.py <thesis.md> --check-only` to verify architecture figure, E-R figure, flow figure, data table, test-case table, and screenshot placeholders exist. When screenshot placeholders exist, use the repository-bundled browser flow: `python tools/extract_screenshot_placeholders.py <thesis.md> --json-out labels.json`, `python tools/build_screenshot_plan.py labels.json output/screenshot-plan.json --base-url <system-url>`, then `python tools/capture_thesis_screenshots.py output/screenshot-plan.json` to produce `image-map.json`. Before DOCX delivery, remove Markdown markers like `**` and links from正文, then generate with `python tools/generate_thesis_docx.py <source.md> <target.docx> --style-profile output/style-profile.json --image-map image-map.json`. Do not silently skip missing images: keep placeholders or warnings in the final DOCX."
lunwen PRD
目标
将中文论文写作 skill 从“单文件规则集合”升级为“可开源、可维护、可复用的仓库型 skill”。
核心用户
- 需要写毕业论文的开发者
- 需要根据真实项目生成技术报告的用户
- 需要按样文或模板稳定交付
.docx的用户
关键问题
1. 初稿容易写得过厚 2. 样文只学结构,不学版式 3. 参考文献比例容易失控 4. Mermaid / PlantUML 图在 Word 中可能只是占位 5. 截图抓取与 Word 交付没有形成稳定闭环
产品目标
1. 支持项目事实抽取 2. 支持样文与模板结构/样式双学习 3. 支持章节控字 4. 支持参考文献约束检查 5. 支持图表与截图闭环 6. 支持最终 .docx 成稿
示例大纲
- 摘 要
- Abstract
- 1 绪论
- 2 开发环境及技术方案
- 3 需求分析
- 4 系统设计
- 5 系统详细设计与实现
- 6 系统测试
- 7 总结与展望
- 参 考 文 献
- 致 谢
{
"base_url": "http://127.0.0.1:3000",
"cdp_url": "",
"output_dir": "output/doc",
"image_map_output": "image-map.json",
"headless": true,
"viewport": {
"width": 1440,
"height": 900
},
"entries": [
{
"label": "系统首页",
"url": "/",
"wait_for_selector": "body",
"wait_for_text": "",
"clip_selector": "",
"filename": "系统首页.png",
"full_page": true,
"actions": []
},
{
"label": "核心功能页面",
"url": "/dashboard",
"wait_for_selector": "[data-page='dashboard']",
"wait_for_text": "",
"clip_selector": "",
"filename": "核心功能页面.png",
"full_page": true,
"actions": []
}
]
}
安装说明
本地技能目录
推荐安装到:
C:\Users\Administrator\.codex\skills\lunwen
或其他 $CODEX_HOME/skills 可发现路径。
Claude Code 兼容
Claude Code 官方支持 SKILL.md 技能目录,也兼容 .claude/commands/ 与 .claude/agents/。
本仓库已经内置:
.claude/commands/lunwen.md.claude/agents/lunwen-writer.md.trae/commands/lunwen.md.trae/agents/lunwen-writer.md
如果你在 Trae 或其他只读取入口 prompt 的环境中使用,建议优先保证宿主至少能读取:
agents/openai.yamltools/analyze_docx.pytools/build_screenshot_plan.pytools/capture_thesis_screenshots.pytools/ensure_thesis_assets.pytools/generate_thesis_docx.py
否则容易出现“读到技能名字,但没执行完整工作流”的情况。
因此有两种接入方式:
1. 作为技能目录放入 .claude/skills/lunwen/ 2. 作为项目仓库放在根目录,直接让 Claude Code 读取其中 .claude/commands/ 和 .claude/agents/
依赖建议
论文文字与 Word 交付建议具备以下能力:
python-docxpdfplumberpypdfNode.js与npm(用于仓库内置 Playwright 自举)
如果需要将 mermaid / plantuml 渲染为真实图片,建议环境额外具备:
@mermaid-js/mermaid-cliplantuml.jar或等效渲染方案
如果需要对 .docx 做逐页渲染检查,建议环境具备:
sofficepdftoppm
浏览器截图自举
仓库内置了 Playwright 执行链路,首次运行:
python tools/capture_thesis_screenshots.py output/screenshot-plan.json会自动在仓库目录执行:
npm install
npm run install:browsers然后再开始截图。用户不需要额外安装浏览器 skill。
编码规则
- 所有 Markdown、YAML、Prompt、脚本文件统一使用
UTF-8 - 不要依赖系统默认编码
- 校验脚本应显式使用
utf-8读取文件
MIT License
Copyright (c) 2026 Doryoku1223
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
{
"name": "lunwen-skill",
"private": true,
"version": "0.1.0",
"description": "Self-contained browser automation helpers for thesis screenshots.",
"scripts": {
"install:browsers": "playwright install chromium",
"capture:screenshots": "node tools/browser/capture_screenshots.mjs"
},
"dependencies": {
"playwright": "^1.53.0"
}
}
Chapter Writer Prompt
用于按样文章节节奏写作。
要求:
- 先看样文对应章节字数
- 控制当前章节篇幅
- 语言风格贴近本科软件类论文
- 第 4 章或“系统详细设计与实现”默认必须是全文最长章节
- 系统详细设计与实现默认按“模块说明 -> 流程图/结构图 -> 关键实现 -> 页面截图”展开
- 系统详细设计与实现中的每个主要模块至少包含 1 张页面截图
- 系统详细设计与实现中的每个主要模块至少包含 1 段简要核心代码、SQL 或关键实现片段
- 第 3 章默认至少包含 5 张设计图
- 第 3 章数据库设计默认至少包含“总表设计图 + 核心表设计图”两张 E-R 图;如果总图过大,必须拆分
- 如果任务书要求与源码实现不一致,正文必须优先按源码事实写,并在相关段落中保持口径一致
- 写完立即统计字数,优先使用
python tools/count_chapter_words.py <thesis.md> - 超出就压缩
- 不能残留
**、反引号、Markdown 链接等标记 - 如果正文缺少架构图、E-R 图、关键流程图、数据表、测试用例表或页面截图占位,必须执行
python tools/ensure_thesis_assets.py <thesis.md> --check-only 并在继续写作前补齐缺失项
- 数据库设计部分必须出现 E-R 图
- 若页面截图总数少于 6 张,默认继续补图,不得直接交付
- 参考文献必须先核验后回填,不能边猜边写
- 文风应贴近本科论文样文,但避免空泛套话和机械排比
- 如果正文存在截图占位,优先执行:
python tools/extract_screenshot_placeholders.py <thesis.md> --json-out labels.json python tools/build_screenshot_plan.py labels.json output/screenshot-plan.json --base-url <system-url> python tools/capture_thesis_screenshots.py output/screenshot-plan.json
- 如果正文存在 Mermaid / PlantUML 图,完成主文稿后必须额外生成附件
.docx,收录这些图的源码版本
DOCX Formatter Prompt
用于将论文源稿转成 .docx。
默认样式:
- 摘要、Abstract、参考文献、致谢标题居中
- 摘要与 Abstract 各自独立分页
- 一级标题分页开始
- 中文正文宋体
- 英文正文 Times New Roman
- 中文关键词单独成段,顶格,不首行缩进,“关键词:”标签黑体小四加粗,内容宋体小四
- 中英文摘要正文除关键词行外,保持正常首行缩进 2 字符
- 英文摘要正文 Times New Roman 小四
- 英文关键词单独成段,顶格,不首行缩进,
Keywords:标签加粗且为 Times New Roman 小四,内容为 Times New Roman 小四 - 图题、表题均使用单倍行距,段前段后为 0
- 参考文献悬挂缩进
必须优先让生成脚本真实实现这些规则,而不是只在说明里写规则。
如果存在截图或渲染图,优先嵌入真实图片。
生成时优先使用:
python tools/generate_thesis_docx.py thesis.md thesis.docx --style-profile output/style-profile.json --image-map image-map.json
导出前必须确认正文已经清理掉 Markdown 痕迹,例如:
**加粗**- `
code` [链接文字](https://example.com)
这些都不能原样进入最终 .docx。
Fact Extractor Prompt
用于从真实项目中提炼论文事实底稿。
输出必须固定:
- 角色集合
- 核心模块
- 关键业务闭环
- 关键数据表
- 关键业务规则
- 当前已完成能力
后续章节只能基于这份底稿扩写。
Final Checker Prompt
交付前必须逐项检查:
1. 章节完整 2. 字数接近样文目标,优先核对 python tools/count_chapter_words.py <thesis.md> 的 APPROX_WORDS 3. 参考文献比例符合要求 4. 图表编号连续 5. 是否残留占位符 6. .docx 是否真实存在 7. 摘要 / Abstract / 参考文献 / 致谢样式是否正确 8. 英文字体是否正确 9. 图表是否真实插入 10. 最终正文是否还残留 **、反引号或 Markdown 链接 11. 主论文文件名是否使用论文标题 12. 附件 .docx 是否真实存在 13. 第 4 章或“系统详细设计与实现”是否为全文最长章节或接近最长章节 14. 第 4 章每个主要模块是否包含截图 15. 第 4 章每个主要模块是否包含简要核心代码 16. 数据库部分是否包含 E-R 图 17. 第 3 章设计图数量是否达到 5 张及以上 18. E-R 图是否在需要时拆分为总表图和核心表图 19. 是否生成 references-verified.json 或等价文献核验清单 20. 文献是否全部在目标时间范围内 21. 是否混入了低可信旧说明文档中的事实 22. 页面截图总数是否达到 6 张及以上
最终检查时,必须至少执行:
python tools/count_chapter_words.py <thesis.md>python tools/ensure_thesis_assets.py <thesis.md> --check-only- 如果存在截图占位,还必须确认
image-map.json已由python tools/capture_thesis_screenshots.py <plan.json>生成 - 如果存在 Mermaid / PlantUML 图,还必须确认附件
.docx已生成 - 还必须确认文献核验清单文件已生成
Intake Prompt
用于锁定用户输入约束。
硬约束:
- 第一次回复时,先收资料,不开写初稿。
- 若用户尚未提供模板、样文、任务书、开题报告等辅助资料,必须优先索要,除非用户明确说“没有”或“按项目自己写”。
- 如果已经给了样文或模板,必须先分析,再回传目录和样式,不得跳过分析直接写正文。
- 输入优先级默认是:学校模板 > 任务书/开题报告 > 用户明确要求 > 样文 > 项目源码 > README > 旧说明文档。
- 旧部署说明、历史项目介绍、演示文档默认视为低可信来源。
必须确认:
1. 论文类型 2. 课题名称 3. 是否有学校模板 4. 是否有往届样文 5. 是否有字数要求 6. 是否必须交付 .docx 7. 是否有开题报告、任务书或封面字段要求 8. 如果需要真实页面截图,系统访问地址或启动方式是什么 9. 主论文文件名是否必须使用论文标题 10. 是否需要额外附件 .docx
首次回复时,必须明确告诉用户可以直接发送本地路径,示例:
D:\论文模板.docxD:\历届样文.pdfD:\开题报告.docx
如果用户一次没给全,不要连续抛很多问题。优先确认:
1. 是否有模板或样文 2. 是否有开题报告或任务书 3. 是否有字数要求 4. 是否最终要 Word 5. 如果需要截图,系统如何打开 6. 论文标题是什么
拿到模板、样文或开题报告后,不要立刻开写,必须先分析:
1. 目录结构是否完整 2. 各章大致字数 3. 标题、正文、摘要、关键词、图题、表题、表格内容的样式 4. 图、表、代码通常出现在什么章节
如果样文或模板是 .docx,必须优先执行:
python tools/analyze_docx.py <sample.docx> --json-out output/style-profile.json
如果样文是 PDF,至少执行:
python tools/analyze_sample_pdf.py <sample.pdf> --json-out output/sample-analysis.json
分析完成后,必须先回给用户:
1. 输入资料表 2. 当前建议目录表 3. 各章字数预算表 4. 版式样式与冲突表 5. 文件命名方案 6. 是否额外生成附件 .docx
并等待用户确认或修改。
Reference Selector Prompt
用于筛选参考文献。
默认约束:
- 2020 年及以后
- 中文 10-12 篇
- 英文 3-5 篇
- 总数约 15 篇
必须优先真实可核验来源:
- DOI
- 出版社官网
- 期刊官网
- 学校数据库公开页面
不确定就不用。
Sample Analyzer Prompt
用于分析往届样文。
必须提取:
- 目录层级
- 各章页数
- 各章大致字数
- 图表分布密度
- 代码放在正文还是附录
- 第 5 章模块展开节奏
- 图、表、代码分别插入在哪些章节
- 摘要、Abstract、参考文献、致谢是否单独分页
- 目录中一二三级标题的命名习惯
- 样文的语言风格与常见表达句式
输出格式必须包含:
1. 样文章节对照表 2. 当前论文建议目标字数表 3. 可直接模仿的表达习惯 4. 当前论文建议目录
如果用户同时给了多个样文,必须比较它们的共同结构与差异,并说明最终建议采用哪一套目录节奏。
Style Extractor Prompt
用于提取模板或样文版式样式。
必须关注:
- 一级、二级、三级标题
- 正文
- 摘要标题
- Abstract 标题
- 关键词格式
- 图题表题
- 表格内容字体字号
- 代码块或代码截图的标题与正文样式
- 参考文献
- 致谢
- 目录样式
- 是否分页
- 是否居中
- 字体、字号、加粗
- 段前段后
- 行间距
- 首行缩进
如果提取出的样式与技能默认规则冲突,必须生成“冲突清单”并提示用户选择。
输出时必须把样式按模块列清楚,至少包括:
1. 一级标题样式 2. 二级标题样式 3. 三级标题样式 4. 中文正文样式 5. 英文正文样式 6. 摘要与 Abstract 样式 7. 关键词样式 8. 图题、表题、表格内容样式 9. 参考文献与致谢样式
样式分析完成后,必须与建议目录一起回传用户确认。
如果样文是 .docx,不能只做口头总结,必须产出结构化样式文件:
python tools/analyze_docx.py <sample.docx> --json-out output/style-profile.json
后续生成 .docx 时,必须继续传入该样式文件,而不是退回固定默认样式。
<div align="center">
论文.skill
“论文不是把字堆满,而是把项目事实、样文规范、图表截图和 Word 交付一次性闭环。”
面向本科计科学生推出的一键论文初稿 Skill。
不会写毕业论文初稿,开题报告和项目代码不知道怎么落成正文? 历届样文、学校模板、格式要求很多,但拼不出完整结构? 流程图、用例图、E-R 图、系统截图和测试内容总是缺一块? 最后要交 .docx,却还停留在零散材料和碎片化笔记里?
把零散材料快速整理成论文初稿的 Skill。 可以结合历届样文、开题报告、学校模板和真实项目内容生成初稿, 自动补充流程图、用例图、E-R 图等论文图表,并补齐项目截图与常见配图位,形成一套可继续精修的论文初稿工作流。
安装 · 快速使用方式 · 使用流程 · 当前能力 · 常用脚本 · 详细安装说明 · GitHub
</div>
安装
方式 1:使用 Codex skill-installer 直接安装
python "${HOME}\.codex\skills\.system\skill-installer\scripts\install-skill-from-github.py" --repo Doryoku1223/lunwen-skill --path . --name lunwen方式 2:手动克隆到本地 skills 目录
git clone https://github.com/Doryoku1223/lunwen-skill.git "${HOME}\.codex\skills\lunwen"安装后重启 Codex,即可让它发现这个 skill。
如果你使用 Claude Code,也可以把仓库放到项目中直接复用其中的 .claude/commands/ 和 .claude/agents/ 兼容层,详细说明见 INSTALL.md。
仓库同时补充了 .trae/commands/ 与 .trae/agents/ 兼容包装,方便在 Trae 这类只读取入口命令/代理定义的环境里尽量完整传递工作流。
快速使用方式
推荐优先在 Codex 中使用本 skill。
最简单的使用方法:
1. 把本项目 GitHub 链接复制给 Codex,让它帮你安装这个 skill。 2. 安装成功后,在 Codex 中打开你的项目文件夹。 3. 直接对 Codex 说:为这个项目写一篇论文 4. 然后按提示继续提供学校模板、历届样文、开题报告或任务书路径。
建议你在第一次使用时优先提供:
- 学校论文模板
- 历届样文
- 开题报告或任务书
- 项目源码目录
这样 Codex 才能按“先分析模板,再确认目录样式,最后生成 .docx 成稿”的正确流程工作。
推荐搭配 Skills
以下是这次实战迭代中证明有价值的搭配:
lunwen
论文主流程,负责模板优先分析、目录设计、正文写作、图表与成稿交付
docx
用于读取、分析和校验 .docx 模板与成稿,尤其适合处理样式、结构和文档对象
planning-with-files
用于长流程任务管理,适合论文这类多阶段任务,避免在模板分析、写作、成稿之间丢状态
systematic-debugging
当成稿出现目录重复、标题编号叠加、图片覆盖正文、附录样式错误等问题时,用它来定位根因,而不是盲改
verification-before-completion
在声称“论文已经生成好”之前,强制执行字数、引用、目录、图片和 .docx 存在性检查
如果论文工作流中还涉及以下任务,也建议按需搭配:
pdf
当学校只给 PDF 模板、PDF 样文或需要逐页视觉检查时使用
doc
当输入是 .doc 而不是 .docx,且需要先转换后再分析时使用
playwright
当需要抓真实系统截图替换论文占位图时使用
当前能力
- 样文目录与字数分析
- 样文 DOCX 样式提取与结构化样式配置
- 模板/样文样式提取
- 项目事实底稿抽取
- 按样文章节节奏控字写作
- 中文/英文参考文献比例控制
- Mermaid / PlantUML 图表策略
- 内置 Playwright / Chrome CDP 浏览器自动截图策略
- doc / docx Word 成稿交付策略
推荐目录结构
lunwen/
├── SKILL.md
├── README.md
├── INSTALL.md
├── agents/
│ └── openai.yaml
├── prompts/
├── references/
├── tools/
├── docs/
└── examples/设计原则
1. 项目事实优先,不凭空补模块。 2. 样文章节体量优先,不默认写厚。 3. 学样文不只学内容,也学样式。 4. 图表、截图、参考文献和 Word 交付都要闭环。 5. 最终对用户展现过程与结果默认使用中文。
推荐工作流
建议按下面顺序使用本 skill:
1. 读取项目代码与文档,冻结项目事实底稿。 2. 分析样文或模板,提取章节体量和版式样式。 3. 生成目标章节字数表。 4. 分章写作并持续控字。 5. 构建参考文献池并检查中英文比例。 6. 处理 Mermaid / PlantUML 图表。 7. 抓取真实系统截图替换占位。 8. 生成 .docx。 9. 做最终检查。
常用脚本
1. 统计论文各章字数
python tools/count_chapter_words.py thesis.md2. 分析样文 PDF
python tools/analyze_sample_pdf.py sample.pdf3. 检查参考文献池
python tools/build_reference_pool.py thesis.md4. 分析样文 DOCX 并生成样式配置
python tools/analyze_docx_styles.py sample.docx output/style-profile.json5. 提取截图占位并生成截图计划
python tools/extract_screenshot_placeholders.py thesis.md --json-out labels.json
python tools/build_image_map.py labels.json output/doc image-map.json6. 提取 Mermaid 图块
python tools/extract_mermaid_blocks.py thesis.md tmp/mermaid --manifest tmp/mermaid/manifest.json7. 渲染 Mermaid
python tools/render_mermaid.py tmp/mermaid/diagram-01.mmd tmp/mermaid/diagram-01.png8. 生成 DOCX
python tools/generate_thesis_docx.py thesis.md thesis.docx --style-spec output/style-profile.json --image-map image-map.json当前限制
- 如果环境没有 LibreOffice / Poppler,无法做逐页渲染检查。
- 如果环境没有 Mermaid 渲染能力,流程图可能只能保留源码或占位。
- 浏览器截图与真实系统抓图能力依赖额外环境支持。
兼容性
- Codex:原生适配。
- Claude Code:已补
.claude/commands/与.claude/agents/兼容层。
章节模式
第 1 章
- 研究背景
- 国内外研究现状
- 研究目的和意义
- 论文结构
第 5 章
每个模块默认结构:
1. 模块功能说明 2. 流程图 3. 关键实现代码 4. 页面截图
第 6 章
- 测试环境
- 测试方法
- 功能测试用例
- 构建验证结果
- 测试结果分析
默认论文样式
在没有学校模板时使用以下默认规则:
- 一级标题:黑体小 2 号加粗
- 二级标题:黑体小 3 号
- 三级标题:黑体小 4 号
- 中文正文:宋体 5 号
- 英文正文:Times New Roman
- 正文:1.25 倍行距,首行缩进 2 字符
- 摘要、Abstract、参考文献、致谢标题居中
- 摘要与 Abstract 各占一页
- 中文关键词:单独成段,顶格,不首行缩进;“关键词:”标签使用黑体小 4 号加粗,关键词内容使用宋体小 4 号
- 中英文摘要正文:除关键词行外,首行缩进 2 字符
- 英文摘要正文:Times New Roman 小 4 号
- 英文关键词:单独成段,顶格,不首行缩进;
Keywords:标签使用 Times New Roman 小 4 号加粗,内容使用 Times New Roman 小 4 号 - 图题图下方居中,单倍行距,段前段后 0
- 表题表上方居中,单倍行距,段前段后 0
- 参考文献悬挂缩进 2 字符
pypdf
pdfplumber
python-docx
from __future__ import annotations
import json
import re
import sys
from collections import Counter, defaultdict
from pathlib import Path
from typing import Any
from docx import Document
from docx.enum.text import WD_ALIGN_PARAGRAPH
from docx.oxml.ns import qn
TITLE_PATTERNS = {
"heading1": re.compile(r"^第.+章"),
"heading2": re.compile(r"^\d+\.\d+(?!\.)"),
"heading3": re.compile(r"^\d+\.\d+\.\d+"),
}
def alignment_name(value: int | None) -> str:
mapping = {
None: "left",
WD_ALIGN_PARAGRAPH.LEFT: "left",
WD_ALIGN_PARAGRAPH.CENTER: "center",
WD_ALIGN_PARAGRAPH.RIGHT: "right",
WD_ALIGN_PARAGRAPH.JUSTIFY: "justify",
WD_ALIGN_PARAGRAPH.DISTRIBUTE: "distribute",
}
return mapping.get(value, "left")
def east_asia_font(run) -> str | None:
if run._element.rPr is None or run._element.rPr.rFonts is None:
return None
return run._element.rPr.rFonts.get(qn("w:eastAsia"))
def paragraph_signature(paragraph) -> dict[str, Any]:
run_fonts = []
run_sizes = []
bold_values = []
for run in paragraph.runs:
if not run.text.strip():
continue
run_fonts.append((east_asia_font(run) or run.font.name or "").strip())
if run.font.size:
run_sizes.append(round(run.font.size.pt, 1))
if run.bold is not None:
bold_values.append(bool(run.bold))
first_line_indent = paragraph.paragraph_format.first_line_indent
line_spacing = paragraph.paragraph_format.line_spacing
page_break_before = False
p_pr = paragraph._p.pPr
if p_pr is not None and p_pr.pageBreakBefore is not None:
page_break_before = True
return {
"font": most_common(run_fonts),
"size_pt": most_common(run_sizes),
"bold": most_common(bold_values, default=False),
"alignment": alignment_name(paragraph.alignment),
"first_line_indent_pt": round(first_line_indent.pt, 1) if first_line_indent else 0.0,
"line_spacing": round(float(line_spacing), 2) if isinstance(line_spacing, (int, float)) else None,
"space_before_pt": round(paragraph.paragraph_format.space_before.pt, 1)
if paragraph.paragraph_format.space_before
else 0.0,
"space_after_pt": round(paragraph.paragraph_format.space_after.pt, 1)
if paragraph.paragraph_format.space_after
else 0.0,
"page_break_before": page_break_before,
}
def most_common(values: list[Any], default: Any = None) -> Any:
filtered = [value for value in values if value not in ("", None)]
if not filtered:
return default
return Counter(filtered).most_common(1)[0][0]
def classify_paragraph(text: str, in_reference_section: bool) -> str | None:
normalized = text.strip()
if not normalized:
return None
if normalized == "摘要":
return "abstract_heading_cn"
if normalized.lower() == "abstract":
return "abstract_heading_en"
if normalized == "参考文献":
return "references_heading"
if normalized == "致谢":
return "ack_heading"
if normalized.startswith("关键词"):
return "keywords_cn"
if normalized.startswith("Keywords"):
return "keywords_en"
if normalized.startswith("图"):
return "figure_caption"
if normalized.startswith("表"):
return "table_caption"
if in_reference_section and re.match(r"^\[\d+\]", normalized):
return "references_body"
if TITLE_PATTERNS["heading3"].match(normalized):
return "heading3"
if TITLE_PATTERNS["heading2"].match(normalized):
return "heading2"
if TITLE_PATTERNS["heading1"].match(normalized):
return "heading1"
if re.search(r"[A-Za-z]{3,}", normalized) and not re.search(r"[\u4e00-\u9fff]", normalized):
return "body_en"
return "body_cn"
def aggregate_styles(document: Document) -> dict[str, Any]:
grouped: dict[str, list[dict[str, Any]]] = defaultdict(list)
in_reference_section = False
for paragraph in document.paragraphs:
text = paragraph.text.strip()
if text == "参考文献":
in_reference_section = True
elif re.match(r"^第.+章", text):
in_reference_section = False
category = classify_paragraph(text, in_reference_section)
if category is None:
continue
grouped[category].append(paragraph_signature(paragraph))
result = {}
for category, signatures in grouped.items():
result[category] = {
"font": most_common([item["font"] for item in signatures]),
"size_pt": most_common([item["size_pt"] for item in signatures]),
"bold": most_common([item["bold"] for item in signatures], default=False),
"alignment": most_common([item["alignment"] for item in signatures], default="left"),
"first_line_indent_pt": most_common([item["first_line_indent_pt"] for item in signatures], default=0.0),
"line_spacing": most_common([item["line_spacing"] for item in signatures]),
"space_before_pt": most_common([item["space_before_pt"] for item in signatures], default=0.0),
"space_after_pt": most_common([item["space_after_pt"] for item in signatures], default=0.0),
"page_break_before": most_common([item["page_break_before"] for item in signatures], default=False),
"sample_count": len(signatures),
}
return result
def print_summary(path: Path, styles: dict[str, Any]) -> None:
print(f"FILE\t{path}")
ordered_keys = [
"heading1",
"heading2",
"heading3",
"body_cn",
"body_en",
"abstract_heading_cn",
"abstract_heading_en",
"keywords_cn",
"keywords_en",
"figure_caption",
"table_caption",
"references_body",
]
for key in ordered_keys:
if key not in styles:
continue
style = styles[key]
print(
f"{key}\tfont={style['font']}\tsize_pt={style['size_pt']}\tbold={style['bold']}"
f"\talignment={style['alignment']}\tfirst_line_indent_pt={style['first_line_indent_pt']}"
f"\tline_spacing={style['line_spacing']}\tpage_break_before={style['page_break_before']}"
)
def main() -> int:
if len(sys.argv) < 2:
print("Usage: python analyze_docx.py <docx> [--json-out path]")
return 1
docx_path = Path(sys.argv[1])
document = Document(docx_path)
styles = aggregate_styles(document)
payload = {
"source_file": str(docx_path),
"styles": styles,
}
if len(sys.argv) >= 4 and sys.argv[2] == "--json-out":
output_path = Path(sys.argv[3])
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
print_summary(docx_path, styles)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import json
import re
import sys
from pathlib import Path
from pypdf import PdfReader
def clean_count(text: str) -> int:
return len(re.sub(r"\s+", "", text))
def walk_outline(items, reader, level=1, out=None):
if out is None:
out = []
for item in items:
if isinstance(item, list):
walk_outline(item, reader, level + 1, out)
else:
title = item.get("/Title", "").strip()
if not title:
continue
try:
page = reader.get_destination_page_number(item) + 1
except Exception:
continue
out.append({"level": level, "title": title, "page": page})
return out
def find_end_page(items, idx, total_pages):
cur = items[idx]
for j in range(idx + 1, len(items)):
if items[j]["level"] <= cur["level"]:
return max(cur["page"], items[j]["page"] - 1)
return total_pages
def analyze_pdf(path: Path) -> dict:
reader = PdfReader(str(path))
pages = [page.extract_text() or "" for page in reader.pages]
outline = walk_outline(reader.outline, reader)
sections = []
for idx, item in enumerate(outline):
end_page = find_end_page(outline, idx, len(pages))
text = "\n".join(pages[item["page"] - 1 : end_page])
sections.append(
{
"level": item["level"],
"title": item["title"],
"start_page": item["page"],
"end_page": end_page,
"char_count": clean_count(text),
}
)
return {
"file": str(path),
"total_pages": len(pages),
"sections": sections,
}
def main() -> int:
if len(sys.argv) < 2:
print("Usage: python analyze_sample_pdf.py <pdf> [--json-out path]")
return 1
pdf_path = Path(sys.argv[1])
result = analyze_pdf(pdf_path)
json_out = None
if len(sys.argv) >= 4 and sys.argv[2] == "--json-out":
json_out = Path(sys.argv[3])
if json_out:
json_out.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"FILE\t{result['file']}")
print(f"TOTAL_PAGES\t{result['total_pages']}")
for section in result["sections"]:
if section["level"] <= 2:
print(
f"{section['title']}\t"
f"p.{section['start_page']}-{section['end_page']}\t"
f"{section['char_count']}"
)
return 0
if __name__ == "__main__":
raise SystemExit(main())
import fs from "node:fs";
import path from "node:path";
import { chromium } from "playwright";
function resolveUrl(baseUrl, entryUrl) {
if (!entryUrl) {
return baseUrl;
}
if (/^https?:\/\//i.test(entryUrl)) {
return entryUrl;
}
if (!baseUrl) {
throw new Error(`Entry url "${entryUrl}" requires base_url in the plan.`);
}
return new URL(entryUrl, baseUrl).toString();
}
function slugify(label) {
const normalized = label.replace(/[\\/:*?"<>|]/g, "-").replace(/\s+/g, "-").trim();
return normalized || "screenshot";
}
async function waitForEntry(page, entry) {
if (entry.wait_for_timeout_ms) {
await page.waitForTimeout(entry.wait_for_timeout_ms);
}
if (entry.wait_for_text) {
await page.getByText(entry.wait_for_text, { exact: false }).first().waitFor({ state: "visible", timeout: entry.timeout_ms ?? 15000 });
}
if (entry.wait_for_selector) {
await page.waitForSelector(entry.wait_for_selector, { state: "visible", timeout: entry.timeout_ms ?? 15000 });
}
}
async function applyActions(page, entry) {
for (const action of entry.actions ?? []) {
switch (action.type) {
case "click":
await page.locator(action.selector).first().click();
break;
case "fill":
await page.locator(action.selector).first().fill(action.value ?? "");
break;
case "press":
await page.locator(action.selector).first().press(action.key);
break;
case "wait_for_selector":
await page.waitForSelector(action.selector, { state: action.state ?? "visible", timeout: action.timeout_ms ?? 15000 });
break;
case "wait_for_text":
await page.getByText(action.text, { exact: false }).first().waitFor({ state: "visible", timeout: action.timeout_ms ?? 15000 });
break;
case "wait_for_timeout":
await page.waitForTimeout(action.timeout_ms ?? 1000);
break;
default:
throw new Error(`Unsupported action type: ${action.type}`);
}
}
}
async function main() {
const planPath = process.argv[2];
if (!planPath) {
console.error("Usage: node tools/browser/capture_screenshots.mjs <plan.json>");
process.exit(1);
}
const absolutePlanPath = path.resolve(planPath);
const plan = JSON.parse(fs.readFileSync(absolutePlanPath, "utf8"));
const outputDir = path.resolve(path.dirname(absolutePlanPath), plan.output_dir ?? "output/doc");
fs.mkdirSync(outputDir, { recursive: true });
const browser = plan.cdp_url
? await chromium.connectOverCDP(plan.cdp_url)
: await chromium.launch({ headless: plan.headless !== false });
const context = plan.cdp_url
? browser.contexts()[0]
: await browser.newContext({
viewport: plan.viewport ?? { width: 1440, height: 900 },
ignoreHTTPSErrors: true
});
const page = context.pages()[0] ?? await context.newPage();
const imageMap = {};
try {
for (const entry of plan.entries ?? []) {
const targetUrl = resolveUrl(plan.base_url, entry.url);
await page.goto(targetUrl, { waitUntil: entry.wait_until ?? "networkidle", timeout: entry.timeout_ms ?? 30000 });
await applyActions(page, entry);
await waitForEntry(page, entry);
const filename = entry.filename ?? `${slugify(entry.label)}.png`;
const outputPath = path.resolve(outputDir, filename);
const locator = entry.clip_selector ? page.locator(entry.clip_selector).first() : null;
if (locator) {
await locator.screenshot({ path: outputPath, type: "png" });
} else {
await page.screenshot({ path: outputPath, fullPage: entry.full_page !== false, type: "png" });
}
imageMap[entry.label] = outputPath;
console.log(`CAPTURED\t${entry.label}\t${outputPath}`);
}
} finally {
const imageMapPath = path.resolve(path.dirname(absolutePlanPath), plan.image_map_output ?? "image-map.json");
fs.writeFileSync(imageMapPath, JSON.stringify(imageMap, null, 2), "utf8");
console.log(`IMAGE_MAP\t${imageMapPath}`);
if (plan.cdp_url) {
await browser.close();
} else {
await context.close();
await browser.close();
}
}
}
main().catch((error) => {
console.error(error.stack || String(error));
process.exit(1);
});
from __future__ import annotations
import json
import sys
from pathlib import Path
def main() -> int:
if len(sys.argv) < 4:
print("Usage: python build_image_map.py <labels.json> <image-dir> <output.json> [--manual manual-map.json]")
return 1
labels_path = Path(sys.argv[1])
image_dir = Path(sys.argv[2])
output_path = Path(sys.argv[3])
manual_path = None
if len(sys.argv) >= 6 and sys.argv[4] == "--manual":
manual_path = Path(sys.argv[5])
labels = json.loads(labels_path.read_text(encoding="utf-8")).get("labels", [])
manual_map = {}
if manual_path and manual_path.exists():
manual_map = json.loads(manual_path.read_text(encoding="utf-8"))
result = {}
for label in labels:
if label in manual_map:
manual_target = Path(manual_map[label])
if manual_target.exists():
result[label] = str(manual_target)
continue
candidates = [
image_dir / f"{label}.png",
image_dir / f"{label}.jpg",
image_dir / f"{label}.jpeg",
]
for candidate in candidates:
if candidate.exists():
result[label] = str(candidate)
break
output_path.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
print(output_path)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import json
import re
import sys
from collections import Counter
from pathlib import Path
REF_RE = re.compile(r"^\[(\d+)\]\s*(.+)$")
def classify_language(text: str) -> str:
return "zh" if re.search(r"[\u4e00-\u9fff]", text) else "en"
def extract_year(text: str) -> int | None:
years = re.findall(r"\b(20\d{2})\b", text)
return int(years[0]) if years else None
def parse_references(path: Path) -> list[dict]:
refs = []
for line in path.read_text(encoding="utf-8").splitlines():
line = line.strip()
m = REF_RE.match(line)
if not m:
continue
body = m.group(2).strip()
refs.append(
{
"raw": line,
"index": int(m.group(1)),
"language": classify_language(body),
"year": extract_year(body),
"has_doi": "doi" in body.lower(),
}
)
return refs
def main() -> int:
if len(sys.argv) < 2:
print("Usage: python build_reference_pool.py <reference-markdown-file>")
return 1
path = Path(sys.argv[1])
refs = parse_references(path)
lang_counts = Counter(ref["language"] for ref in refs)
bad_years = [ref for ref in refs if ref["year"] is None or ref["year"] < 2020]
print(f"TOTAL\t{len(refs)}")
print(f"ZH\t{lang_counts.get('zh', 0)}")
print(f"EN\t{lang_counts.get('en', 0)}")
print(f"BAD_YEAR\t{len(bad_years)}")
for ref in bad_years:
print(f"BAD\t{ref['raw']}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import json
import re
from pathlib import Path
def safe_filename(label: str) -> str:
return re.sub(r'[\\/:*?"<>|]+', "-", label).strip() or "screenshot"
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Build a screenshot capture plan from thesis screenshot placeholders.")
parser.add_argument("labels_json", type=Path)
parser.add_argument("output_plan", type=Path)
parser.add_argument("--base-url", dest="base_url", default="")
parser.add_argument("--output-dir", dest="output_dir", default="output/doc")
parser.add_argument("--image-map", dest="image_map_output", default="image-map.json")
parser.add_argument("--cdp-url", dest="cdp_url", default="")
return parser.parse_args()
def main() -> int:
args = parse_args()
payload = json.loads(args.labels_json.read_text(encoding="utf-8"))
labels = payload.get("labels", [])
entries = []
for label in labels:
entries.append(
{
"label": label,
"url": "",
"wait_for_selector": "",
"wait_for_text": "",
"clip_selector": "",
"filename": f"{safe_filename(label)}.png",
"full_page": True,
"actions": [],
}
)
plan = {
"base_url": args.base_url,
"cdp_url": args.cdp_url,
"output_dir": args.output_dir,
"image_map_output": args.image_map_output,
"headless": True,
"viewport": {"width": 1440, "height": 900},
"entries": entries,
}
args.output_plan.parent.mkdir(parents=True, exist_ok=True)
args.output_plan.write_text(json.dumps(plan, ensure_ascii=False, indent=2), encoding="utf-8")
print(args.output_plan)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import shutil
import subprocess
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
PACKAGE_JSON = REPO_ROOT / "package.json"
def npm_command() -> list[str]:
npm = shutil.which("npm")
if not npm:
raise RuntimeError("npm is required for Playwright bootstrap, but it was not found in PATH.")
return [npm]
def ensure_playwright_installed() -> None:
node_modules = REPO_ROOT / "node_modules" / "playwright"
if node_modules.exists():
return
subprocess.run(npm_command() + ["install"], cwd=REPO_ROOT, check=True)
subprocess.run(npm_command() + ["run", "install:browsers"], cwd=REPO_ROOT, check=True)
def run_capture(plan_path: Path) -> None:
node = shutil.which("node")
if not node:
raise RuntimeError("Node.js is required for browser capture, but it was not found in PATH.")
script = REPO_ROOT / "tools" / "browser" / "capture_screenshots.mjs"
subprocess.run([node, str(script), str(plan_path)], cwd=REPO_ROOT, check=True)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Capture thesis screenshots with the repository-bundled Playwright flow.")
parser.add_argument("plan_json", type=Path)
parser.add_argument("--skip-bootstrap", action="store_true")
return parser.parse_args()
def main() -> int:
args = parse_args()
if not PACKAGE_JSON.exists():
print("package.json not found; Playwright bootstrap is unavailable.")
return 1
try:
if not args.skip_bootstrap:
ensure_playwright_installed()
run_capture(args.plan_json)
except subprocess.CalledProcessError as exc:
print(f"Command failed with exit code {exc.returncode}")
return exc.returncode
except RuntimeError as exc:
print(str(exc))
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import re
import sys
from pathlib import Path
from markdown_utils import compute_text_metrics
CHAPTER_RE = re.compile(r"^##\s+(.+)$", flags=re.M)
def chapter_spans(text: str) -> list[tuple[str, int, int]]:
matches = list(CHAPTER_RE.finditer(text))
spans: list[tuple[str, int, int]] = []
for index, match in enumerate(matches):
start = match.end()
end = matches[index + 1].start() if index + 1 < len(matches) else len(text)
spans.append((match.group(1).strip(), start, end))
return spans
def format_metrics(label: str, metrics: dict[str, int]) -> str:
return (
f"{label}\t"
f"APPROX_WORDS={metrics['approx_word_count']}\t"
f"CHAR_NO_SPACES={metrics['char_no_spaces']}\t"
f"CHAR_WITH_SPACES={metrics['char_with_spaces']}\t"
f"CJK_CHARS={metrics['chinese_chars']}\t"
f"NON_CJK_WORDS={metrics['non_chinese_words']}\t"
f"EN_WORDS={metrics['english_words']}"
)
def main() -> int:
if len(sys.argv) < 2:
print("Usage: python count_chapter_words.py <markdown-file>")
return 1
path = Path(sys.argv[1])
text = path.read_text(encoding="utf-8")
print(format_metrics("TOTAL", compute_text_metrics(text)))
for title, start, end in chapter_spans(text):
print(format_metrics(title, compute_text_metrics(text[start:end])))
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import re
from pathlib import Path
REQUIRED_ITEMS = [
{
"id": "architecture",
"label": "系统总体架构图",
"kind": "figure",
"keywords": ["架构图", "总体架构", "系统架构"],
"template": """```mermaid
flowchart LR
U[用户] --> W[Web/桌面前端]
W --> S[业务服务层]
S --> D[(数据库)]
S --> F[文件存储/缓存]
```
图 {number} {label}""",
},
{
"id": "er",
"label": "数据库 E-R 图",
"kind": "figure",
"keywords": ["E-R 图", "ER图", "实体关系"],
"template": """```mermaid
erDiagram
USER ||--o{{ NOTE : creates
NOTE ||--o{{ TAG : tagged_with
USER {{
int id
string name
}}
NOTE {{
int id
string title
}}
TAG {{
int id
string name
}}
```
图 {number} {label}""",
},
{
"id": "flow",
"label": "关键业务流程图",
"kind": "figure",
"keywords": ["流程图", "业务流程", "关键流程"],
"template": """```mermaid
flowchart TD
A[进入系统] --> B[输入业务参数]
B --> C[执行核心处理]
C --> D[保存结果]
D --> E[反馈执行状态]
```
图 {number} {label}""",
},
{
"id": "data_table",
"label": "核心数据表设计",
"kind": "table",
"keywords": ["数据表", "表结构", "数据库设计"],
"template": """表 {number} {label}
| 表名 | 说明 | 关键字段 |
| --- | --- | --- |
| user | 用户信息表 | id, username, password |
| note | 业务主表 | id, title, content |
| tag | 分类标签表 | id, name |""",
},
{
"id": "test_table",
"label": "系统测试用例表",
"kind": "table",
"keywords": ["测试用例", "测试表", "功能测试"],
"template": """表 {number} {label}
| 用例编号 | 测试目标 | 输入/操作 | 预期结果 |
| --- | --- | --- | --- |
| TC-01 | 正常流程 | 输入合法数据并提交 | 系统提示成功 |
| TC-02 | 异常流程 | 输入缺失字段 | 系统给出校验提示 |""",
},
{
"id": "screenshot",
"label": "核心功能页面截图",
"kind": "screenshot",
"keywords": ["此处插入截图", "页面截图", "系统截图"],
"template": "[此处插入截图:系统首页]\n[此处插入截图:核心功能页面]",
},
]
FIGURE_RE = re.compile(r"^图\s*(\d+(?:\.\d+)?)", flags=re.M)
TABLE_RE = re.compile(r"^表\s*(\d+(?:\.\d+)?)", flags=re.M)
def next_number(text: str, pattern: re.Pattern[str]) -> str:
matches = pattern.findall(text)
if not matches:
return "1.1"
last = matches[-1]
if "." in last:
major, minor = last.split(".", 1)
if minor.isdigit():
return f"{major}.{int(minor) + 1}"
if last.isdigit():
return str(int(last) + 1)
return "1.1"
def analyze_missing_items(text: str) -> list[dict]:
missing = []
for item in REQUIRED_ITEMS:
if any(keyword in text for keyword in item["keywords"]):
continue
missing.append(item)
return missing
def append_templates(text: str, missing: list[dict]) -> str:
if not missing:
return text
figure_number = next_number(text, FIGURE_RE)
table_number = next_number(text, TABLE_RE)
blocks = ["", "## 图表与截图补全草稿", ""]
for item in missing:
if item["kind"] == "figure":
blocks.append(item["template"].format(number=figure_number, label=item["label"]))
major, minor = figure_number.split(".") if "." in figure_number else (figure_number, "0")
figure_number = f"{major}.{int(minor) + 1}"
elif item["kind"] == "table":
blocks.append(item["template"].format(number=table_number, label=item["label"]))
major, minor = table_number.split(".") if "." in table_number else (table_number, "0")
table_number = f"{major}.{int(minor) + 1}"
else:
blocks.append(item["template"])
blocks.append("")
return text.rstrip() + "\n" + "\n".join(blocks)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Ensure mandatory figures, tables and screenshots exist in thesis markdown.")
parser.add_argument("markdown_file", type=Path)
parser.add_argument("--check-only", action="store_true")
parser.add_argument("--in-place", action="store_true")
return parser.parse_args()
def main() -> int:
args = parse_args()
text = args.markdown_file.read_text(encoding="utf-8")
missing = analyze_missing_items(text)
print(f"MISSING_COUNT\t{len(missing)}")
for item in missing:
print(f"MISSING\t{item['id']}\t{item['label']}")
if args.check_only or not args.in_place or not missing:
return 2 if missing else 0
updated = append_templates(text, missing)
args.markdown_file.write_text(updated, encoding="utf-8")
print(f"UPDATED\t{args.markdown_file}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import json
import re
import sys
from pathlib import Path
CAPTION_RE = re.compile(r"^(图\s*\d+(?:\.\d+)?\s+.+)$")
def safe_name(text: str) -> str:
cleaned = re.sub(r"[^\w\u4e00-\u9fff\-]+", "-", text).strip("-")
return cleaned[:80] or "diagram"
def main() -> int:
if len(sys.argv) < 3:
print("Usage: python extract_mermaid_blocks.py <markdown-file> <output-dir> [--manifest path]")
return 1
source = Path(sys.argv[1])
out_dir = Path(sys.argv[2])
manifest_path = None
if len(sys.argv) >= 5 and sys.argv[3] == "--manifest":
manifest_path = Path(sys.argv[4])
out_dir.mkdir(parents=True, exist_ok=True)
lines = source.read_text(encoding="utf-8").splitlines()
in_mermaid = False
buffer: list[str] = []
pending_blocks: list[dict] = []
results: list[dict] = []
index = 0
for line in lines:
stripped = line.strip()
if in_mermaid:
if stripped.startswith("```"):
pending_blocks.append({"content": "\n".join(buffer)})
in_mermaid = False
buffer = []
else:
buffer.append(line.rstrip())
continue
if stripped.startswith("```mermaid"):
in_mermaid = True
buffer = []
continue
cap = CAPTION_RE.match(stripped)
if cap and pending_blocks:
block = pending_blocks.pop(0)
index += 1
caption = cap.group(1).strip()
filename = f"{index:02d}-{safe_name(caption)}.mmd"
file_path = out_dir / filename
file_path.write_text(block["content"], encoding="utf-8")
results.append(
{
"caption": caption,
"mmd": str(file_path),
"png": str(file_path.with_suffix(".png")),
"svg": str(file_path.with_suffix(".svg")),
}
)
if manifest_path:
manifest_path.parent.mkdir(parents=True, exist_ok=True)
manifest_path.write_text(json.dumps(results, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"COUNT\t{len(results)}")
for item in results:
print(f"{item['caption']}\t{item['mmd']}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import json
import re
import sys
from pathlib import Path
PLACEHOLDER_RE = re.compile(r"^\[此处插入截图:(.+?)\]$")
def main() -> int:
if len(sys.argv) < 2:
print("Usage: python extract_screenshot_placeholders.py <markdown-file> [--json-out path]")
return 1
source = Path(sys.argv[1])
labels = []
for line in source.read_text(encoding="utf-8").splitlines():
m = PLACEHOLDER_RE.match(line.strip())
if m:
labels.append(m.group(1).strip())
print("COUNT\t" + str(len(labels)))
for label in labels:
print(label)
if len(sys.argv) >= 4 and sys.argv[2] == "--json-out":
Path(sys.argv[3]).write_text(
json.dumps({"labels": labels}, ensure_ascii=False, indent=2),
encoding="utf-8",
)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import re
from pathlib import Path
from docx import Document
from docx.enum.text import WD_ALIGN_PARAGRAPH
from docx.oxml import OxmlElement
from docx.oxml.ns import qn
from docx.shared import Cm, Pt
CODE_BLOCK_RE = re.compile(r"^```(mermaid|plantuml)\s*$", re.I)
CAPTION_RE = re.compile(r"^(图\s*\d+(?:\.\d+)?\s+.+)$")
def set_run_fonts(run, east_asia_font: str, latin_font: str, size_pt: float, *, bold: bool = False) -> None:
run.bold = bold
run.font.size = Pt(size_pt)
run.font.name = latin_font
r_pr = run._element.get_or_add_rPr()
r_fonts = r_pr.rFonts
if r_fonts is None:
r_fonts = OxmlElement("w:rFonts")
r_pr.append(r_fonts)
r_fonts.set(qn("w:eastAsia"), east_asia_font)
r_fonts.set(qn("w:ascii"), latin_font)
r_fonts.set(qn("w:hAnsi"), latin_font)
def add_paragraph(doc: Document, text: str, *, font_cn: str = "宋体", font_en: str = "Times New Roman", size: float = 10.5, bold: bool = False, center: bool = False) -> None:
paragraph = doc.add_paragraph()
paragraph.alignment = WD_ALIGN_PARAGRAPH.CENTER if center else WD_ALIGN_PARAGRAPH.LEFT
run = paragraph.add_run(text)
set_run_fonts(run, font_cn, font_en, size, bold=bold)
if not center:
paragraph.paragraph_format.first_line_indent = Pt(21)
paragraph.paragraph_format.line_spacing = 1.25
def add_code_block(doc: Document, code: str) -> None:
for line in code.splitlines() or [""]:
paragraph = doc.add_paragraph()
paragraph.alignment = WD_ALIGN_PARAGRAPH.LEFT
paragraph.paragraph_format.first_line_indent = Pt(0)
paragraph.paragraph_format.line_spacing = 1.1
run = paragraph.add_run(line.rstrip())
set_run_fonts(run, "Consolas", "Consolas", 9)
def extract_blocks(markdown_path: Path) -> list[dict[str, str]]:
lines = markdown_path.read_text(encoding="utf-8").splitlines()
blocks: list[dict[str, str]] = []
in_code = False
lang = ""
buffer: list[str] = []
pending: list[dict[str, str]] = []
for line in lines:
stripped = line.strip()
if in_code:
if stripped.startswith("```"):
pending.append({"lang": lang.lower(), "code": "\n".join(buffer)})
in_code = False
lang = ""
buffer = []
else:
buffer.append(line.rstrip())
continue
code_match = CODE_BLOCK_RE.match(stripped)
if code_match:
in_code = True
lang = code_match.group(1)
buffer = []
continue
caption_match = CAPTION_RE.match(stripped)
if caption_match and pending:
item = pending.pop(0)
item["caption"] = caption_match.group(1).strip()
blocks.append(item)
return blocks
def build_doc(title: str, markdown_path: Path, output_path: Path) -> Path:
blocks = extract_blocks(markdown_path)
doc = Document()
section = doc.sections[0]
section.top_margin = Cm(2.54)
section.bottom_margin = Cm(2.54)
section.left_margin = Cm(3.17)
section.right_margin = Cm(3.17)
add_paragraph(doc, f"{title}-附件", font_cn="黑体", font_en="Times New Roman", size=18, bold=True, center=True)
add_paragraph(doc, "流程图、E-R 图及其他 Mermaid / PlantUML 源码附件", font_cn="黑体", font_en="Times New Roman", size=14, center=True)
if not blocks:
add_paragraph(doc, "正文中未提取到 Mermaid 或 PlantUML 代码块。", size=10.5)
else:
for index, block in enumerate(blocks, start=1):
add_paragraph(doc, f"{index}. {block['caption']}", font_cn="黑体", font_en="Times New Roman", size=14, bold=True)
add_paragraph(doc, f"图类型:{block['lang']}", size=10.5)
add_code_block(doc, block["code"])
output_path.parent.mkdir(parents=True, exist_ok=True)
doc.save(str(output_path))
return output_path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate appendix DOCX containing Mermaid / PlantUML source blocks from thesis markdown.")
parser.add_argument("title", help="Thesis title used for the appendix heading")
parser.add_argument("markdown_file", type=Path)
parser.add_argument("target_docx", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
output = build_doc(args.title, args.markdown_file, args.target_docx)
print(output)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import json
import re
from pathlib import Path
from typing import Any
from docx import Document
from docx.enum.table import WD_ALIGN_VERTICAL, WD_TABLE_ALIGNMENT
from docx.enum.text import WD_ALIGN_PARAGRAPH, WD_LINE_SPACING
from docx.oxml import OxmlElement
from docx.oxml.ns import qn
from docx.shared import Cm, Pt
from markdown_utils import SCREENSHOT_PLACEHOLDER_RE, cleanup_inline_markdown
SPECIAL_CENTERED_HEADINGS = {
"摘要": "abstract_heading_cn",
"abstract": "abstract_heading_en",
"参考文献": "references_heading",
"致谢": "ack_heading",
}
ALIGNMENT_MAP = {
"left": WD_ALIGN_PARAGRAPH.LEFT,
"center": WD_ALIGN_PARAGRAPH.CENTER,
"right": WD_ALIGN_PARAGRAPH.RIGHT,
"justify": WD_ALIGN_PARAGRAPH.JUSTIFY,
"distribute": WD_ALIGN_PARAGRAPH.DISTRIBUTE,
}
def default_style_profile() -> dict[str, Any]:
return {
"styles": {
"title": centered_style("黑体", "Times New Roman", 18, bold=True),
"heading1": paragraph_style("黑体", "Times New Roman", 18, bold=True, space_before=10, space_after=10),
"heading2": paragraph_style("黑体", "Times New Roman", 15, bold=True, space_before=10, space_after=10),
"heading3": paragraph_style("黑体", "Times New Roman", 12, bold=True, space_before=10, space_after=10),
"abstract_heading_cn": centered_style("黑体", "Times New Roman", 18, bold=True, page_break_before=True),
"abstract_heading_en": centered_style(
"Times New Roman",
"Times New Roman",
18,
bold=True,
page_break_before=True,
),
"references_heading": centered_style("黑体", "Times New Roman", 18, bold=True, page_break_before=True),
"ack_heading": centered_style("黑体", "Times New Roman", 18, bold=True, page_break_before=True),
"body_cn": paragraph_style("宋体", "Times New Roman", 10.5, first_line_indent=21),
"body_en": paragraph_style("Times New Roman", "Times New Roman", 12, first_line_indent=21),
"keywords_cn_label": run_style("黑体", "Times New Roman", 12, bold=True),
"keywords_cn_content": run_style("宋体", "Times New Roman", 12),
"keywords_en_label": run_style("Times New Roman", "Times New Roman", 12, bold=True),
"keywords_en_content": run_style("Times New Roman", "Times New Roman", 12),
"keywords_paragraph": paragraph_style("宋体", "Times New Roman", 10.5, first_line_indent=0),
"figure_caption": centered_style("宋体", "Times New Roman", 10.5, line_spacing=1, line_spacing_rule="single"),
"table_caption": centered_style("宋体", "Times New Roman", 10.5, line_spacing=1, line_spacing_rule="single"),
"table_text": centered_style("宋体", "Times New Roman", 10.5),
"references_body": paragraph_style("宋体", "Times New Roman", 10.5, first_line_indent=-21, left_indent=21),
"missing_asset": centered_style("楷体", "Times New Roman", 10.5),
"code": paragraph_style("Consolas", "Consolas", 9, first_line_indent=0),
}
}
def paragraph_style(
east_asia_font: str,
latin_font: str,
size_pt: float,
*,
bold: bool = False,
alignment: str = "left",
line_spacing: float = 1.25,
line_spacing_rule: str = "multiple",
first_line_indent: float = 0,
left_indent: float = 0,
space_before: float = 0,
space_after: float = 0,
page_break_before: bool = False,
) -> dict[str, Any]:
return {
"east_asia_font": east_asia_font,
"latin_font": latin_font,
"size_pt": size_pt,
"bold": bold,
"alignment": alignment,
"line_spacing": line_spacing,
"line_spacing_rule": line_spacing_rule,
"first_line_indent_pt": first_line_indent,
"left_indent_pt": left_indent,
"space_before_pt": space_before,
"space_after_pt": space_after,
"page_break_before": page_break_before,
}
def centered_style(
east_asia_font: str,
latin_font: str,
size_pt: float,
*,
bold: bool = False,
line_spacing: float = 1.25,
line_spacing_rule: str = "multiple",
page_break_before: bool = False,
) -> dict[str, Any]:
return paragraph_style(
east_asia_font,
latin_font,
size_pt,
bold=bold,
alignment="center",
line_spacing=line_spacing,
line_spacing_rule=line_spacing_rule,
first_line_indent=0,
space_before=0,
space_after=0,
page_break_before=page_break_before,
)
def run_style(east_asia_font: str, latin_font: str, size_pt: float, *, bold: bool = False) -> dict[str, Any]:
return {
"east_asia_font": east_asia_font,
"latin_font": latin_font,
"size_pt": size_pt,
"bold": bold,
}
def infer_fonts(style: dict[str, Any], fallback: dict[str, Any]) -> dict[str, Any]:
font = style.get("font")
if font and "east_asia_font" not in style:
style["east_asia_font"] = font
if font and "latin_font" not in style:
style["latin_font"] = font if re.search(r"[A-Za-z]", str(font)) and not re.search(r"[\u4e00-\u9fff]", str(font)) else fallback["latin_font"]
return style
def deep_merge(base: dict[str, Any], override: dict[str, Any]) -> dict[str, Any]:
merged = json.loads(json.dumps(base))
for key, value in override.items():
if isinstance(value, dict) and isinstance(merged.get(key), dict):
merged[key] = deep_merge(merged[key], value)
else:
merged[key] = value
return merged
def load_style_profile(path: Path | None) -> dict[str, Any]:
profile = default_style_profile()
if path is None or not path.exists():
return profile
raw = json.loads(path.read_text(encoding="utf-8"))
if "styles" not in raw:
raw = {"styles": raw}
profile = deep_merge(profile, raw)
styles = profile["styles"]
fallback_map = default_style_profile()["styles"]
for name, style in list(styles.items()):
if isinstance(style, dict):
styles[name] = infer_fonts(style, fallback_map.get(name, fallback_map["body_cn"]))
# Allow simplified analyzer output to drive keyword/caption styles.
if "keywords_cn" in styles:
styles["keywords_paragraph"] = deep_merge(styles["keywords_paragraph"], styles["keywords_cn"])
styles["keywords_cn_label"] = deep_merge(styles["keywords_cn_label"], styles["keywords_cn"])
styles["keywords_cn_content"] = deep_merge(styles["keywords_cn_content"], styles["keywords_cn"])
if "keywords_en" in styles:
styles["keywords_paragraph"] = deep_merge(styles["keywords_paragraph"], styles["keywords_en"])
styles["keywords_en_label"] = deep_merge(styles["keywords_en_label"], styles["keywords_en"])
styles["keywords_en_content"] = deep_merge(styles["keywords_en_content"], styles["keywords_en"])
if "figure_caption" in styles:
styles["figure_caption"] = infer_fonts(styles["figure_caption"], fallback_map["figure_caption"])
if "table_caption" in styles:
styles["table_caption"] = infer_fonts(styles["table_caption"], fallback_map["table_caption"])
return profile
def set_run_fonts(run, east_asia_font: str, latin_font: str, size_pt: float, *, bold: bool = False) -> None:
run.bold = bold
run.font.size = Pt(size_pt)
run.font.name = latin_font
r_pr = run._element.get_or_add_rPr()
r_fonts = r_pr.rFonts
if r_fonts is None:
r_fonts = OxmlElement("w:rFonts")
r_pr.append(r_fonts)
r_fonts.set(qn("w:eastAsia"), east_asia_font)
r_fonts.set(qn("w:ascii"), latin_font)
r_fonts.set(qn("w:hAnsi"), latin_font)
def apply_paragraph_style(paragraph, style: dict[str, Any]) -> None:
paragraph.alignment = ALIGNMENT_MAP.get(style.get("alignment", "left"), WD_ALIGN_PARAGRAPH.LEFT)
paragraph.paragraph_format.line_spacing_rule = (
WD_LINE_SPACING.SINGLE if style.get("line_spacing_rule") == "single" else WD_LINE_SPACING.MULTIPLE
)
paragraph.paragraph_format.line_spacing = style.get("line_spacing", 1.25)
paragraph.paragraph_format.first_line_indent = Pt(style.get("first_line_indent_pt", 0))
paragraph.paragraph_format.left_indent = Pt(style.get("left_indent_pt", 0))
paragraph.paragraph_format.space_before = Pt(style.get("space_before_pt", 0))
paragraph.paragraph_format.space_after = Pt(style.get("space_after_pt", 0))
if style.get("page_break_before"):
add_page_break_before(paragraph)
for run in paragraph.runs:
set_run_fonts(
run,
style["east_asia_font"],
style["latin_font"],
style["size_pt"],
bold=style.get("bold", False),
)
def add_page_break_before(paragraph) -> None:
p_pr = paragraph._p.get_or_add_pPr()
page_break_before = OxmlElement("w:pageBreakBefore")
p_pr.append(page_break_before)
def load_image_map(path: Path | None) -> dict[str, Path]:
if path is None or not path.exists():
return {}
data = json.loads(path.read_text(encoding="utf-8"))
return {k: Path(v) for k, v in data.items()}
def add_image(doc: Document, path: Path) -> None:
paragraph = doc.add_paragraph()
paragraph.alignment = WD_ALIGN_PARAGRAPH.CENTER
paragraph.add_run().add_picture(str(path), width=Cm(14.5))
def add_missing_asset_placeholder(doc: Document, label: str, styles: dict[str, Any]) -> None:
paragraph = doc.add_paragraph()
paragraph.add_run(f"【待补素材:{label}】")
apply_paragraph_style(paragraph, styles["missing_asset"])
def normalize_text(text: str) -> str:
return cleanup_inline_markdown(text)
def add_markdown_table(doc: Document, lines: list[str], styles: dict[str, Any]) -> None:
rows = []
for row in lines:
cells = [normalize_text(cell.strip()) for cell in row.strip().strip("|").split("|")]
if all(re.fullmatch(r"[:\- ]+", cell or "") for cell in cells):
continue
rows.append(cells)
if len(rows) < 2:
return
headers = rows[0]
data_rows = rows[1:]
table = doc.add_table(rows=1 + len(data_rows), cols=len(headers))
table.style = "Table Grid"
table.alignment = WD_TABLE_ALIGNMENT.CENTER
for idx, text in enumerate(headers):
table.rows[0].cells[idx].text = text
for row_index, row in enumerate(data_rows, start=1):
for cell_index, text in enumerate(row):
if cell_index < len(table.rows[row_index].cells):
table.rows[row_index].cells[cell_index].text = text
for row in table.rows:
for cell in row.cells:
cell.vertical_alignment = WD_ALIGN_VERTICAL.CENTER
for paragraph in cell.paragraphs:
apply_paragraph_style(paragraph, styles["table_text"])
def apply_keyword_runs(paragraph, label: str, content: str, styles: dict[str, Any]) -> None:
apply_paragraph_style(paragraph, styles["keywords_paragraph"])
if label.startswith("关键词"):
label_style = styles["keywords_cn_label"]
content_style = styles["keywords_cn_content"]
else:
label_style = styles["keywords_en_label"]
content_style = styles["keywords_en_content"]
label_run = paragraph.add_run(label)
set_run_fonts(
label_run,
label_style["east_asia_font"],
label_style["latin_font"],
label_style["size_pt"],
bold=label_style.get("bold", False),
)
if content:
spacer = "" if label.endswith((":", ":")) else " "
content_run = paragraph.add_run(f"{spacer}{normalize_text(content)}")
set_run_fonts(
content_run,
content_style["east_asia_font"],
content_style["latin_font"],
content_style["size_pt"],
bold=content_style.get("bold", False),
)
def build_doc(source: Path, image_map: dict[str, Path], profile: dict[str, Any]) -> Document:
styles = profile["styles"]
doc = Document()
section = doc.sections[0]
section.top_margin = Cm(2.54)
section.bottom_margin = Cm(2.54)
section.left_margin = Cm(3.17)
section.right_margin = Cm(3.17)
lines = source.read_text(encoding="utf-8").splitlines()
current_section = ""
in_code = False
code_lang = ""
pending_mermaid = False
seen_content = False
index = 0
while index < len(lines):
line = lines[index]
stripped = line.strip()
if in_code:
if stripped.startswith("```"):
in_code = False
if code_lang == "mermaid":
pending_mermaid = True
code_lang = ""
elif code_lang != "mermaid":
paragraph = doc.add_paragraph()
paragraph.add_run(line.rstrip())
apply_paragraph_style(paragraph, styles["code"])
index += 1
continue
if not stripped:
index += 1
continue
if stripped == "---":
index += 1
continue
if stripped.startswith("```"):
in_code = True
code_lang = stripped[3:].strip().lower()
index += 1
continue
if stripped.startswith("# "):
paragraph = doc.add_paragraph()
paragraph.add_run(normalize_text(stripped[2:].strip()))
apply_paragraph_style(paragraph, styles["title"])
seen_content = True
index += 1
continue
if stripped.startswith("## "):
text = normalize_text(stripped[3:].strip())
normalized = text.replace(" ", "").lower()
style_key = SPECIAL_CENTERED_HEADINGS.get(normalized, "heading1")
paragraph = doc.add_paragraph()
paragraph.add_run(text)
if seen_content and style_key == "heading1":
style = deep_merge(styles["heading1"], {"page_break_before": True})
else:
style = styles[style_key]
apply_paragraph_style(paragraph, style)
current_section = text
seen_content = True
index += 1
continue
if stripped.startswith("### "):
paragraph = doc.add_paragraph()
paragraph.add_run(normalize_text(stripped[4:].strip()))
apply_paragraph_style(paragraph, styles["heading2"])
seen_content = True
index += 1
continue
if stripped.startswith("#### "):
paragraph = doc.add_paragraph()
paragraph.add_run(normalize_text(stripped[5:].strip()))
apply_paragraph_style(paragraph, styles["heading3"])
seen_content = True
index += 1
continue
keyword_match = re.match(r"^(关键词[::]|Keywords[::])\s*(.*)$", stripped)
if keyword_match:
paragraph = doc.add_paragraph()
apply_keyword_runs(paragraph, keyword_match.group(1), keyword_match.group(2), styles)
seen_content = True
index += 1
continue
image_match = SCREENSHOT_PLACEHOLDER_RE.match(stripped)
if image_match:
label = image_match.group(1).strip()
image_path = image_map.get(label)
if image_path and image_path.exists():
add_image(doc, image_path)
else:
add_missing_asset_placeholder(doc, label, styles)
seen_content = True
index += 1
continue
if re.match(r"^图\s*\d+(\.\d+)?", stripped):
caption = normalize_text(stripped)
if pending_mermaid:
image_path = image_map.get(caption)
if image_path and image_path.exists():
add_image(doc, image_path)
else:
add_missing_asset_placeholder(doc, caption, styles)
pending_mermaid = False
paragraph = doc.add_paragraph()
paragraph.add_run(caption)
apply_paragraph_style(paragraph, styles["figure_caption"])
seen_content = True
index += 1
continue
if re.match(r"^表\s*\d+(\.\d+)?", stripped):
paragraph = doc.add_paragraph()
paragraph.add_run(normalize_text(stripped))
apply_paragraph_style(paragraph, styles["table_caption"])
seen_content = True
index += 1
continue
if stripped.startswith("|"):
table_lines = []
while index < len(lines) and lines[index].strip().startswith("|"):
table_lines.append(lines[index].strip())
index += 1
add_markdown_table(doc, table_lines, styles)
seen_content = True
continue
paragraph = doc.add_paragraph()
paragraph.add_run(normalize_text(stripped))
if current_section.replace(" ", "") == "参考文献" and re.match(r"^\[\d+\]", stripped):
apply_paragraph_style(paragraph, styles["references_body"])
elif current_section.lower() == "abstract":
apply_paragraph_style(paragraph, styles["body_en"])
else:
apply_paragraph_style(paragraph, styles["body_cn"])
seen_content = True
index += 1
return doc
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate thesis DOCX from markdown source.")
parser.add_argument("source", type=Path)
parser.add_argument("target", type=Path)
parser.add_argument("legacy_image_map", nargs="?", type=Path, help="Optional image map for backward compatibility.")
parser.add_argument("--image-map", dest="image_map", type=Path, default=None)
parser.add_argument("--style-profile", dest="style_profile", type=Path, default=None)
return parser.parse_args()
def main() -> int:
args = parse_args()
image_map_path = args.image_map or args.legacy_image_map
profile = load_style_profile(args.style_profile)
image_map = load_image_map(image_map_path)
args.target.parent.mkdir(parents=True, exist_ok=True)
document = build_doc(args.source, image_map, profile)
document.save(args.target)
print(args.target)
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import re
SCREENSHOT_PLACEHOLDER_RE = re.compile(r"^\[此处插入截图:(.+?)\]$")
LINK_RE = re.compile(r"\[([^\]]+)\]\([^)]+\)")
AUTOLINK_RE = re.compile(r"<(https?://[^>]+)>")
def cleanup_inline_markdown(text: str) -> str:
cleaned = text.replace("`", "")
cleaned = LINK_RE.sub(r"\1", cleaned)
cleaned = AUTOLINK_RE.sub(r"\1", cleaned)
# Remove emphasis markers but keep their content.
for pattern in (
(r"\*\*(.+?)\*\*", r"\1"),
(r"__(.+?)__", r"\1"),
(r"(?<!\*)\*(?!\s)(.+?)(?<!\s)\*(?!\*)", r"\1"),
(r"(?<!_)_(?!\s)(.+?)(?<!\s)_(?!_)", r"\1"),
):
cleaned = re.sub(pattern[0], pattern[1], cleaned)
cleaned = re.sub(r"\s+", " ", cleaned)
return cleaned.strip()
def markdown_to_visible_text(text: str) -> str:
lines: list[str] = []
in_code = False
for raw_line in text.splitlines():
stripped = raw_line.strip()
if stripped.startswith("```"):
in_code = not in_code
continue
if in_code or not stripped:
continue
if stripped == "---":
continue
if SCREENSHOT_PLACEHOLDER_RE.match(stripped):
continue
if stripped.startswith("|"):
cells = [cleanup_inline_markdown(cell) for cell in stripped.strip("|").split("|")]
if all(re.fullmatch(r"[:\- ]+", cell or "") for cell in cells):
continue
lines.append(" ".join(cell for cell in cells if cell))
continue
if stripped.startswith("#"):
stripped = stripped.lstrip("#").strip()
lines.append(cleanup_inline_markdown(stripped))
return "\n".join(line for line in lines if line)
def compute_text_metrics(text: str) -> dict[str, int]:
visible = markdown_to_visible_text(text)
char_with_spaces = len(visible)
char_no_spaces = len(re.sub(r"\s+", "", visible))
chinese_chars = len(re.findall(r"[\u4e00-\u9fff]", visible))
non_chinese_words = len(re.findall(r"[A-Za-z0-9]+(?:[._/\-'][A-Za-z0-9]+)*", visible))
english_words = len(re.findall(r"[A-Za-z]+(?:[-'][A-Za-z]+)*", visible))
approx_word_count = chinese_chars + non_chinese_words
return {
"char_with_spaces": char_with_spaces,
"char_no_spaces": char_no_spaces,
"chinese_chars": chinese_chars,
"non_chinese_words": non_chinese_words,
"english_words": english_words,
"approx_word_count": approx_word_count,
}
from __future__ import annotations
import argparse
import shutil
import subprocess
import sys
from pathlib import Path
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("input")
parser.add_argument("output")
parser.add_argument("--cwd", default=None, help="Working directory for npx execution")
parser.add_argument("--puppeteer-config", default=None, help="Path to puppeteer config json")
args = parser.parse_args()
input_path = Path(args.input)
output_path = Path(args.output)
cwd = args.cwd
npx_path = shutil.which("npx") or shutil.which("npx.cmd")
if npx_path is None:
print("ERROR\tMissing npx. Install Node.js first.")
return 2
cmd = [npx_path, "@mermaid-js/mermaid-cli", "-i", str(input_path), "-o", str(output_path)]
if args.puppeteer_config:
cmd.extend(["-p", str(Path(args.puppeteer_config))])
result = subprocess.run(cmd, capture_output=True, text=True, cwd=cwd)
if result.returncode != 0:
print("ERROR\tmmdc failed")
if result.stdout.strip():
print(result.stdout.strip())
if result.stderr.strip():
print(result.stderr.strip())
return result.returncode
print(str(output_path))
return 0
if __name__ == "__main__":
raise SystemExit(main())
from __future__ import annotations
import argparse
import json
from pathlib import Path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Create an empty reference verification checklist template for thesis work.")
parser.add_argument("target_json", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
args.target_json.parent.mkdir(parents=True, exist_ok=True)
template = {
"references": [
{
"title": "",
"authors": [],
"year": "",
"source": "",
"doi_or_url": "",
"citation_count_if_available": "",
"relevance_note": "",
"status": "verified"
}
]
}
args.target_json.write_text(json.dumps(template, ensure_ascii=False, indent=2), encoding="utf-8")
print(args.target_json)
return 0
if __name__ == "__main__":
raise SystemExit(main())