
Citation Manager
- 8 installs
- 33 repo stars
- Updated April 26, 2026
- bighardperson/computer-science-skills-collection
Citation-manager is a skill that adds real references and standardizes citations for research papers and theses with Crossref integration.
About
Citation-manager is a skill for managing academic citations in research papers and theses. It fetches real bibliographic metadata from Crossref by DOI, ISBN, or title, generates in-text citations and bibliography lists, and converts between formats such as APA, MLA, Chicago, GB/T 7714, IEEE, and Harvard. It also checks that in-text citations match the reference list.
- Adds real references and standardizes citations for papers and theses
- Supports APA, MLA, Chicago, GB/T 7714, IEEE, and Harvard formats
- Fetches bibliographic metadata from Crossref via DOI, ISBN, or title
Citation Manager by the numbers
- 8 all-time installs (skills.sh)
- Ranked #1,167 of 1,879 Documentation skills by installs in the Skillselion catalog
- Data as of Jul 30, 2026 (Skillselion catalog sync)
citation-manager capabilities & compatibility
- Capabilities
- documentation
- Use cases
- documentation · research
What citation-manager says it does
Add real references and standardize citations for research papers and theses. Supports CrossRef integration, multiple citation formats (APA/MLA/Chicago/GB-T), batch import, and auto-detection of citat
Academic citation manager for papers and theses with CrossRef integration
Citation integrity checking: Automatically checks consistency between in-text citations and bibliography lists
npx skills add https://github.com/bighardperson/computer-science-skills-collection --skill citation-managerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 8 |
|---|---|
| repo stars | ★ 33 |
| Last updated | April 26, 2026 |
| Repository | bighardperson/computer-science-skills-collection ↗ |
What it does
Add real references and generate standardized citations for research papers across APA, MLA, IEEE, and other formats.
Who is it for?
Adding real references and formatting citations for academic papers and theses.
Skip if: General code documentation or non-academic writing.
When should I use this skill?
You need to add, format, or convert citations for a research paper or thesis.
What you get
Generates in-text citations and a formatted bibliography from real metadata.
- In-text citations
- Formatted bibliography list
- Citation integrity report
By the numbers
- Supports 6 citation formats: APA, MLA, Chicago, GB/T 7714, IEEE, Harvard
Files
Academic Citation Manager | 学术引用管理器
功能说明 | Function Description
【中文】 学术引用管理器是一个专业的科研论文引用管理工具,旨在为科研论文和毕业论文添加真实参考文献并规范引用标注。该工具支持多种国际通用的引用格式,包括APA、MLA、Chicago、GB/T 7714、IEEE、Harvard等,并与Crossref API集成以获取真实的文献元数据。
核心功能包括:
- 多格式引用支持:支持APA 7th、MLA 9th、Chicago 17th、GB/T 7714-2015、IEEE、Harvard等多种引用格式
- 智能元数据获取:通过DOI、ISBN、标题等信息从Crossref等权威数据库自动获取文献元数据
- 文内引用管理:自动生成和插入文中引用标注,支持作者-年份制和序号制
- 参考文献列表生成:自动生成格式规范的参考文献列表,支持多种排序方式
- 批量文献导入:支持批量导入和管理参考文献,提高工作效率
- 格式转换:在不同引用格式间进行快速转换,满足不同期刊要求
- 引用完整性检查:自动检查文中引用与参考文献列表的一致性
- 中英文双语支持:完整支持中文和英文文献的处理和格式化
【English】 Academic Citation Manager is a professional research paper citation management tool designed to add real references and standardize citations for research papers and theses. The tool supports multiple internationally used citation formats, including APA, MLA, Chicago, GB/T 7714, IEEE, Harvard, etc., and integrates with Crossref API to retrieve authentic bibliographic metadata.
Core features include:
- Multi-format citation support: Supports APA 7th, MLA 9th, Chicago 17th, GB/T 7714-2015, IEEE, Harvard and other citation formats
- Intelligent metadata retrieval: Automatically retrieves bibliographic metadata from authoritative databases like Crossref via DOI, ISBN, title, etc.
- In-text citation management: Automatically generates and inserts in-text citations, supporting author-year and numeric systems
- Bibliography list generation: Automatically generates format-compliant bibliography lists with various sorting options
- Batch reference import: Supports batch importing and managing references to improve efficiency
- Format conversion: Fast conversion between different citation formats to meet various journal requirements
- Citation integrity checking: Automatically checks consistency between in-text citations and bibliography lists
- Bilingual support: Complete support for processing and formatting Chinese and English references
支持的引用格式 | Supported Citation Formats
【中文】
本工具支持以下主要引用格式:
APA 7th Edition
- 适用领域:心理学、教育学、社会科学
- 文内引用:(Smith, 2023) 或 Smith (2023)
- 特点:作者-年份制,强调出版年,DOI强制推荐
MLA 9th Edition
- 适用领域:人文学、语言文学
- 文内引用:(Smith 23) 或 Smith (23)
- 特点:作者-页码制,对"容器"层次要求明确
Chicago 17th Edition
- 适用领域:历史、艺术、人文学
- 注脚-参考文献制:Superscript¹ 或 Author, Title, p. xx
- 作者-日期制:(Smith 2023, 45)
- 特点:支持两种体系,历史文献注释丰富
GB/T 7714-2015
- 适用领域:中文学位论文、期刊
- 文内引用:[1] 或 (张, 2023)
- 特点:支持序号与作者-日期双制,强制文献类型标识
IEEE
- 适用领域:电气、电子、计算机
- 文内引用:[1]
- 特点:方括号数字制,强调出版年、卷期页码
Harvard
- 适用领域:英联邦高校通用
- 文内引用:(Smith, 1999, p. 23)
- 特点:作者-年份制,页码必加
【English】
This tool supports the following major citation formats:
APA 7th Edition
- Fields: Psychology, Education, Social Sciences
- In-text citation: (Smith, 2023) or Smith (2023)
- Features: Author-year system, emphasizes publication year, DOI recommended
MLA 9th Edition
- Fields: Humanities, Language & Literature
- In-text citation: (Smith 23) or Smith (23)
- Features: Author-page system, clear container hierarchy requirements
Chicago 17th Edition
- Fields: History, Arts, Humanities
- Notes-bibliography: Superscript¹ or Author, Title, p. xx
- Author-date: (Smith 2023, 45)
- Features: Supports two systems, rich historical document annotations
GB/T 7714-2015
- Fields: Chinese academic papers, journals
- In-text citation: [1] or (Zhang, 2023)
- Features: Supports both numeric and author-date systems, mandatory document type identifiers
IEEE
- Fields: Electrical, Electronics, Computer Science
- In-text citation: [1]
- Features: Bracketed numeric system, emphasizes publication year, volume, issue, page numbers
Harvard
- Fields: Commonwealth universities
- In-text citation: (Smith, 1999, p. 23)
- Features: Author-year system, page numbers mandatory
使用方法 | Usage
【中文】
Python API 使用
from academic_citation_skill import AcademicCitationManager
# 创建管理器实例
manager = AcademicCitationManager()
# 方法1:通过DOI获取文献信息
doi = "10.1000/xyz123"
reference = manager.fetch_reference_by_doi(doi)
print(f"标题: {reference['title']}")
print(f"作者: {reference['authors']}")
# 方法2:通过ISBN获取图书信息
isbn = "9780262033848"
reference = manager.fetch_reference_by_isbn(isbn)
# 方法3:通过标题和作者搜索
results = manager.search_references(
title="Artificial Intelligence",
author="Russell",
year=2020
)
# 添加到本地文献库
manager.add_to_library(reference)
# 生成文中引用
citation = manager.generate_citation(
reference_id=reference['id'],
style='apa',
citation_type='author-date'
)
print(f"文中引用: {citation}")
# 生成参考文献列表
bibliography = manager.generate_bibliography(
style='apa',
sort_by='alphabetical'
)
# 检查引用完整性
issues = manager.check_citation_integrity(document_text="your document content")
for issue in issues:
print(f"问题: {issue}")
# 格式转换
converted = manager.convert_citation_style(
references=[reference],
from_style='apa',
to_style='ieee'
)命令行使用
# 通过DOI获取文献信息
python academic_citation_skill.py --fetch-doi 10.1000/xyz123
# 通过ISBN获取图书信息
python academic_citation_skill.py --fetch-isbn 9780262033848
# 批量导入参考文献
python batch_import.py --input references.bib --format bibtex
# 格式转换
python format_converter.py --input apa_refs.txt --from-style apa --to-style ieee
# 检查引用完整性
python citation_checker.py --document paper.docx --bibliography refs.txt
# 生成参考文献列表
python academic_citation_skill.py --generate-bib --style apa --input citations.txt集成到写作工作流
# 在LaTeX中使用
from academic_citation_skill import AcademicCitationManager
manager = AcademicCitationManager()
# 生成BibTeX格式
bibtex = manager.export_bibtex()
with open('references.bib', 'w') as f:
f.write(bibtex)
# 在Word中使用
from academic_citation_skill import WordCitationPlugin
plugin = WordCitationPlugin()
plugin.insert_citation(reference_id="ref1", style="apa")
plugin.generate_bibliography(style="apa")
# 在Markdown中使用
from academic_citation_skill import MarkdownCitationFormatter
formatter = MarkdownCitationFormatter()
formatted_text = formatter.format_markdown(
text="Your document with @[ref1] citations",
style="apa"
)【English】
Python API Usage
from academic_citation_skill import AcademicCitationManager
# Create manager instance
manager = AcademicCitationManager()
# Method 1: Fetch reference by DOI
doi = "10.1000/xyz123"
reference = manager.fetch_reference_by_doi(doi)
print(f"Title: {reference['title']}")
print(f"Authors: {reference['authors']}")
# Method 2: Fetch book information by ISBN
isbn = "9780262033848"
reference = manager.fetch_reference_by_isbn(isbn)
# Method 3: Search by title and author
results = manager.search_references(
title="Artificial Intelligence",
author="Russell",
year=2020
)
# Add to local library
manager.add_to_library(reference)
# Generate in-text citation
citation = manager.generate_citation(
reference_id=reference['id'],
style='apa',
citation_type='author-date'
)
print(f"In-text citation: {citation}")
# Generate bibliography list
bibliography = manager.generate_bibliography(
style='apa',
sort_by='alphabetical'
)
# Check citation integrity
issues = manager.check_citation_integrity(document_text="your document content")
for issue in issues:
print(f"Issue: {issue}")
# Convert citation style
converted = manager.convert_citation_style(
references=[reference],
from_style='apa',
to_style='ieee'
)Command Line Usage
# Fetch reference by DOI
python academic_citation_skill.py --fetch-doi 10.1000/xyz123
# Fetch book information by ISBN
python academic_citation_skill.py --fetch-isbn 9780262033848
# Batch import references
python batch_import.py --input references.bib --format bibtex
# Format conversion
python format_converter.py --input apa_refs.txt --from-style apa --to-style ieee
# Check citation integrity
python citation_checker.py --document paper.docx --bibliography refs.txt
# Generate bibliography
python academic_citation_skill.py --generate-bib --style apa --input citations.txtIntegration into Writing Workflow
# Using with LaTeX
from academic_citation_skill import AcademicCitationManager
manager = AcademicCitationManager()
# Generate BibTeX format
bibtex = manager.export_bibtex()
with open('references.bib', 'w') as f:
f.write(bibtex)
# Using with Word
from academic_citation_skill import WordCitationPlugin
plugin = WordCitationPlugin()
plugin.insert_citation(reference_id="ref1", style="apa")
plugin.generate_bibliography(style="apa")
# Using with Markdown
from academic_citation_skill import MarkdownCitationFormatter
formatter = MarkdownCitationFormatter()
formatted_text = formatter.format_markdown(
text="Your document with @[ref1] citations",
style="apa"
)配置选项 | Configuration Options
【中文】
引用格式配置 (citation_styles.json)
主要配置项:
{
"apa": {
"name": "APA 7th Edition",
"in_text": {
"format": "parenthetical",
"separator": ", ",
"year_position": "after_author",
"page_number_format": "p. {page}"
},
"bibliography": {
"sort_by": "alphabetical",
"author_format": "last_first",
"title_format": "sentence_case",
"doi_required": true,
"url_format": "available_at"
},
"document_type_codes": {
"book": "",
"journal": "",
"conference": "",
"thesis": "",
"report": ""
}
},
"gbt7714": {
"name": "GB/T 7714-2015",
"in_text": {
"format": "numeric_bracket",
"separator": ", ",
"prefix": "[",
"suffix": "]"
},
"bibliography": {
"sort_by": "citation_order",
"author_format": "chinese_name_order",
"document_type_codes": {
"book": "[M]",
"journal": "[J]",
"conference": "[C]",
"thesis": "[D]",
"report": "[R]",
"newspaper": "[N]"
}
},
"chinese_authors": {
"name_order": "last_first",
"separator": ", "
}
}
}Crossref API配置 (crossref_config.json)
{
"api_base": "https://api.crossref.org",
"endpoints": {
"works": "/works",
"journals": "/journals",
"types": "/types",
"fields": "/fields"
},
"rate_limiting": {
"requests_per_second": 10,
"backoff_strategy": "exponential",
"max_retries": 3
},
"filters": {
"default": {
"has_full_text": true,
"state": "active"
},
"journals": {
"type": "journal-article"
},
"books": {
"type": "monograph"
}
},
"cache": {
"enabled": true,
"ttl_seconds": 86400,
"max_size": 1000
}
}本地文献库配置 (reference_database.json)
{
"metadata": {
"version": "1.0.0",
"created_date": "2026-03-01",
"last_updated": "2026-03-01"
},
"references": {
"ref_001": {
"id": "ref_001",
"type": "journal_article",
"title": "Deep Learning",
"authors": [
{
"given": "Yann",
"family": "LeCun",
"sequence": "first"
},
{
"given": "Yoshua",
"family": "Bengio",
"sequence": "additional"
}
],
"container_title": "Nature",
"volume": "521",
"issue": "7553",
"page": "436-444",
"published_date": "2015-05-27",
"year": 2015,
"doi": "10.1038/nature14539",
"issn": ["0028-0836", "1476-4687"],
"language": "en",
"tags": ["machine learning", "neural networks"]
}
},
"citation_mappings": {
"ref_001": [
{"document_id": "doc1", "positions": [45, 89, 234]},
{"document_id": "doc2", "positions": [12, 156]}
]
}
}【English】
Citation Style Configuration (citation_styles.json)
Main configuration items:
{
"apa": {
"name": "APA 7th Edition",
"in_text": {
"format": "parenthetical",
"separator": ", ",
"year_position": "after_author",
"page_number_format": "p. {page}"
},
"bibliography": {
"sort_by": "alphabetical",
"author_format": "last_first",
"title_format": "sentence_case",
"doi_required": true,
"url_format": "available_at"
},
"document_type_codes": {
"book": "",
"journal": "",
"conference": "",
"thesis": "",
"report": ""
}
},
"gbt7714": {
"name": "GB/T 7714-2015",
"in_text": {
"format": "numeric_bracket",
"separator": ", ",
"prefix": "[",
"suffix": "]"
},
"bibliography": {
"sort_by": "citation_order",
"author_format": "chinese_name_order",
"document_type_codes": {
"book": "[M]",
"journal": "[J]",
"conference": "[C]",
"thesis": "[D]",
"report": "[R]",
"newspaper": "[N]"
}
},
"chinese_authors": {
"name_order": "last_first",
"separator": ", "
}
}
}Crossref API Configuration (crossref_config.json)
{
"api_base": "https://api.crossref.org",
"endpoints": {
"works": "/works",
"journals": "/journals",
"types": "/types",
"fields": "/fields"
},
"rate_limiting": {
"requests_per_second": 10,
"backoff_strategy": "exponential",
"max_retries": 3
},
"filters": {
"default": {
"has_full_text": true,
"state": "active"
},
"journals": {
"type": "journal-article"
},
"books": {
"type": "monograph"
}
},
"cache": {
"enabled": true,
"ttl_seconds": 86400,
"max_size": 1000
}
}Local Reference Database Configuration (reference_database.json)
{
"metadata": {
"version": "1.0.0",
"created_date": "2026-03-01",
"last_updated": "2026-03-01"
},
"references": {
"ref_001": {
"id": "ref_001",
"type": "journal_article",
"title": "Deep Learning",
"authors": [
{
"given": "Yann",
"family": "LeCun",
"sequence": "first"
},
{
"given": "Yoshua",
"family": "Bengio",
"sequence": "additional"
}
],
"container_title": "Nature",
"volume": "521",
"issue": "7553",
"page": "436-444",
"published_date": "2015-05-27",
"year": 2015,
"doi": "10.1038/nature14539",
"issn": ["0028-0836", "1476-4687"],
"language": "en",
"tags": ["machine learning", "neural networks"]
}
},
"citation_mappings": {
"ref_001": [
{"document_id": "doc1", "positions": [45, 89, 234]},
{"document_id": "doc2", "positions": [12, 156]}
]
}
}使用示例 | Usage Examples
【中文】
示例1:通过DOI获取文献并生成APA格式引用
from academic_citation_skill import AcademicCitationManager
manager = AcademicCitationManager()
# 获取文献信息
ref = manager.fetch_reference_by_doi("10.1038/nature14539")
# 生成文中引用
in_text = manager.generate_citation(
reference_id=ref['id'],
style='apa',
citation_type='author-date'
)
print(f"文中引用: {in_text}")
# 输出: (LeCun, Bengio, & Hinton, 2015)
# 生成参考文献条目
bib_entry = manager.format_bibliography_entry(ref, style='apa')
print(f"参考文献条目: {bib_entry}")
# 输出: LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539示例2:批量导入和格式转换
# 从BibTeX文件导入
from batch_import import BatchImporter
importer = BatchImporter()
references = importer.import_from_file('references.bib', format='bibtex')
# 添加到文献库
for ref in references:
manager.add_to_library(ref)
# 转换为IEEE格式
from format_converter import FormatConverter
converter = FormatConverter()
ieee_refs = converter.convert_references(
references=references,
from_style='bibtex',
to_style='ieee'
)
# 保存到文件
converter.save_to_file(ieee_refs, 'references_ieee.txt')示例3:检查论文引用完整性
from citation_checker import CitationChecker
checker = CitationChecker()
# 加载论文和参考文献
with open('paper.txt', 'r') as f:
paper_text = f.read()
# 检查引用完整性
report = checker.check_paper(
document_text=paper_text,
bibliography_file='references.txt'
)
# 打印报告
print(f"总引用数: {report['total_citations']}")
print(f"参考文献数: {report['total_references']}")
print(f"不一致项: {report['inconsistencies']}")
if report['missing_in_bibliography']:
print("\n文中引用但参考文献列表缺失:")
for item in report['missing_in_bibliography']:
print(f" - {item}")
if report['unused_references']:
print("\n参考文献列表中未使用的条目:")
for item in report['unused_references']:
print(f" - {item}")示例4:处理中文文献(GB/T 7714格式)
# 添加中文文献
chinese_ref = {
'type': 'journal_article',
'title': '深度学习在自然语言处理中的应用',
'authors': [
{'family': '张', 'given': '三'},
{'family': '李', 'given': '四'}
],
'container_title': '计算机学报',
'year': 2023,
'volume': 46,
'issue': 3,
'page': '1-15',
'language': 'zh'
}
manager.add_to_library(chinese_ref)
# 生成GB/T 7714格式引用
bib_entry = manager.format_bibliography_entry(
chinese_ref,
style='gbt7714'
)
print(f"GB/T 7714格式: {bib_entry}")
# 输出: 张三, 李四. 深度学习在自然语言处理中的应用[J]. 计算机学报, 2023, 46(3): 1-15.【English】
Example 1: Fetch Reference by DOI and Generate APA Format Citation
from academic_citation_skill import AcademicCitationManager
manager = AcademicCitationManager()
# Fetch reference information
ref = manager.fetch_reference_by_doi("10.1038/nature14539")
# Generate in-text citation
in_text = manager.generate_citation(
reference_id=ref['id'],
style='apa',
citation_type='author-date'
)
print(f"In-text citation: {in_text}")
# Output: (LeCun, Bengio, & Hinton, 2015)
# Generate bibliography entry
bib_entry = manager.format_bibliography_entry(ref, style='apa')
print(f"Bibliography entry: {bib_entry}")
# Output: LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539Example 2: Batch Import and Format Conversion
# Import from BibTeX file
from batch_import import BatchImporter
importer = BatchImporter()
references = importer.import_from_file('references.bib', format='bibtex')
# Add to library
for ref in references:
manager.add_to_library(ref)
# Convert to IEEE format
from format_converter import FormatConverter
converter = FormatConverter()
ieee_refs = converter.convert_references(
references=references,
from_style='bibtex',
to_style='ieee'
)
# Save to file
converter.save_to_file(ieee_refs, 'references_ieee.txt')Example 3: Check Paper Citation Integrity
from citation_checker import CitationChecker
checker = CitationChecker()
# Load paper and references
with open('paper.txt', 'r') as f:
paper_text = f.read()
# Check citation integrity
report = checker.check_paper(
document_text=paper_text,
bibliography_file='references.txt'
)
# Print report
print(f"Total citations: {report['total_citations']}")
print(f"Total references: {report['total_references']}")
print(f"Inconsistencies: {report['inconsistencies']}")
if report['missing_in_bibliography']:
print("\nCited in text but missing from bibliography:")
for item in report['missing_in_bibliography']:
print(f" - {item}")
if report['unused_references']:
print("\nUnused entries in bibliography:")
for item in report['unused_references']:
print(f" - {item}")Example 4: Process Chinese Literature (GB/T 7714 Format)
# Add Chinese reference
chinese_ref = {
'type': 'journal_article',
'title': '深度学习在自然语言处理中的应用',
'authors': [
{'family': '张', 'given': '三'},
{'family': '李', 'given': '四'}
],
'container_title': '计算机学报',
'year': 2023,
'volume': 46,
'issue': 3,
'page': '1-15',
'language': 'zh'
}
manager.add_to_library(chinese_ref)
# Generate GB/T 7714 format citation
bib_entry = manager.format_bibliography_entry(
chinese_ref,
style='gbt7714'
)
print(f"GB/T 7714 format: {bib_entry}")
# Output: 张三, 李四. 深度学习在自然语言处理中的应用[J]. 计算机学报, 2023, 46(3): 1-15.注意事项 | Notes
【中文】
使用建议
1. DOI使用
- DOI(数字对象标识符)是获取准确文献信息的最佳方式
- 建议优先使用DOI而不是仅用标题搜索
- 如果DOI无法解析,可以尝试使用标题和作者信息搜索
2. 文献类型选择
- 正确选择文献类型(期刊、图书、会议论文、学位论文等)
- 不同文献类型在引用格式上有重要差异
- GB/T 7714要求使用文献类型标识符如[M][J][C][D]
3. 中文文献处理
- GB/T 7714-2015是中文文献的标准格式
- 注意中文作者姓名的正确格式(姓前名后)
- 期刊名称建议使用标准中文期刊名称
4. 引用完整性
- 使用引用检查功能确保所有文中引用都有对应的参考文献条目
- 定期检查以避免遗漏或重复
- 删除未使用的参考文献条目
5. 格式转换注意事项
- 不同引用格式的规则差异较大,转换后需人工检查
- 特殊字段(如DOI、URL)在不同格式中的显示方式不同
- 建议在最终定稿前再次核对期刊的具体格式要求
限制说明
1. API限制
- Crossref API有速率限制(建议每秒不超过10次请求)
- 大批量查询时建议使用本地缓存
- 某些出版商的文献可能无法通过Crossref获取
2. 元数据准确性
- Crossref返回的元数据可能存在错误或不完整
- 重要文献建议人工核实关键信息
- 中文文献的元数据可能不如英文文献完整
3. 格式兼容性
- 某些期刊可能使用自定义格式
- 特殊字符(如重音符号)可能需要特别处理
- 非拉丁文字符(中文、阿拉伯文等)的显示可能因格式而异
4. 性能考虑
- 批量处理大量文献时,建议分批进行
- 使用本地缓存可以显著提高处理速度
- 网络延迟会影响在线查询的速度
常见问题
Q: 如何获取文献的DOI? A: DOI通常可以在以下位置找到:
- 文献的首页或PDF文件的第一页
- 期刊网站的文献页面
- Crossref或出版社网站
- 如果找不到,可以尝试使用标题和作者信息搜索
Q: 支持哪些文献导入格式? A: 当前支持:
- BibTeX (.bib)
- EndNote (.enw, .xml)
- RIS (.ris)
- CSV (.csv)
- JSON (.json)
Q: 如何处理没有DOI的文献? A: 可以使用以下方法: 1. 使用ISBN(仅适用于图书) 2. 使用标题和作者信息搜索 3. 手动输入完整文献信息 4. 从PDF或其他来源提取元数据
Q: 转换格式后需要人工检查吗? A: 是的,强烈建议: 1. 核对所有作者姓名 2. 检查标题的大小写格式 3. 验证卷号、期号、页码的格式 4. 确认DOI或URL的显示方式 5. 检查特殊字符和非拉丁文字符的显示
【English】
Usage Recommendations
1. DOI Usage
- DOI (Digital Object Identifier) is the best way to obtain accurate bibliographic information
- It is recommended to prioritize using DOI rather than just searching by title
- If DOI cannot be resolved, try searching using title and author information
2. Document Type Selection
- Correctly select document type (journal article, book, conference paper, thesis, etc.)
- Different document types have significant differences in citation formats
- GB/T 7714 requires document type identifiers such as [M][J][C][D]
3. Chinese Literature Processing
- GB/T 7714-2015 is the standard format for Chinese literature
- Pay attention to the correct format of Chinese author names (surname before given name)
- Use standard Chinese journal names for journal titles
4. Citation Integrity
- Use citation checking function to ensure all in-text citations have corresponding bibliography entries
- Check regularly to avoid omissions or duplicates
- Remove unused bibliography entries
5. Format Conversion Considerations
- Rules vary significantly between different citation formats, manual review after conversion is recommended
- Special fields (such as DOI, URL) display differently in different formats
- It is recommended to verify the journal's specific format requirements before final submission
Limitations
1. API Limitations
- Crossref API has rate limits (recommended not exceeding 10 requests per second)
- Use local cache for large batch queries
- Some publishers' works may not be available through Crossref
2. Metadata Accuracy
- Metadata returned by Crossref may contain errors or be incomplete
- It is recommended to manually verify key information for important references
- Metadata for Chinese literature may not be as complete as for English literature
3. Format Compatibility
- Some journals may use custom formats
- Special characters (such as diacritics) may require special handling
- Display of non-Latin characters (Chinese, Arabic, etc.) may vary by format
4. Performance Considerations
- When processing large volumes of literature in batch, it is recommended to process in batches
- Using local cache can significantly improve processing speed
- Network latency affects the speed of online queries
Frequently Asked Questions
Q: How to get the DOI of a reference? A: DOIs can usually be found in the following locations:
- The first page of the article or PDF file
- The article page on the journal website
- Crossref or publisher websites
- If not found, try searching using title and author information
Q: Which bibliography import formats are supported? A: Currently supported:
- BibTeX (.bib)
- EndNote (.enw, .xml)
- RIS (.ris)
- CSV (.csv)
- JSON (.json)
Q: How to handle references without DOI? A: You can use the following methods: 1. Use ISBN (applicable only to books) 2. Search using title and author information 3. Manually enter complete reference information 4. Extract metadata from PDF or other sources
Q: Is manual review needed after format conversion? A: Yes, strongly recommended: 1. Verify all author names 2. Check title case formatting 3. Validate volume, issue, and page number formats 4. Confirm DOI or URL display 5. Check display of special characters and non-Latin characters
更新日志 | Changelog
Version 1.0.0 (2026-03-01)
【中文】
- 初始版本发布
- 支持6种主要引用格式(APA、MLA、Chicago、GB/T 7714、IEEE、Harvard)
- 集成Crossref API获取真实文献元数据
- 支持DOI、ISBN、标题搜索等多种获取方式
- 实现文中引用和参考文献列表生成
- 提供批量导入、格式转换、引用检查等辅助功能
- 完整支持中英文文献
- 提供Python API和命令行接口
- 包含完整的测试和文档
【English】
- Initial release
- Support for 6 major citation formats (APA, MLA, Chicago, GB/T 7714, IEEE, Harvard)
- Integration with Crossref API to retrieve authentic bibliographic metadata
- Support for multiple retrieval methods including DOI, ISBN, title search
- Implementation of in-text citations and bibliography list generation
- Provision of auxiliary functions such as batch import, format conversion, citation checking
- Complete support for Chinese and English references
- Provision of Python API and command line interface
- Complete testing and documentation included
许可证 | License
【中文】 本项目采用 MIT 许可证。详见 LICENSE 文件。
【English】 This project is licensed under MIT License. See LICENSE file for details.
联系方式 | Contact
【中文】
- 项目主页:https://github.com/YouStudyeveryday/academic-citation-manager
- 问题反馈:https://github.com/YouStudyeveryday/academic-citation-manager/issues
- 技术支持:youstudyeveryday@example.com
【English】
- Project homepage: https://github.com/YouStudyeveryday/academic-citation-manager
- Issue tracker: https://github.com/YouStudyeveryday/academic-citation-manager/issues
- Technical support: youstudyeveryday@example.com
致谢 | Acknowledgments
【中文】 感谢Crossref提供免费的DOI元数据查询服务,以及所有为学术引用管理工具做出贡献的开发者。特别感谢开源社区在citation-style-language等项目上的贡献。
【English】 Thanks to Crossref for providing free DOI metadata query services, and all developers who have contributed to academic citation management tools. Special thanks to the open source community for their contributions to projects like citation-style-language.
#!/usr/bin/env python
# -*- coding: utf-8 -*-
"""
Academic Citation Manager - 学术引用管理器
专业的科研论文引用管理工具,支持多种引用格式和Crossref API集成
"""
import re
import json
import requests
import hashlib
import time
import argparse
from typing import Dict, List, Optional, Union, Any, Tuple
from datetime import datetime
from pathlib import Path
from dataclasses import dataclass, field
from enum import Enum
from abc import ABC, abstractmethod
import logging
# 配置日志
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class CitationStyle(Enum):
"""引用格式枚举"""
APA = "apa"
MLA = "mla"
CHICAGO = "chicago"
CHICAGO_AUTHOR_DATE = "chicago-author-date"
GBT7714 = "gbt7714"
IEEE = "ieee"
HARVARD = "harvard"
BIBTEX = "bibtex"
class DocumentType(Enum):
"""文献类型枚举"""
JOURNAL = "journal_article"
BOOK = "book"
CONFERENCE = "conference_paper"
THESIS = "thesis"
REPORT = "report"
NEWSPAPER = "newspaper_article"
WEBPAGE = "webpage"
class ChineseDocumentType(Enum):
"""中文文献类型枚举(GB/T 7714)"""
MONOGRAPH = "[M]"
JOURNAL_ARTICLE = "[J]"
CONFERENCE_PAPER = "[C]"
DISSERTATION = "[D]"
REPORT = "[R]"
NEWSPAPER = "[N]"
STANDARD = "[S]"
PATENT = "[P]"
DATABASE = "[DB]"
COMPILER = "[CP]"
@dataclass
class Author:
"""作者信息"""
given: str
family: str
sequence: str = "additional"
orcid: Optional[str] = None
affiliation: Optional[str] = None
def __str__(self) -> str:
return f"{self.family} {self.given}"
def format_name(self, style: CitationStyle = CitationStyle.APA) -> str:
"""格式化作者姓名"""
if style == CitationStyle.GBT7714:
return f"{self.family}{self.given}"
elif style in [CitationStyle.IEEE, CitationStyle.CHICAGO]:
return f"{self.family}, {self.given[0]}"
elif style == CitationStyle.MLA:
return f"{self.family}, {self.given}"
else:
return f"{self.family}, {self.given[0]}."
@dataclass
class Reference:
"""参考文献信息"""
id: str
type: str
title: str
authors: List[Author]
year: int
doi: Optional[str] = None
isbn: Optional[str] = None
container_title: Optional[str] = None
volume: Optional[str] = None
issue: Optional[str] = None
page: Optional[str] = None
publisher: Optional[str] = None
url: Optional[str] = None
published_date: Optional[str] = None
language: str = "en"
tags: List[str] = field(default_factory=list)
abstract: Optional[str] = None
issn: Optional[List[str]] = None
def __post_init__(self):
"""初始化后处理"""
if not self.id:
self.id = self._generate_id()
def _generate_id(self) -> str:
"""生成唯一ID"""
content = f"{self.title}_{self.year}"
if self.doi:
content += f"_{self.doi}"
elif self.isbn:
content += f"_{self.isbn}"
return f"ref_{hashlib.md5(content.encode()).hexdigest()[:8]}"
def to_dict(self) -> Dict:
"""转换为字典"""
return {
'id': self.id,
'type': self.type,
'title': self.title,
'authors': [
{'given': a.given, 'family': a.family, 'sequence': a.sequence}
for a in self.authors
],
'year': self.year,
'doi': self.doi,
'isbn': self.isbn,
'container_title': self.container_title,
'volume': self.volume,
'issue': self.issue,
'page': self.page,
'publisher': self.publisher,
'url': self.url,
'published_date': self.published_date,
'language': self.language,
'tags': self.tags,
'abstract': self.abstract,
'issn': self.issn
}
class BaseCitationFormatter(ABC):
"""引用格式化器基类"""
@abstractmethod
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
"""格式化文中引用"""
pass
@abstractmethod
def format_bibliography_entry(self, reference: Reference) -> str:
"""格式化参考文献条目"""
pass
class APAFormatter(BaseCitationFormatter):
"""APA 7th Edition格式化器"""
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
if not reference.authors:
return "Unknown"
first_author = reference.authors[0]
authors_count = len(reference.authors)
if citation_type == "parenthetical":
if authors_count == 1:
citation = f"({first_author.family}, {reference.year})"
elif authors_count == 2:
second_author = reference.authors[1]
citation = f"({first_author.family} & {second_author.family}, {reference.year})"
else:
citation = f"({first_author.family} et al., {reference.year})"
else:
citation = f"{first_author.family} ({reference.year})"
if page:
citation = f"{citation[:-1]}, p. {page})"
return citation
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_authors(reference.authors)
year = reference.year
title = reference.title
if reference.type == "journal_article":
entry = f"{authors_text} ({year}). {title}. {reference.container_title}, {reference.volume}({reference.issue}), {reference.page}."
if reference.doi:
entry += f" https://doi.org/{reference.doi}"
elif reference.type == "book":
entry = f"{authors_text} ({year}). {title}. {reference.publisher}."
elif reference.type == "conference_paper":
entry = f"{authors_text} ({year}). {title}. In {reference.container_title} (pp. {reference.page}). {reference.publisher}."
else:
entry = f"{authors_text} ({year}). {title}. {reference.container_title}."
return entry
def _format_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
if len(authors) == 1:
return f"{authors[0].family}, {authors[0].given[0]}."
elif len(authors) == 2:
return f"{authors[0].family}, {authors[0].given[0]}., & {authors[1].family}, {authors[1].given[0]}."
elif len(authors) <= 7:
formatted = ", ".join([f"{a.family}, {a.given[0]}." for a in authors[:-1]])
formatted += f", & {authors[-1].family}, {authors[-1].given[0]}."
return formatted
else:
return f"{authors[0].family}, {authors[0].given[0]}., et al."
class MLAFormatter(BaseCitationFormatter):
"""MLA 9th Edition格式化器"""
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
if not reference.authors:
return "Unknown"
first_author = reference.authors[0]
authors_count = len(reference.authors)
year_short = str(reference.year)[-2:]
if authors_count == 1:
citation = f"({first_author.family} {year_short})"
elif authors_count <= 3:
authors_text = "-".join([f"{a.family}" for a in reference.authors])
citation = f"({authors_text} {year_short})"
else:
citation = f"({first_author.family} et al. {year_short})"
if page:
citation = f"{citation[:-1]}, {page})"
return citation
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_authors(reference.authors)
title = f'"{reference.title}"'
year = reference.year
if reference.type == "journal_article":
entry = f"{authors_text}. {title}. {reference.container_title}, vol. {reference.volume}, no. {reference.issue}, {year}, pp. {reference.page}."
elif reference.type == "book":
entry = f"{authors_text}. {title}. {reference.publisher}, {year}."
else:
entry = f"{authors_text}. {title}. {reference.container_title}, {year}."
return entry
def _format_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
if len(authors) <= 3:
return ", ".join([f"{a.family}, {a.given}" for a in authors])
else:
return f"{authors[0].family}, {authors[0].given}, et al."
class ChicagoFormatter(BaseCitationFormatter):
"""Chicago 17th Edition格式化器"""
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
if not reference.authors:
return "Unknown"
first_author = reference.authors[0]
authors_count = len(reference.authors)
if citation_type == "notes":
citation = f"{first_author.family}, {reference.title}"
if page:
citation += f", {page}"
citation += "."
else:
if authors_count == 1:
citation = f"({first_author.family} {reference.year})"
if page:
citation = f"{citation[:-1]}, {page})"
else:
citation = f"({first_author.family} et al. {reference.year})"
if page:
citation = f"{citation[:-1]}, {page})"
return citation
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_authors(reference.authors)
year = reference.year
title = f'"{reference.title}"'
if reference.type == "journal_article":
entry = f"{authors_text}. {year}. {title}. {reference.container_title} {reference.volume}, no. {reference.issue} (p. {reference.page})."
elif reference.type == "book":
entry = f"{authors_text}. {year}. {title}. {reference.publisher}: {reference.page}."
else:
entry = f"{authors_text}. {year}. {title}. {reference.container_title}."
return entry
def _format_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
if len(authors) == 1:
return f"{authors[0].family}, {authors[0].given}"
elif len(authors) <= 3:
return ", ".join([f"{a.family}, {a.given}" for a in authors])
else:
return f"{authors[0].family}, {authors[0].given}, et al."
class IEEEFormatter(BaseCitationFormatter):
"""IEEE格式化器"""
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
return f"[{reference.id}]"
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_authors(reference.authors)
year = reference.year
title = f'"{reference.title}"'
if reference.type == "journal_article":
entry = f"[{reference.id}] {authors_text}, {title}, {reference.container_title}, vol. {reference.volume}, no. {reference.issue}, pp. {reference.page}, {year}."
elif reference.type == "book":
entry = f"[{reference.id}] {authors_text}, {title}. {reference.publisher}, {year}."
elif reference.type == "conference_paper":
entry = f"[{reference.id}] {authors_text}, {title}, in {reference.container_title}, {year}, pp. {reference.page}."
else:
entry = f"[{reference.id}] {authors_text}, {title}. {reference.container_title}, {year}."
return entry
def _format_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
author_list = ", ".join([f"{a.family} {a.given[0]}." for a in authors[:3]])
if len(authors) > 3:
author_list += ", et al."
return author_list
class GBT7714Formatter(BaseCitationFormatter):
"""GB/T 7714-2015格式化器(中文文献)"""
def __init__(self):
"""初始化格式化器"""
self.type_codes = {
"book": ChineseDocumentType.MONOGRAPH.value,
"journal_article": ChineseDocumentType.JOURNAL_ARTICLE.value,
"conference_paper": ChineseDocumentType.CONFERENCE_PAPER.value,
"thesis": ChineseDocumentType.DISSERTATION.value,
"report": ChineseDocumentType.REPORT.value,
"newspaper_article": ChineseDocumentType.NEWSPAPER.value,
"standard": ChineseDocumentType.STANDARD.value,
"patent": ChineseDocumentType.PATENT.value,
"database": ChineseDocumentType.DATABASE.value,
"compiler": ChineseDocumentType.COMPILER.value
}
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
if not reference.authors:
return "Unknown"
first_author = reference.authors[0]
authors_count = len(reference.authors)
if citation_type == "numeric":
return f"[{reference.id}]"
else:
if authors_count == 1:
citation = f"({first_author.family}, {reference.year})"
elif authors_count <= 3:
authors_text = ",".join([f"{a.family}" for a in reference.authors])
citation = f"({authors_text}, {reference.year})"
else:
citation = f"({first_author.family} et al., {reference.year})"
if page:
citation = f"{citation[:-1]}, {page})"
return citation
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_chinese_authors(reference.authors)
year = reference.year
title = reference.title
type_code = self.type_codes.get(reference.type, "")
if reference.type == "journal_article":
entry = f"{authors_text}. {title}{type_code}. {reference.container_title}, {year}, {reference.volume}({reference.issue}): {reference.page}."
elif reference.type == "book":
entry = f"{authors_text}. {title}{type_code}. {reference.publisher}, {year}: {reference.page}."
elif reference.type == "conference_paper":
entry = f"{authors_text}. {title}{type_code}. // {reference.container_title}, {year}, {reference.page}."
elif reference.type == "thesis":
entry = f"{authors_text}. {title}{type_code}. {reference.container_title}, {year}."
elif reference.type == "report":
entry = f"{authors_text}. {title}{type_code}. {reference.container_title}, {year}."
elif reference.type == "standard":
entry = f"{title}{type_code}. {reference.container_title}, {year}."
elif reference.type == "patent":
entry = f"{title}{type_code}. {reference.container_title}, {year}."
else:
entry = f"{authors_text}. {title}{type_code}. {reference.container_title}, {year}."
if reference.doi:
entry += f" doi: {reference.doi}."
if reference.url:
entry += f" {reference.url}."
return entry
def _format_chinese_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
formatted = ",".join([f"{a.family}{a.given}" for a in authors])
return formatted
class HarvardFormatter(BaseCitationFormatter):
"""Harvard格式化器"""
def format_in_text_citation(self, reference: Reference,
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
if not reference.authors:
return "Unknown"
first_author = reference.authors[0]
citation = f"({first_author.family}, {reference.year}"
if page:
citation = f"{citation}, p. {page})"
return citation
def format_bibliography_entry(self, reference: Reference) -> str:
authors_text = self._format_authors(reference.authors)
year = reference.year
title = reference.title
if reference.type == "journal_article":
entry = f"{authors_text} ({year}) '{reference.title}', {reference.container_title}, {reference.volume}({reference.issue}), p. {reference.page}."
elif reference.type == "book":
entry = f"{authors_text} ({year}) {title}. {reference.publisher}."
else:
entry = f"{authors_text} ({year}) '{title}', {reference.container_title}."
return entry
def _format_authors(self, authors: List[Author]) -> str:
if not authors:
return "Unknown"
if len(authors) == 1:
return f"{authors[0].family}, {authors[0].given}"
elif len(authors) == 2:
return f"{authors[0].family}, {authors[0].given} and {authors[1].family}, {authors[1].given}"
elif len(authors) <= 3:
return ", ".join([f"{a.family}, {a.given}" for a in authors])
else:
return f"{authors[0].family}, {authors[0].given}, et al."
class CrossrefClient:
"""Crossref API客户端"""
def __init__(self, config_file: str = "crossref_config.json"):
"""初始化Crossref客户端"""
self.config = self._load_config(config_file)
self.cache = {}
self.session = requests.Session()
self.session.headers.update({
'User-Agent': 'AcademicCitationManager/1.0.0 (mailto:youstudyeveryday@example.com)'
})
self.last_request_time = 0
self.min_request_interval = 0.1
def _load_config(self, config_file: str) -> Dict:
"""加载配置文件"""
config_path = Path(__file__).parent / config_file
try:
with open(config_path, 'r', encoding='utf-8') as f:
return json.load(f)
except FileNotFoundError:
return {
"api_base": "https://api.crossref.org",
"rate_limit": 10,
"cache_ttl": 86400
}
def _rate_limit_wait(self):
"""遵守速率限制"""
current_time = time.time()
time_since_last = current_time - self.last_request_time
if time_since_last < self.min_request_interval:
time.sleep(self.min_request_interval - time_since_last)
self.last_request_time = time.time()
def fetch_by_doi(self, doi: str) -> Optional[Dict]:
"""通过DOI获取文献信息"""
cache_key = f"doi_{doi}"
if cache_key in self.cache:
logger.info(f"Using cached result for DOI: {doi}")
return self.cache[cache_key]
self._rate_limit_wait()
try:
url = f"{self.config['api_base']}/works/{doi}"
logger.info(f"Fetching DOI from: {url}")
response = self.session.get(url, timeout=15)
if response.status_code == 200:
data = response.json()
result = self._parse_crossref_response(data)
self.cache[cache_key] = result
logger.info(f"Successfully fetched reference: {result.get('title', 'Unknown')}")
return result
elif response.status_code == 404:
logger.warning(f"DOI not found: {doi}")
else:
logger.error(f"Failed to fetch DOI {doi}: HTTP {response.status_code}")
except requests.exceptions.Timeout:
logger.error(f"Timeout fetching DOI: {doi}")
except Exception as e:
logger.error(f"Error fetching DOI {doi}: {str(e)}")
return None
def search_by_title_author(self, title: str, author: str = None,
year: int = None, max_results: int = 10) -> List[Dict]:
"""通过标题和作者搜索文献"""
self._rate_limit_wait()
try:
params = {
'query.title': title,
'rows': max_results,
'select': 'doi,title,author,type,published-print,container-title,volume,issue,page,publisher,ISSN'
}
if author:
params['query.author'] = author
if year:
params['filter'] = f'from-pub-date:{year},until-pub-date:{year}'
url = f"{self.config['api_base']}/works"
logger.info(f"Searching with params: {params}")
response = self.session.get(url, params=params, timeout=15)
if response.status_code == 200:
data = response.json()
items = data.get('message', {}).get('items', [])
logger.info(f"Found {len(items)} references")
return [self._parse_crossref_response(item) for item in items[:max_results]]
else:
logger.error(f"Search failed: HTTP {response.status_code}")
except Exception as e:
logger.error(f"Error searching: {str(e)}")
return []
def _parse_crossref_response(self, data: Dict) -> Dict:
"""解析Crossref API响应"""
message = data.get('message', {})
try:
authors = []
for item in message.get('author', []):
author = Author(
given=item.get('given', ''),
family=item.get('family', ''),
sequence=item.get('sequence', 'additional')
)
authors.append(author)
title = ''
if 'title' in message and message['title']:
title = message['title'][0]
elif 'short-container-title' in message and message['short-container-title']:
title = message['short-container-title'][0]
year = datetime.now().year
published = message.get('published-print', {})
if 'date-parts' in published and published['date-parts']:
year = published['date-parts'][0][0]
return {
'type': message.get('type', 'journal-article'),
'title': title,
'authors': authors,
'year': year,
'doi': message.get('DOI'),
'container_title': message.get('container-title', [''])[0] if message.get('container-title') else None,
'volume': message.get('volume'),
'issue': message.get('issue'),
'page': message.get('page'),
'publisher': message.get('publisher'),
'published_date': str(published.get('date', '')),
'issn': message.get('ISSN'),
'language': 'en'
}
except Exception as e:
logger.error(f"Error parsing Crossref response: {str(e)}")
return {}
class ReferenceDatabase:
"""本地文献数据库"""
def __init__(self, db_file: str = "reference_database.json"):
"""初始化数据库"""
self.db_file = Path(__file__).parent / db_file
self.data = self._load_database()
self._ensure_db_structure()
def _ensure_db_structure(self):
"""确保数据库结构完整"""
if 'metadata' not in self.data:
self.data['metadata'] = {
'version': '1.0.0',
'created_date': datetime.now().isoformat(),
'last_updated': datetime.now().isoformat()
}
if 'references' not in self.data:
self.data['references'] = {}
if 'citation_mappings' not in self.data:
self.data['citation_mappings'] = {}
def _load_database(self) -> Dict:
"""加载数据库"""
if self.db_file.exists():
try:
with open(self.db_file, 'r', encoding='utf-8') as f:
return json.load(f)
except Exception as e:
logger.error(f"Error loading database: {str(e)}")
return {}
def save_database(self):
"""保存数据库"""
self.data["metadata"]["last_updated"] = datetime.now().isoformat()
try:
with open(self.db_file, 'w', encoding='utf-8') as f:
json.dump(self.data, f, ensure_ascii=False, indent=2)
logger.info(f"Database saved: {self.db_file}")
except Exception as e:
logger.error(f"Error saving database: {str(e)}")
def add_reference(self, reference: Union[Reference, Dict]) -> str:
"""添加参考文献"""
if isinstance(reference, Reference):
ref_data = reference.to_dict()
else:
ref_data = reference
ref_id = ref_data.get('id', '')
if not ref_id:
ref = Reference(**ref_data)
ref_id = ref.id
ref_data['id'] = ref_id
self.data["references"][ref_id] = ref_data
self.save_database()
logger.info(f"Reference added: {ref_id}")
return ref_id
def get_reference(self, ref_id: str) -> Optional[Dict]:
"""获取参考文献"""
return self.data["references"].get(ref_id)
def get_all_references(self) -> List[Dict]:
"""获取所有参考文献"""
return list(self.data["references"].values())
def update_reference(self, ref_id: str, reference: Union[Reference, Dict]):
"""更新参考文献"""
if ref_id in self.data["references"]:
if isinstance(reference, Reference):
ref_data = reference.to_dict()
else:
ref_data = reference
self.data["references"][ref_id] = ref_data
self.save_database()
logger.info(f"Reference updated: {ref_id}")
def delete_reference(self, ref_id: str):
"""删除参考文献"""
if ref_id in self.data["references"]:
del self.data["references"][ref_id]
if ref_id in self.data["citation_mappings"]:
del self.data["citation_mappings"][ref_id]
self.save_database()
logger.info(f"Reference deleted: {ref_id}")
def search_references(self, query: str, field: str = "title") -> List[Dict]:
"""搜索参考文献"""
results = []
query_lower = query.lower()
for ref in self.data["references"].values():
if field in ref and query_lower in str(ref[field]).lower():
results.append(ref)
logger.info(f"Found {len(results)} references matching: {query}")
return results
def add_citation_mapping(self, ref_id: str, document_id: str, position: int):
"""添加引用映射"""
if ref_id not in self.data["citation_mappings"]:
self.data["citation_mappings"][ref_id] = []
mapping = next((m for m in self.data["citation_mappings"][ref_id]
if m["document_id"] == document_id), None)
if mapping:
if position not in mapping["positions"]:
mapping["positions"].append(position)
else:
self.data["citation_mappings"][ref_id].append({
"document_id": document_id,
"positions": [position]
})
self.save_database()
def get_citation_count(self, ref_id: str) -> int:
"""获取引用次数"""
mappings = self.data["citation_mappings"].get(ref_id, [])
return sum(len(m.get("positions", [])) for m in mappings)
def export_to_json(self, output_file: str):
"""导出为JSON格式"""
with open(output_file, 'w', encoding='utf-8') as f:
json.dump(self.data, f, ensure_ascii=False, indent=2)
logger.info(f"Exported to JSON: {output_file}")
def import_from_json(self, input_file: str):
"""从JSON格式导入"""
with open(input_file, 'r', encoding='utf-8') as f:
data = json.load(f)
for ref_id, ref_data in data.get('references', {}).items():
self.data["references"][ref_id] = ref_data
self.data["citation_mappings"] = data.get("citation_mappings", {})
self.save_database()
logger.info(f"Imported from JSON: {input_file}")
class CitationIntegrityChecker:
"""引用完整性检查器"""
def __init__(self):
"""初始化检查器"""
self.citation_patterns = {
'numeric_bracket': re.compile(r'\[(\w+)\]'),
'at_sign': re.compile(r'@(\w+)'),
'parenthetical': re.compile(r'\(([A-Z][a-z]+,\s*\d{4})')
}
self.ref_pattern = re.compile(r'^\d+\.|\[\d+\]')
def check_document(self, document_text: str,
bibliography: List[Dict]) -> Dict:
"""检查文档引用完整性"""
in_text_citations = self._extract_citations(document_text)
bib_ids = {ref['id'] for ref in bibliography}
missing_in_bib = [c for c in in_text_citations if c not in bib_ids]
unused_refs = [ref_id for ref_id in bib_ids if ref_id not in in_text_citations]
report = {
'total_citations': len(in_text_citations),
'total_references': len(bib_ids),
'missing_in_bibliography': missing_in_bib,
'unused_references': unused_refs,
'inconsistencies': len(missing_in_bib) + len(unused_refs),
'issues': []
}
if missing_in_bib:
report['issues'].append({
'type': 'missing_in_bibliography',
'count': len(missing_in_bib),
'items': missing_in_bib
})
if unused_refs:
report['issues'].append({
'type': 'unused_references',
'count': len(unused_refs),
'items': unused_refs
})
logger.info(f"Integrity check: {report['total_citations']} citations, {report['total_references']} references")
return report
def _extract_citations(self, text: str) -> List[str]:
"""提取文中引用"""
citations = []
for pattern_name, pattern in self.citation_patterns.items():
matches = pattern.findall(text)
citations.extend(matches)
return list(set(citations))
class FormatConverter:
"""格式转换器"""
def __init__(self):
"""初始化转换器"""
self.formatters = {
CitationStyle.APA: APAFormatter(),
CitationStyle.MLA: MLAFormatter(),
CitationStyle.CHICAGO: ChicagoFormatter(),
CitationStyle.IEEE: IEEEFormatter(),
CitationStyle.GBT7714: GBT7714Formatter(),
CitationStyle.HARVARD: HarvardFormatter()
}
def convert_reference(self, reference: Dict,
from_style: CitationStyle,
to_style: CitationStyle) -> Dict:
"""转换单个引用格式"""
ref_obj = self._dict_to_reference(reference)
from_formatter = self.formatters.get(from_style)
to_formatter = self.formatters.get(to_style)
if not to_formatter:
raise ValueError(f"Unsupported citation style: {to_style}")
original_entry = from_formatter.format_bibliography_entry(ref_obj) if from_formatter else str(reference)
converted_entry = to_formatter.format_bibliography_entry(ref_obj)
return {
'id': reference.get('id', ''),
'original': original_entry,
'converted': converted_entry,
'from_style': from_style.value,
'to_style': to_style.value
}
def convert_batch(self, references: List[Dict],
from_style: CitationStyle,
to_style: CitationStyle) -> List[Dict]:
"""批量转换引用格式"""
results = []
for ref in references:
try:
result = self.convert_reference(ref, from_style, to_style)
results.append(result)
except Exception as e:
logger.error(f"Error converting reference: {str(e)}")
return results
def _dict_to_reference(self, ref_dict: Dict) -> Reference:
"""将字典转换为Reference对象"""
authors = []
for author_data in ref_dict.get('authors', []):
author = Author(
given=author_data.get('given', ''),
family=author_data.get('family', ''),
sequence=author_data.get('sequence', 'additional')
)
authors.append(author)
return Reference(
id=ref_dict.get('id', ''),
type=ref_dict.get('type', 'journal_article'),
title=ref_dict.get('title', ''),
authors=authors,
year=ref_dict.get('year', datetime.now().year),
doi=ref_dict.get('doi'),
isbn=ref_dict.get('isbn'),
container_title=ref_dict.get('container_title'),
volume=ref_dict.get('volume'),
issue=ref_dict.get('issue'),
page=ref_dict.get('page'),
publisher=ref_dict.get('publisher'),
url=ref_dict.get('url'),
published_date=ref_dict.get('published_date'),
language=ref_dict.get('language', 'en'),
tags=ref_dict.get('tags', []),
abstract=ref_dict.get('abstract'),
issn=ref_dict.get('issn')
)
def export_bibtex(self, references: List[Dict]) -> str:
"""导出为BibTeX格式"""
bibtex_lines = []
for ref in references:
key = ref.get('id', 'unknown')
entry_type = self._get_bibtex_type(ref.get('type', 'misc'))
bibtex = f"@{entry_type}{{{key},\n"
fields = []
for k, v in ref.items():
if k not in ['id', 'type']:
if v:
if k == 'authors':
authors = ", and ".join([f"{a.get('family', '')}, {a.get('given', '')}"
for a in v])
fields.append(f"author = {{{authors}}}")
elif k == 'container_title':
fields.append(f"journal = {{{v}}}")
elif k == 'year':
fields.append(f"year = {{{v}}}")
else:
fields.append(f"{k} = {{{v}}}")
bibtex += ",\n".join(fields)
bibtex += "\n}\n\n"
bibtex_lines.append(bibtex)
return ''.join(bibtex_lines)
def _get_bibtex_type(self, doc_type: str) -> str:
"""获取BibTeX条目类型"""
type_mapping = {
'journal_article': 'article',
'book': 'book',
'conference_paper': 'inproceedings',
'thesis': 'phdthesis',
'report': 'techreport',
'newspaper_article': 'article'
}
return type_mapping.get(doc_type, 'misc')
class AcademicCitationManager:
"""学术引用管理器 - 主类"""
def __init__(self, config_dir: Optional[str] = None):
"""初始化管理器"""
if config_dir:
config_path = Path(config_dir)
else:
config_path = Path(__file__).parent
self.crossref_client = CrossrefClient(str(config_path / "crossref_config.json"))
self.database = ReferenceDatabase(str(config_path / "reference_database.json"))
self.checker = CitationIntegrityChecker()
self.converter = FormatConverter()
logger.info("AcademicCitationManager initialized")
def fetch_reference_by_doi(self, doi: str) -> Optional[Dict]:
"""通过DOI获取文献信息"""
logger.info(f"Fetching reference by DOI: {doi}")
data = self.crossref_client.fetch_by_doi(doi)
if data:
ref_id = self.database.add_reference(data)
data['id'] = ref_id
return data
return None
def fetch_reference_by_isbn(self, isbn: str) -> Optional[Dict]:
"""通过ISBN获取图书信息"""
logger.info(f"Fetching reference by ISBN: {isbn}")
try:
url = f"https://openlibrary.org/api/books?bibkeys=ISBN:{isbn}&format=json&jscmd=data"
response = requests.get(url, timeout=15)
if response.status_code == 200:
data = response.json()
if data:
book_data = list(data.values())[0] if data else {}
ref_data = self._parse_openlibrary_data(book_data)
ref_id = self.database.add_reference(ref_data)
ref_data['id'] = ref_id
return ref_data
except Exception as e:
logger.error(f"Error fetching ISBN {isbn}: {str(e)}")
return None
def _parse_openlibrary_data(self, data: Dict) -> Dict:
"""解析OpenLibrary数据"""
try:
authors = []
for author_data in data.get('authors', []):
author = Author(
given=author_data.get('name', '').split(' ')[-1],
family=author_data.get('name', '').split(' ')[0]
)
authors.append(author)
return {
'type': 'book',
'title': data.get('title', ''),
'authors': authors,
'year': int(data.get('publish_date', '2026')[:4]) if data.get('publish_date') else 2026,
'isbn': data.get('isbn_13', [None])[0] or data.get('isbn_10', [None])[0],
'publisher': data.get('publishers', [''])[0] if data.get('publishers') else None,
'page': data.get('number_of_pages'),
'language': 'en'
}
except Exception as e:
logger.error(f"Error parsing OpenLibrary data: {str(e)}")
return {}
def search_references(self, title: str, author: Optional[str] = None,
year: Optional[int] = None, max_results: int = 10) -> List[Dict]:
"""搜索参考文献"""
logger.info(f"Searching references: {title}")
results = self.crossref_client.search_by_title_author(title, author, year, max_results)
for result in results:
ref_id = self.database.add_reference(result)
result['id'] = ref_id
return results
def add_to_library(self, reference: Union[Reference, Dict]) -> str:
"""添加文献到本地库"""
if isinstance(reference, Reference):
ref_data = reference.to_dict()
else:
ref_data = reference
ref_id = self.database.add_reference(ref_data)
logger.info(f"Reference added to library: {ref_id}")
return ref_id
def generate_citation(self, reference_id: str,
style: str = "apa",
citation_type: str = "parenthetical",
page: Optional[str] = None) -> str:
"""生成文中引用"""
try:
style_enum = CitationStyle(style)
except ValueError:
logger.error(f"Unsupported citation style: {style}")
return f"[{reference_id}]"
ref_data = self.database.get_reference(reference_id)
if not ref_data:
logger.warning(f"Reference not found: {reference_id}")
return f"[{reference_id}]"
reference = self.converter._dict_to_reference(ref_data)
formatter = self.converter.formatters.get(style_enum, self.converter.formatters[CitationStyle.APA])
citation = formatter.format_in_text_citation(reference, citation_type, page)
return citation
def generate_bibliography(self, style: str = "apa",
sort_by: str = "alphabetical") -> List[str]:
"""生成参考文献列表"""
try:
style_enum = CitationStyle(style)
except ValueError:
logger.error(f"Unsupported citation style: {style}")
style_enum = CitationStyle.APA
references = self.database.get_all_references()
if sort_by == "alphabetical":
references.sort(key=lambda x: (x.get('title', ''), x.get('year', 0)))
elif sort_by == "citation_order":
references.sort(key=lambda x: x.get('id', ''))
elif sort_by == "year":
references.sort(key=lambda x: x.get('year', 0), reverse=True)
formatter = self.converter.formatters.get(style_enum, self.converter.formatters[CitationStyle.APA])
bibliography = []
for ref_dict in references:
try:
reference = self.converter._dict_to_reference(ref_dict)
entry = formatter.format_bibliography_entry(reference)
bibliography.append(entry)
except Exception as e:
logger.error(f"Error formatting reference: {str(e)}")
logger.info(f"Generated bibliography: {len(bibliography)} entries")
return bibliography
def check_citation_integrity(self, document_text: str) -> Dict:
"""检查引用完整性"""
references = self.database.get_all_references()
return self.checker.check_document(document_text, references)
def convert_citation_style(self, references: List[Dict],
from_style: str, to_style: str) -> List[Dict]:
"""转换引用格式"""
try:
from_style_enum = CitationStyle(from_style)
to_style_enum = CitationStyle(to_style)
except ValueError as e:
logger.error(f"Unsupported citation style: {str(e)}")
return []
return self.converter.convert_batch(references, from_style_enum, to_style_enum)
def export_bibtex(self, output_file: str) -> bool:
"""导出为BibTeX格式"""
references = self.database.get_all_references()
bibtex_content = self.converter.export_bibtex(references)
try:
with open(output_file, 'w', encoding='utf-8') as f:
f.write(bibtex_content)
logger.info(f"Exported BibTeX: {output_file}")
return True
except Exception as e:
logger.error(f"Error exporting BibTeX: {str(e)}")
return False
def import_from_bibtex(self, input_file: str) -> int:
"""从BibTeX格式导入"""
try:
with open(input_file, 'r', encoding='utf-8') as f:
bibtex_content = f.read()
imported_count = 0
entries = self._parse_bibtex(bibtex_content)
for entry in entries:
ref_id = self.database.add_reference(entry)
imported_count += 1
logger.info(f"Imported {imported_count} references from BibTeX")
return imported_count
except Exception as e:
logger.error(f"Error importing BibTeX: {str(e)}")
return 0
def _parse_bibtex(self, bibtex_content: str) -> List[Dict]:
"""解析BibTeX内容"""
entries = []
entry_pattern = re.compile(r'@(\w+)\s*\{([^,]+),\s*(.*?)\n\s*\}', re.DOTALL)
for match in entry_pattern.finditer(bibtex_content):
entry_type = match.group(1)
entry_key = match.group(2)
entry_content = match.group(3)
fields = {}
field_pattern = re.compile(r'(\w+)\s*=\s*\{(.*?)\}')
for field_match in field_pattern.finditer(entry_content):
field_name = field_match.group(1)
field_value = field_match.group(2).strip()
fields[field_name] = field_value
entry = {
'id': entry_key,
'type': self._map_bibtex_type(entry_type),
'title': fields.get('title', ''),
'authors': self._parse_bibtex_authors(fields.get('author', '')),
'year': int(fields.get('year', datetime.now().year)),
'journal': fields.get('journal') or fields.get('journaltitle'),
'volume': fields.get('volume'),
'number': fields.get('number'),
'pages': fields.get('pages'),
'publisher': fields.get('publisher'),
'doi': fields.get('doi'),
'isbn': fields.get('isbn') or fields.get('isbn13'),
'language': fields.get('language', 'en')
}
entries.append(entry)
return entries
def _map_bibtex_type(self, bibtex_type: str) -> str:
"""映射BibTeX类型到内部类型"""
type_mapping = {
'article': 'journal_article',
'book': 'book',
'inproceedings': 'conference_paper',
'phdthesis': 'thesis',
'mastersthesis': 'thesis',
'techreport': 'report',
'misc': 'webpage'
}
return type_mapping.get(bibtex_type.lower(), 'webpage')
def _parse_bibtex_authors(self, author_string: str) -> List[Dict]:
"""解析BibTeX作者字符串"""
authors = []
if not author_string:
return authors
author_parts = re.split(r'\s+and\s+', author_string)
for part in author_parts:
name_parts = re.split(r',\s*', part.strip())
if len(name_parts) >= 2:
family = name_parts[0].strip()
given = name_parts[1].strip()
authors.append({
'family': family,
'given': given,
'sequence': 'additional'
})
return authors
def get_library_stats(self) -> Dict:
"""获取文献库统计信息"""
references = self.database.get_all_references()
type_counts = {}
year_counts = {}
language_counts = {}
for ref in references:
ref_type = ref.get('type', 'unknown')
type_counts[ref_type] = type_counts.get(ref_type, 0) + 1
year = ref.get('year', 0)
year_counts[year] = year_counts.get(year, 0) + 1
language = ref.get('language', 'en')
language_counts[language] = language_counts.get(language, 0) + 1
return {
'total_references': len(references),
'type_distribution': type_counts,
'year_distribution': year_counts,
'language_distribution': language_counts
}
def export_to_json(self, output_file: str) -> bool:
"""导出为JSON格式"""
try:
self.database.export_to_json(output_file)
return True
except Exception as e:
logger.error(f"Error exporting to JSON: {str(e)}")
return False
def import_from_json(self, input_file: str) -> int:
"""从JSON格式导入"""
try:
self.database.import_from_json(input_file)
references = self.database.get_all_references()
return len(references)
except Exception as e:
logger.error(f"Error importing from JSON: {str(e)}")
return 0
def main():
"""主函数 - 命令行接口"""
parser = argparse.ArgumentParser(
description='Academic Citation Manager - 为科研论文和毕业论文添加真实参考文献并规范引用标注',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
示例:
通过DOI获取文献信息:
python academic_citation_skill.py --fetch-doi 10.1038/nature14539
搜索文献:
python academic_citation_skill.py --search "Deep Learning" --author "LeCun"
生成参考文献列表:
python academic_citation_skill.py --generate-bib --style apa --sort alphabetical
检查引用完整性:
python academic_citation_skill.py --check document.txt
格式转换:
python academic_citation_skill.py --convert --from-style apa --to-style ieee --input refs.txt
更多信息请访问: https://github.com/YouStudyeveryday/academic-citation-manager
"""
)
parser.add_argument('--fetch-doi', type=str, help='通过DOI获取文献信息')
parser.add_argument('--fetch-isbn', type=str, help='通过ISBN获取图书信息')
parser.add_argument('--search', type=str, help='搜索文献(标题)')
parser.add_argument('--author', type=str, help='作者名称(用于搜索)')
parser.add_argument('--year', type=int, help='出版年份(用于搜索)')
parser.add_argument('--max-results', type=int, default=10, help='最大结果数(默认10)')
parser.add_argument('--style', type=str, default='apa',
help='引用格式 (apa, mla, chicago, ieee, gbt7714, harvard)')
parser.add_argument('--sort', type=str, default='alphabetical',
help='排序方式 (alphabetical, citation_order, year)')
parser.add_argument('--generate-bib', action='store_true', help='从文献库生成参考文献列表')
parser.add_argument('--output', type=str, help='输出文件路径')
parser.add_argument('--convert', action='store_true', help='转换引用格式')
parser.add_argument('--from-style', type=str, help='源格式(用于转换)')
parser.add_argument('--to-style', type=str, help='目标格式(用于转换)')
parser.add_argument('--input', type=str, help='输入文件路径')
parser.add_argument('--check', type=str, help='检查文档引用完整性')
parser.add_argument('--import-bibtex', type=str, help='从BibTeX文件导入')
parser.add_argument('--export-bibtex', type=str, help='导出为BibTeX文件')
parser.add_argument('--import-json', type=str, help='从JSON文件导入')
parser.add_argument('--export-json', type=str, help='导出为JSON文件')
parser.add_argument('--stats', action='store_true', help='显示文献库统计信息')
args = parser.parse_args()
manager = AcademicCitationManager()
try:
if args.fetch_doi:
print(f"正在获取DOI文献信息: {args.fetch_doi}")
ref = manager.fetch_reference_by_doi(args.fetch_doi)
if ref:
print(f"\n文献信息:")
print(f" 标题: {ref.get('title', 'Unknown')}")
print(f" 作者: {', '.join([f\"{a.get('family', '')} {a.get('given', '')}\" for a in ref.get('authors', [])])}")
print(f" 年份: {ref.get('year', 'Unknown')}")
print(f" 类型: {ref.get('type', 'Unknown')}")
print(f" DOI: {ref.get('doi', 'Unknown')}")
print(f" 容器: {ref.get('container_title', 'N/A')}")
print(f" 卷期: {ref.get('volume', 'N/A')}({ref.get('issue', 'N/A')})")
print(f" 页码: {ref.get('page', 'N/A')}")
else:
print("未找到文献信息")
elif args.fetch_isbn:
print(f"正在获取ISBN图书信息: {args.fetch_isbn}")
ref = manager.fetch_reference_by_isbn(args.fetch_isbn)
if ref:
print(f"\n图书信息:")
print(f" 标题: {ref.get('title', 'Unknown')}")
print(f" 作者: {', '.join([f\"{a.get('family', '')} {a.get('given', '')}\" for a in ref.get('authors', [])])}")
print(f" 年份: {ref.get('year', 'Unknown')}")
print(f" ISBN: {ref.get('isbn', 'Unknown')}")
print(f" 出版社: {ref.get('publisher', 'N/A')}")
print(f" 页数: {ref.get('page', 'N/A')}")
else:
print("未找到图书信息")
elif args.search:
print(f"正在搜索文献: {args.search}")
if args.author:
print(f" 作者: {args.author}")
if args.year:
print(f" 年份: {args.year}")
results = manager.search_references(
args.search, args.author, args.year, args.max_results
)
print(f"\n找到 {len(results)} 篇文献:\n")
for i, ref in enumerate(results, 1):
print(f"{i}. {ref.get('title', 'Unknown')} ({ref.get('year', 'Unknown')})")
print(f" 作者: {', '.join([f\"{a.get('family', '')} {a.get('given', '')}\" for a in ref.get('authors', [])[:3]])}")
if len(ref.get('authors', [])) > 3:
print(" 等")
print(f" DOI: {ref.get('doi', 'N/A')}")
print()
elif args.generate_bib:
print(f"正在生成参考文献列表 ({args.style}格式, {args.sort}排序):")
bibliography = manager.generate_bibliography(args.style, args.sort)
output_text = "\n".join([f"{i}. {entry}" for i, entry in enumerate(bibliography, 1)])
if args.output:
with open(args.output, 'w', encoding='utf-8') as f:
f.write(output_text)
print(f"\n参考文献列表已保存到: {args.output}")
else:
print(f"\n{output_text}")
elif args.check:
print(f"正在检查文档引用完整性: {args.check}")
with open(args.check, 'r', encoding='utf-8') as f:
document_text = f.read()
report = manager.check_citation_integrity(document_text)
print(f"\n引用完整性报告:")
print(f" 文中引用总数: {report['total_citations']}")
print(f" 参考文献总数: {report['total_references']}")
print(f" 不一致项数量: {report['inconsistencies']}")
if report['missing_in_bibliography']:
print(f"\n 文中引用但参考文献列表缺失 ({len(report['missing_in_bibliography'])} 项):")
for item in report['missing_in_bibliography']:
print(f" - {item}")
if report['unused_references']:
print(f"\n 参考文献列表中未使用 ({len(report['unused_references'])} 项):")
for item in report['unused_references']:
print(f" - {item}")
if report['inconsistencies'] == 0:
print("\n 引用完整性检查通过!")
elif args.convert and args.input:
print(f"正在转换引用格式: {args.from_style} -> {args.to_style}")
with open(args.input, 'r', encoding='utf-8') as f:
data = json.load(f)
references = data if isinstance(data, list) else data.get('references', [])
converted = manager.convert_citation_style(references, args.from_style, args.to_style)
print(f"\n转换结果 ({len(converted)} 项):\n")
for item in converted:
print(f"ID: {item.get('id', 'unknown')}")
print(f" 原始: {item.get('original', '')[:100]}...")
print(f" 转换后: {item.get('converted', '')[:100]}...")
print()
elif args.import_bibtex:
print(f"正在从BibTeX导入: {args.import_bibtex}")
count = manager.import_from_bibtex(args.import_bibtex)
print(f"\n成功导入 {count} 篇文献")
elif args.export_bibtex:
print(f"正在导出为BibTeX: {args.export_bibtex}")
success = manager.export_bibtex(args.export_bibtex)
if success:
print(f"\n成功导出到: {args.export_bibtex}")
elif args.import_json:
print(f"正在从JSON导入: {args.import_json}")
count = manager.import_from_json(args.import_json)
print(f"\n成功导入 {count} 篇文献")
elif args.export_json:
print(f"正在导出为JSON: {args.export_json}")
success = manager.export_to_json(args.export_json)
if success:
print(f"\n成功导出到: {args.export_json}")
elif args.stats:
print("文献库统计信息:")
stats = manager.get_library_stats()
print(f"\n 总文献数: {stats['total_references']}")
print(f"\n 文献类型分布:")
for type_name, count in stats['type_distribution'].items():
print(f" {type_name}: {count}")
print(f"\n 年份分布:")
for year, count in sorted(stats['year_distribution'].items(), reverse=True)[:10]:
print(f" {year}: {count}")
print(f"\n 语言分布:")
for lang, count in stats['language_distribution'].items():
print(f" {lang}: {count}")
else:
print("未指定操作。使用 --help 查看帮助信息。")
except KeyboardInterrupt:
print("\n操作已取消")
except Exception as e:
logger.error(f"错误: {str(e)}")
import traceback
traceback.print_exc()
if __name__ == '__main__':
main()#!/usr/bin/env python
# -*- coding: utf-8 -*-
"""
批量导入参考文献
支持从多种格式批量导入参考文献到学术引用管理器
"""
import argparse
import json
import re
import time
from pathlib import Path
from typing import List, Dict, Optional, Tuple
import logging
# 配置日志
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class BibTeXParser:
"""BibTeX解析器"""
def __init__(self):
"""初始化解析器"""
self.entry_pattern = re.compile(r'@(\w+)\s*\{([^,]+),\s*(.*?)\n\s*\}', re.DOTALL)
self.field_pattern = re.compile(r'(\w+)\s*=\s*\{(.*?)\}', re.DOTALL)
self.type_mapping = {
'article': 'journal_article',
'book': 'book',
'inproceedings': 'conference_paper',
'phdthesis': 'thesis',
'mastersthesis': 'thesis',
'techreport': 'report',
'misc': 'webpage'
}
def parse_file(self, file_path: str) -> List[Dict]:
"""解析BibTeX文件"""
logger.info(f"解析BibTeX文件: {file_path}")
with open(file_path, 'r', encoding='utf-8') as f:
content = f.read()
entries = []
for match in self.entry_pattern.finditer(content):
entry_type = match.group(1)
entry_key = match.group(2)
entry_content = match.group(3)
entry = {
'id': entry_key,
'type': self.type_mapping.get(entry_type, entry_type),
'original_type': entry_type,
'original_key': entry_key
}
# 解析字段
for field_match in self.field_pattern.finditer(entry_content):
field_name = field_match.group(1)
field_value = field_match.group(2).strip('{}')
# 处理特殊字段
if field_name == 'author' or field_name == 'editor':
entry['authors'] = self._parse_authors(field_value)
elif field_name == 'journal' or field_name == 'journaltitle':
entry['container_title'] = field_value
elif field_name == 'year':
entry['year'] = int(field_value) if field_value.isdigit() else 2026
else:
entry[field_name] = field_value
entries.append(entry)
logger.info(f"成功解析{len(entries)}条BibTeX条目")
return entries
def _parse_authors(self, author_string: str) -> List[Dict]:
"""解析作者字符串"""
authors = []
# 分割多个作者
author_parts = re.split(r'\s+and\s+', author_string)
for part in author_parts:
part = part.strip()
if not part:
continue
# 分析"姓, 名"格式
name_parts = re.split(r',\s*', part)
if len(name_parts) >= 2:
family = name_parts[0].strip()
given = name_parts[1].strip()
authors.append({
'family': family,
'given': given,
'sequence': 'additional'
})
else:
# 其他格式,尝试提取名字
words = part.split()
if len(words) >= 2:
authors.append({
'family': words[-1],
'given': ' '.join(words[:-1]),
'sequence': 'additional'
})
elif words:
authors.append({
'family': words[0],
'given': '',
'sequence': 'additional'
})
# 设置第一个作者为序列"first"
if authors:
authors[0]['sequence'] = 'first'
return authors
class RISParser:
"""RIS格式解析器"""
def __init__(self):
"""初始化解析器"""
self.type_mapping = {
'JOUR': 'journal_article',
'BOOK': 'book',
'CONF': 'conference_paper',
'THES': 'thesis',
'RPRT': 'report',
'NEWS': 'newspaper_article'
}
self.field_mapping = {
'TY': 'type',
'AU': 'authors',
'TI': 'title',
'PY': 'year',
'JO': 'container_title',
'VL': 'volume',
'IS': 'issue',
'SP': 'page',
'PB': 'publisher',
'UR': 'url',
'DO': 'doi',
'SN': 'issn',
'AB': 'abstract'
}
def parse_file(self, file_path: str) -> List[Dict]:
"""解析RIS文件"""
logger.info(f"解析RIS文件: {file_path}")
with open(file_path, 'r', encoding='utf-8') as f:
content = f.read()
# 分割条目
entries_text = re.split(r'\n\s*ER\s*\n', content)
entries = []
entry_id = 0
for entry_text in entries_text:
if not entry_text.strip():
continue
entry_id += 1
entry = {'id': f'imported_ris_{entry_id:04d}'}
# 解析字段
lines = entry_text.split('\n')
for line in lines:
line = line.strip()
if not line or not '=' in line:
continue
parts = line.split('=', 1)
if len(parts) != 2:
continue
field_code = parts[0].strip()
field_value = parts[1].strip()
# 映射字段名
if field_code in self.field_mapping:
mapped_name = self.field_mapping[field_code]
if field_code == 'AU':
entry['authors'] = self._parse_ris_authors(field_value)
elif field_code == 'TY' and field_value in self.type_mapping:
entry['type'] = self.type_mapping[field_value]
elif field_code == 'PY':
try:
entry['year'] = int(field_value[:4])
except ValueError:
entry['year'] = 2026
else:
entry[mapped_name] = field_value
if 'type' not in entry:
entry['type'] = 'webpage'
entries.append(entry)
logger.info(f"成功解析{len(entries)}条RIS条目")
return entries
def _parse_ris_authors(self, author_string: str) -> List[Dict]:
"""解析RIS作者字符串"""
authors = []
# RIS格式中,作者通常由AU字段提供,多个作者用换行分隔
for author_line in author_string.split('\n'):
author_line = author_line.strip()
if not author_line:
continue
# 尝试"姓, 名"格式
name_parts = re.split(r',\s*', author_line)
if len(name_parts) >= 2:
family = name_parts[0].strip()
given = name_parts[1].strip()
authors.append({
'family': family,
'given': given,
'sequence': 'additional'
})
elif name_parts and name_parts[0]:
authors.append({
'family': name_parts[0],
'given': '',
'sequence': 'additional'
})
if authors:
authors[0]['sequence'] = 'first'
return authors
class BatchImporter:
"""批量导入器"""
def __init__(self, manager):
"""初始化导入器"""
self.manager = manager
self.stats = {
'total': 0,
'success': 0,
'failed': 0,
'skipped': 0,
'errors': []
}
def import_from_bibtex(self, file_path: str, validate_doi: bool = True) -> Dict:
"""从BibTeX文件导入"""
logger.info(f"开始从BibTeX导入: {file_path}")
try:
parser = BibTeXParser()
entries = parser.parse_file(file_path)
results = self._process_entries(entries, 'bibtex', validate_doi)
logger.info(f"BibTeX导入完成: 成功{results['success']}, 失败{results['failed']}, 跳过{results['skipped']}")
return results
except Exception as e:
error_msg = f"BibTeX导入失败: {str(e)}"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
def import_from_ris(self, file_path: str, validate_doi: bool = True) -> Dict:
"""从RIS文件导入"""
logger.info(f"开始从RIS导入: {file_path}")
try:
parser = RISParser()
entries = parser.parse_file(file_path)
results = self._process_entries(entries, 'ris', validate_doi)
logger.info(f"RIS导入完成: 成功{results['success']}, 失败{results['failed']}, 跳过{results['skipped']}")
return results
except Exception as e:
error_msg = f"RIS导入失败: {str(e)}"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
def import_from_json(self, file_path: str, validate_doi: bool = True) -> Dict:
"""从JSON文件导入"""
logger.info(f"开始从JSON导入: {file_path}")
try:
with open(file_path, 'r', encoding='utf-8') as f:
data = json.load(f)
entries = data if isinstance(data, list) else data.get('references', [])
if not isinstance(entries, list):
error_msg = f"JSON格式错误: references应为列表"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
results = self._process_entries(entries, 'json', validate_doi)
logger.info(f"JSON导入完成: 成功{results['success']}, 失败{results['failed']}, 跳过{results['skipped']}")
return results
except json.JSONDecodeError as e:
error_msg = f"JSON解析失败: {str(e)}"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
except Exception as e:
error_msg = f"JSON导入失败: {str(e)}"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
def import_from_csv(self, file_path: str, validate_doi: bool = True) -> Dict:
"""从CSV文件导入"""
logger.info(f"开始从CSV导入: {file_path}")
try:
import csv
entries = []
with open(file_path, 'r', encoding='utf-8') as f:
reader = csv.DictReader(f)
for row in reader:
# 映射CSV列名到内部字段名
entry = {
'id': row.get('id', ''),
'type': row.get('type', 'webpage'),
'title': row.get('title', ''),
'authors': self._parse_csv_authors(row.get('authors', '')),
'year': int(row.get('year', 2026)) if row.get('year', '').isdigit() else 2026,
'doi': row.get('doi', ''),
'container_title': row.get('journal', row.get('container_title', '')),
'volume': row.get('volume', ''),
'issue': row.get('issue', ''),
'page': row.get('page', ''),
'publisher': row.get('publisher', ''),
'url': row.get('url', ''),
'language': row.get('language', 'en')
}
entries.append(entry)
results = self._process_entries(entries, 'csv', validate_doi)
logger.info(f"CSV导入完成: 成功{results['success']}, 失败{results['failed']}, 跳过{results['skipped']}")
return results
except Exception as e:
error_msg = f"CSV导入失败: {str(e)}"
logger.error(error_msg)
self.stats['errors'].append(error_msg)
return {'total': 0, 'success': 0, 'failed': 1, 'skipped': 0, 'errors': [error_msg]}
def import_from_directory(self, directory: str,
file_pattern: str = "*.{bib,ris,json,csv}",
validate_doi: bool = True) -> Dict:
"""从目录批量导入"""
logger.info(f"开始从目录导入: {directory}")
dir_path = Path(directory)
if not dir_path.exists():
error_msg = f"目录不存在: {directory}"
logger.error(error_msg)
return {'total': 0, 'success': 0, 'failed': 0, 'skipped': 0, 'errors': [error_msg]}
# 查找所有匹配的文件
files = list(dir_path.glob(file_pattern))
logger.info(f"找到{len(files)}个文件")
total_results = {
'total': 0,
'success': 0,
'failed': 0,
'skipped': 0,
'errors': [],
'file_results': {}
}
for file_path in files:
file_str = str(file_path)
suffix = file_path.suffix.lower()
try:
if suffix == '.bib':
result = self.import_from_bibtex(file_str, validate_doi)
elif suffix == '.ris':
result = self.import_from_ris(file_str, validate_doi)
elif suffix == '.json':
result = self.import_from_json(file_str, validate_doi)
elif suffix == '.csv':
result = self.import_from_csv(file_str, validate_doi)
else:
logger.warning(f"跳过不支持的文件类型: {file_str}")
result = {'total': 0, 'success': 0, 'failed': 0, 'skipped': 1, 'errors': []}
total_results['total'] += result['total']
total_results['success'] += result['success']
total_results['failed'] += result['failed']
total_results['skipped'] += result['skipped']
total_results['errors'].extend(result['errors'])
total_results['file_results'][file_path] = result
except Exception as e:
error_msg = f"导入文件{file_path}失败: {str(e)}"
logger.error(error_msg)
total_results['errors'].append(error_msg)
logger.info(f"目录导入完成: 总计{total_results['total']}, 成功{total_results['success']}, 失败{total_results['failed']}, 跳过{total_results['skipped']}")
return total_results
def _process_entries(self, entries: List[Dict],
source_type: str,
validate_doi: bool) -> Dict:
"""处理导入的条目"""
results = {
'total': len(entries),
'success': 0,
'failed': 0,
'skipped': 0,
'errors': [],
'imported_ids': []
}
for entry in entries:
try:
# 验证必需字段
if not entry.get('title'):
results['skipped'] += 1
results['errors'].append(f"跳过无标题条目: {entry.get('id', 'unknown')}")
continue
# 生成ID
if not entry.get('id'):
content = entry.get('title', '')
if entry.get('year'):
content += f"_{entry['year']}"
if entry.get('doi'):
content += f"_{entry['doi']}"
import hashlib
entry['id'] = f"import_{hashlib.md5(content.encode()).hexdigest()[:8]}"
# 验证DOI
if validate_doi and entry.get('doi'):
if not self._validate_doi(entry['doi']):
results['failed'] += 1
results['errors'].append(f"无效DOI: {entry['doi']} (ID: {entry['id']})")
continue
# 添加到数据库
ref_id = self.manager.database.add_reference(entry)
results['success'] += 1
results['imported_ids'].append(ref_id)
logger.debug(f"成功导入: {entry['title']} (ID: {ref_id})")
except Exception as e:
error_msg = f"导入条目失败 (ID: {entry.get('id', 'unknown')}): {str(e)}"
logger.error(error_msg)
results['failed'] += 1
results['errors'].append(error_msg)
return results
def _parse_csv_authors(self, author_string: str) -> List[Dict]:
"""解析CSV作者字符串"""
authors = []
if not author_string:
return authors
# CSV中作者可能用分号或and分隔
parts = re.split(r'[;,]\s*and\s*', author_string)
for i, part in enumerate(parts):
part = part.strip()
if not part:
continue
# 尝试提取姓名
words = part.split()
if len(words) >= 2:
family = words[-1]
given = ' '.join(words[:-1])
elif words:
family = words[0]
given = ' '.join(words[1:])
else:
continue
authors.append({
'family': family,
'given': given,
'sequence': 'first' if i == 0 else 'additional'
})
return authors
def _validate_doi(self, doi: str) -> bool:
"""验证DOI格式"""
if not doi:
return True
# 基本DOI格式验证
doi_pattern = r'^10\.\d{4,}/.+'
return bool(re.match(doi_pattern, doi))
def generate_report(self) -> str:
"""生成导入报告"""
report = []
report.append("=" * 70)
report.append("批量导入报告")
report.append("=" * 70)
report.append(f"\n总计: {self.stats['total']}")
report.append(f"成功: {self.stats['success']}")
report.append(f"失败: {self.stats['failed']}")
report.append(f"跳过: {self.stats['skipped']}")
if self.stats['errors']:
report.append("\n错误:")
for error in self.stats['errors'][:20]: # 只显示前20个错误
report.append(f" - {error}")
if len(self.stats['errors']) > 20:
report.append(f" ... 还有{len(self.stats['errors']) - 20}个错误未显示")
report.append("\n" + "=" * 70)
return '\n'.join(report)
def export_report(self, output_file: str):
"""导出报告到文件"""
report = self.generate_report()
try:
with open(output_file, 'w', encoding='utf-8') as f:
f.write(report)
logger.info(f"报告已导出: {output_file}")
return True
except Exception as e:
logger.error(f"导出报告失败: {str(e)}")
return False
def main():
"""主函数"""
import argparse
parser = argparse.ArgumentParser(
description='批量导入参考文献到学术引用管理器',
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
示例:
导入BibTeX文件:
python batch_import.py --bibtex references.bib
导入RIS文件:
python batch_import.py --ris references.ris
导入JSON文件:
python batch_import.py --json references.json
导入整个目录:
python batch_import.py --directory ./references
验证DOI并生成报告:
python batch_import.py --bibtex references.bib --validate-doi --report import_report.txt
更多信息请访问: https://github.com/YouStudyeveryday/academic-citation-manager
"""
)
parser.add_argument('--bibtex', type=str, help='BibTeX文件路径')
parser.add_argument('--ris', type=str, help='RIS文件路径')
parser.add_argument('--json', type=str, help='JSON文件路径')
parser.add_argument('--csv', type=str, help='CSV文件路径')
parser.add_argument('--directory', type=str, help='导入目录路径')
parser.add_argument('--pattern', type=str, default='*.{bib,ris,json,csv}',
help='目录中的文件模式(默认: *.{bib,ris,json,csv})')
parser.add_argument('--validate-doi', action='store_true',
help='验证DOI格式')
parser.add_argument('--report', type=str,
help='导出导入报告到指定文件')
parser.add_argument('--verbose', '-v', action='store_true',
help='显示详细输出')
args = parser.parse_args()
# 设置日志级别
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
# 导入管理器
try:
from academic_citation_skill import AcademicCitationManager
manager = AcademicCitationManager()
importer = BatchImporter(manager)
# 执行导入
if args.bibtex:
results = importer.import_from_bibtex(args.bibtex, args.validate_doi)
importer.stats.update(results)
elif args.ris:
results = importer.import_from_ris(args.ris, args.validate_doi)
importer.stats.update(results)
elif args.json:
results = importer.import_from_json(args.json, args.validate_doi)
importer.stats.update(results)
elif args.csv:
results = importer.import_from_csv(args.csv, args.validate_doi)
importer.stats.update(results)
elif args.directory:
results = importer.import_from_directory(args.directory, args.pattern, args.validate_doi)
importer.stats.update(results)
else:
print("错误: 必须指定输入文件或目录")
print("使用 --help 查看帮助信息")
return 1
# 生成报告
if args.report:
success = importer.export_report(args.report)
return 0 if success else 1
else:
print(importer.generate_report())
return 0
except ImportError as e:
print(f"错误: 无法导入AcademicCitationManager: {str(e)}")
print("请确保academic_citation_skill.py在同一目录下")
return 1
except Exception as e:
logging.error(f"错误: {str(e)}")
import traceback
traceback.print_exc()
return 1
if __name__ == '__main__':
import sys
sys.exit(main()){
"api_base": "https://api.crossref.org",
"api_version": "v1",
"endpoints": {
"works": "/works",
"works_with_doi": "/works/{doi}",
"journals": "/journals",
"types": "/types",
"fields": "/fields",
"funders": "/funders",
"members": "/members",
"prefixes": "/prefixes",
"deposits": "/deposits"
},
"request_config": {
"user_agent": "AcademicCitationManager/1.0.0 (mailto:youstudyeveryday@example.com)",
"timeout": 15,
"max_retries": 3,
"retry_backoff": 1.5,
"default_rows": 20,
"select": "doi,title,author,type,published-print,container-title,volume,issue,page,publisher,ISSN,ISBN"
},
"rate_limiting": {
"enabled": true,
"requests_per_second": 10,
"requests_per_minute": 600,
"backoff_strategy": "exponential",
"max_retries": 3,
"initial_backoff": 1.0,
"max_backoff": 60.0
},
"filters": {
"default": {
"has_full_text": true,
"state": "active"
},
"journal_articles": {
"type": "journal-article"
},
"books": {
"type": "monograph"
},
"conferences": {
"type": "proceedings"
},
"datasets": {
"type": "dataset"
},
"reports": {
"type": "report"
},
"standards": {
"type": "standard"
},
"components": {
"type": "component"
}
},
"sort_options": {
"published": "published",
"published-print": "published-print",
"is-referenced-by-count": "is-referenced-by-count",
"score": "score"
},
"cache": {
"enabled": true,
"type": "memory",
"ttl_seconds": 86400,
"max_size": 1000,
"clear_on_start": false
},
"metadata_fields": {
"author": {
"fields": ["given", "family", "sequence", "ORCID", "affiliation"],
"prefix": "author"
},
"title": {
"fields": ["title", "short-container-title", "subtitle", "original-title"],
"prefix": "title"
},
"publication": {
"fields": ["type", "published", "published-print", "published-online", "deposited", "indexed"],
"prefix": ""
},
"container": {
"fields": ["container-title", "short-container-title", "ISSN", "ISBN", "publisher", "member"],
"prefix": ""
},
"details": {
"fields": ["volume", "issue", "page", "article-number", "publisher-location", "license"],
"prefix": ""
},
"identifiers": {
"fields": ["DOI", "URL", "PMID", "PMCID", "arXiv"],
"prefix": ""
}
},
"response_parsing": {
"handle_empty_author": "Unknown",
"handle_missing_title": "Untitled",
"handle_multiple_titles": "use_first",
"normalize_author_names": true,
"extract_year_from_date": true
},
"error_handling": {
"retry_on_timeout": true,
"retry_on_5xx": true,
"log_errors": true,
"error_log_file": "crossref_errors.log",
"max_error_log_size": 1048576
}
}{
"metadata": {
"version": "1.0.0",
"created_date": "2026-03-01T00:00:00Z",
"last_updated": "2026-03-01T00:00:00Z",
"schema_version": "1.0",
"total_references": 8,
"supported_languages": ["zh", "en"],
"supported_styles": ["apa", "mla", "chicago", "ieee", "gbt7714", "harvard"]
},
"references": {
"ref_001": {
"id": "ref_001",
"type": "journal_article",
"title": "Deep Learning",
"authors": [
{
"given": "Yann",
"family": "LeCun",
"sequence": "first"
},
{
"given": "Yoshua",
"family": "Bengio",
"sequence": "additional"
},
{
"given": "Geoffrey",
"family": "Hinton",
"sequence": "additional"
}
],
"container_title": "Nature",
"volume": "521",
"issue": "7553",
"page": "436-444",
"published_date": "2015-05-27",
"year": 2015,
"doi": "10.1038/nature14539",
"issn": ["0028-0836", "1476-4687"],
"publisher": "Nature Publishing Group",
"language": "en",
"tags": ["machine learning", "neural networks", "deep learning"],
"abstract": "Deep learning allows computational models that are composed of multiple processing layers to learn representations of data with multiple levels of abstraction."
},
"ref_002": {
"id": "ref_002",
"type": "book",
"title": "Artificial Intelligence: A Modern Approach",
"authors": [
{
"given": "Stuart",
"family": "Russell",
"sequence": "first"
},
{
"given": "Peter",
"family": "Norvig",
"sequence": "additional"
}
],
"container_title": "",
"volume": "",
"issue": "",
"page": "1150",
"published_date": "2020-12-01",
"year": 2020,
"isbn": "9780262046255",
"doi": "",
"issn": [],
"publisher": "Pearson Education",
"language": "en",
"tags": ["artificial intelligence", "textbook", "reference"],
"abstract": "The most comprehensive, up-to-date introduction to the theory and practice of artificial intelligence."
},
"ref_003": {
"id": "ref_003",
"type": "journal_article",
"title": "Attention Is All You Need",
"authors": [
{
"given": "Ashish",
"family": "Vaswani",
"sequence": "first"
},
{
"given": "Noam",
"family": "Shazeer",
"sequence": "additional"
},
{
"given": "Niki",
"family": "Parmar",
"sequence": "additional"
},
{
"given": "Jakob",
"family": "Uszkoreit",
"sequence": "additional"
},
{
"given": "Llion",
"family": "Jones",
"sequence": "additional"
},
{
"given": "Aidan",
"family": "Gomez",
"sequence": "additional"
},
{
"given": "Łukasz",
"family": "Kaiser",
"sequence": "additional"
}
],
"container_title": "NeurIPS",
"volume": "30",
"issue": "",
"page": "5998-6024",
"published_date": "2017-12-06",
"year": 2017,
"doi": "10.5555/2020.FPTR.17450",
"issn": ["1049-5258"],
"publisher": "Curran Associates",
"language": "en",
"tags": ["attention mechanism", "transformer", "neural networks", "machine learning"],
"abstract": "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks."
},
"ref_004": {
"id": "ref_004",
"type": "conference_paper",
"title": "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding",
"authors": [
{
"given": "Jacob",
"family": "Devlin",
"sequence": "first"
},
{
"given": "Ming-Wei",
"family": "Chang",
"sequence": "additional"
},
{
"given": "Kenton",
"family": "Lee",
"sequence": "additional"
},
{
"given": "Kristina",
"family": "Toutanova",
"sequence": "additional"
}
],
"container_title": "Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics",
"volume": "",
"issue": "",
"page": "4171-4186",
"published_date": "2019-06-01",
"year": 2019,
"doi": "10.18653/v1/2019.naaci-1.1",
"issn": [],
"publisher": "Association for Computational Linguistics",
"language": "en",
"tags": ["BERT", "transformer", "NLP", "pre-training"],
"abstract": "We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers."
},
"ref_005": {
"id": "ref_005",
"type": "journal_article",
"title": "深度学习在自然语言处理中的应用研究",
"authors": [
{
"given": "三",
"family": "张",
"sequence": "first"
},
{
"given": "四",
"family": "李",
"sequence": "additional"
}
],
"container_title": "计算机学报",
"volume": "46",
"issue": "3",
"page": "1-15",
"published_date": "2023-03-01",
"year": 2023,
"doi": "10.3724/SP.J.2023.001234",
"issn": ["0254-8389"],
"publisher": "中国计算机学会",
"language": "zh",
"tags": ["深度学习", "自然语言处理", "NLP", "中文文献"],
"abstract": "本文研究了深度学习技术在自然语言处理领域的应用现状和发展趋势。"
},
"ref_006": {
"id": "ref_006",
"type": "journal_article",
"title": "基于卷积神经网络的文本分类方法研究",
"authors": [
{
"given": "伟",
"family": "王",
"sequence": "first"
},
{
"given": "明",
"family": "刘",
"sequence": "additional"
}
],
"container_title": "软件学报",
"volume": "42",
"issue": "6",
"page": "567-579",
"published_date": "2021-06-15",
"year": 2021,
"doi": "10.3724/SP.J.2021.04567",
"issn": ["0254-8389"],
"publisher": "中国计算机学会",
"language": "zh",
"tags": ["卷积神经网络", "文本分类", "CNN", "深度学习"],
"abstract": "提出了一种基于卷积神经网络的文本分类方法,在多个数据集上取得了良好的效果。"
},
"ref_007": {
"id": "ref_007",
"type": "thesis",
"title": "多模态深度学习技术研究",
"authors": [
{
"given": "丽",
"family": "陈",
"sequence": "first"
}
],
"container_title": "清华大学",
"volume": "",
"issue": "",
"page": "150",
"published_date": "2022-05-01",
"year": 2022,
"doi": "",
"issn": [],
"publisher": "清华大学",
"language": "zh",
"tags": ["多模态学习", "深度学习", "学位论文", "博士学位"],
"abstract": "本文研究了多模态深度学习技术,提出了新的融合方法和架构设计。"
},
"ref_008": {
"id": "ref_008",
"type": "book",
"title": "机器学习",
"authors": [
{
"given": "周",
"family": "志华",
"sequence": "first"
},
{
"given": "王",
"family": "亚",
"sequence": "additional"
}
],
"container_title": "",
"volume": "",
"issue": "",
"page": "285",
"published_date": "2016-01-01",
"year": 2016,
"isbn": "9787111390224",
"doi": "",
"issn": [],
"publisher": "清华大学出版社",
"language": "zh",
"tags": ["机器学习", "教材", "中文图书", "入门教材"],
"abstract": "本书系统地介绍了机器学习的基本概念、主要算法和应用案例。"
}
},
"citation_mappings": {
"ref_001": [
{
"document_id": "doc_001",
"document_title": "Deep Learning Survey Paper",
"positions": [45, 89, 234, 456],
"citation_type": "parenthetical",
"style": "apa"
},
{
"document_id": "doc_002",
"document_title": "Neural Networks Tutorial",
"positions": [12, 156, 289],
"citation_type": "numeric",
"style": "ieee"
}
],
"ref_002": [
{
"document_id": "doc_001",
"document_title": "AI Textbook Survey",
"positions": [78, 201, 345, 523],
"citation_type": "author-date",
"style": "harvard"
}
],
"ref_005": [
{
"document_id": "doc_zh_001",
"document_title": "中文文献综述",
"positions": [23, 67, 134],
"citation_type": "author-date",
"style": "gbt7714"
}
],
"ref_008": [
{
"document_id": "doc_zh_002",
"document_title": "中文教材",
"positions": [56, 123, 289, 412],
"citation_type": "author-date",
"style": "gbt7714"
}
]
},
"statistics": {
"total_references": 8,
"total_documents": 5,
"total_citations": 13,
"language_distribution": {
"zh": 3,
"en": 5
},
"type_distribution": {
"journal_article": 4,
"book": 2,
"conference_paper": 1,
"thesis": 1
},
"style_usage": {
"apa": 2,
"ieee": 1,
"harvard": 1,
"gbt7714": 2
},
"year_distribution": {
"2015": 1,
"2016": 1,
"2020": 1,
"2021": 1,
"2022": 1,
"2023": 1
}
},
"settings": {
"auto_save": true,
"auto_backup": true,
"backup_interval_hours": 24,
"max_cache_size": 10000,
"default_style": "apa",
"default_language": "zh",
"enable_crossref_cache": true,
"enable_doi_validation": true,
"format_on_export": true,
"include_abstract_in_bib": false,
"include_url_in_bib": true,
"abbreviate_journal_names": false,
"use_et_al_for_multiple_authors": true,
"max_authors_in_citation": 3,
"sort_bibliography_by": "alphabetical",
"use_ibid_for_consecutive": false,
"disambiguate_same_author_same_year": true
}
}Related skills
FAQ
Which citation formats are supported?
APA 7th, MLA 9th, Chicago 17th, GB/T 7714-2015, IEEE, and Harvard.
Where does it get metadata?
From authoritative databases like Crossref via DOI, ISBN, or title.