图书馆杂志

图书馆杂志 ›› 2022, Vol. 41 ›› Issue (12): 96-103.

• 数字人文 • 上一篇    下一篇

中国古代书目提要结构功能识别研究——以《四库全书总目》著录的古代科技文献为例

郑 翔1, 2 李明杰1, 3 (1 武汉大学信息管理学院 2 武汉大学图书情报国家级实验教学示范中心 3 武汉大学文化遗产智能计算实验室)   

  • 出版日期:2022-12-15 发布日期:2023-01-03
  • 作者简介:郑 翔 女,武汉大学信息管理学院,博士研究生。研究方向:古籍整理与数字人文。作者贡献:进行实验,起草论文。E-mail:zhengxiang059@whu.edu.cn 湖北武汉 430072 李明杰 武汉大学信息管理学院,教授。研究方向:文献整理与保护、中国图书文化史。作者贡献:提出选题,修订论文。湖北武汉 430072

Research on Structure Function Recognition of Ancient Bibliographic Synopses: A Case Study of Ancient Scientific and Technical Books in the Siku Quanshu Zongmu

Zheng Xiang1, 2, Li Mingjie1, 3 (1 School of Information Management, Wuhan University; 2 National Demonstration Center for Experimental Library and Information Science Education, Wuhan University; 3 Intellectual Computing Laboratory for Cultural Heritage, Wuhan University)   

  • Online:2022-12-15 Published:2023-01-03
  • About author:Zheng Xiang1, 2, Li Mingjie1, 3 (1 School of Information Management, Wuhan University; 2 National Demonstration Center for Experimental Library and Information Science Education, Wuhan University; 3 Intellectual Computing Laboratory for Cultural Heritage, Wuhan University)

摘要: 从古代书目提要中准确识别提要结构功能模块,对查询古籍、研究学问、鉴别版本、考证佚亡有重要作用。针对当前古代书目提要结构功能内涵把握的现实需求,从提要结构功能智能识别角度出发,提出细粒度六分类提要结构功能划分策略,并使用RoBERTa模型构建提要结构功能识别模型。面向《四库全书总目》中古代科技文献提要的实验表明:本文所提划分策略与识别模型具有一定的领域适应性与有效性;通过结合预训练模型、自动学习古代书目提要语义内涵特征,能够有效进行提要结构功能识别。本文有效降低了提要利用门槛、提升了提要内容呈现效率,有利于读者对古籍文献的精准检索、快速了解与深度利用。

Abstract: Accurate identification of the structure function modules from the ancient bibliographic synopses is important for searching ancient books, studying knowledge, identifying editions and examining the loss of ancient books. This paper proposes a fine-grained six-category synopses structure function classification strategy and construct a synopses structure function recognition model using the RoBERTa to meet the current practical needs of grasping the structure function connotation of ancient bibliographic synopses, from the perspective of intelligent synopses structure function recognition. The experiments on the synopses of ancient scientific and technical books in the Siku Quanshu Zongmu show that the proposed classification strategy and recognition model has certain domain adaptability and effectiveness. By combining pre-training models and automatic learning of semantic connotation features of ancient bibliographic synopses, we can effectively perform synopses structure function recognition. It effectively reduces the threshold of synopses utilization, improves the speed of synopses content presentation, and facilitates readers’ accurate search, rapid understanding and in-depth utilization of ancient books.