Knowledge base construction from scientific literature

Wang, Han
Thumbnail Image
Other Contributors
Fox, Peter A.
Hendler, James A.
Ji, Heng
Stephan, Eric
Lewis, Daniel
Issue Date
Multidisciplinary science
Terms of Use
Attribution-NonCommercial-NoDerivs 3.0 United States
This electronic version is a licensed copy owned by Rensselaer Polytechnic Institute, Troy, NY. Copyright of original work retained by author.
Full Citation
Knowledge Bases (KBs) have become a functional utility as a repository of information for both humans and software agents to seek confirmed facts about the world. With the wide-ranging application of KBs, automatically constructing either generic KBs or domain-specific KBs using information extracted from multiple sources such as web pages, reports, and research papers has grown into an interesting task for both academia and industry.
SciKB adopts an open information extraction approach to extract fact triples from the input documents, then jointly learns the distributed representations of the involved entities and relations in an unsupervised fashion, and finally utilizes the obtained representations to organize the entities and relations into hierarchical clusters. Experiments are conducted to evaluate each component of the SciKB pipeline and the results demonstrate its effectiveness in two scientific domains: Biomedical Science and Earth Science.
This dissertation presents SciKB, an end-to-end Knowledge Base Construction system, which takes in a collection of research articles within a certain scientific domain and outputs a domain-specific KB. The resultant KB contains fact triples extracted from the input documents as well as hierarchical clusters of the entities and relations involved in the facts. Each cluster aggregates entities or relations with similar semantic meanings, and the hierarchies serve as an implicit schema of the KB.
December 2016
School of Science
Multidisciplinary Science Program
Rensselaer Polytechnic Institute, Troy, NY
Rensselaer Theses and Dissertations Online Collection
CC BY-NC-ND. Users may download and share copies with attribution in accordance with a Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 License. No commercial use or derivatives are permitted without the explicit approval of the author.