DSpace at KOASAS: An efficient pre-processing method to identify logical components from PDF documents

DSpace at KOASAS

College of Engineering(공과대학)Dept. of Nuclear and Quantum Engineering(원자력및양자공학과)NE-Conference Papers(학술회의논문)

An efficient pre-processing method to identify logical components from PDF documents

Cited 7 time in

Cited 0 time in

Hit : 558
Download : 0

Export

Liu, Ying researcher / Bai, K. / Gao, L.

As the rapid growth of the scientific documents in digital libraries, the search demands for the documents as well as specific components increase dramatically. Accurately detecting the component boundary is of vital importance to all the further information extraction and applications. However, document component boundary detection (especially the table, figure, and equation) is a challenging problem because there is no standardized formats and layouts across diverse documents. This paper presents an efficient document preprocessing technique to improve the document component boundary detection performance by taking advantage of the nature of document lines. Our method easily simplifies the component boundary detection problem into the sparse line analysis problem with much less noise. We define eight document line label types and apply machine learning techniques as well as the heuristic rule-based method on identifying multiple document components. Combining with different heuristic rules, we extract the multiple components in a batch way by filtering out massive noises as early as possible. Our method focus on an important un-tagged document format - PDF documents. The experimental results prove the effectiveness of the sparse line analysis.

Publisher: PAKDD'11

Issue Date: 2011-05-24

Language: English

Citation: 15th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2011, pp.500 - 511

ISSN: 0302-9743

URI: http://hdl.handle.net/10203/164991

Appears in Collection: KSE-Conference Papers(학술회의논문)

Files in This Item: There are no files associated with this item.

This item is cited by other documents in WoS

⊙ Detail Information in WoSⓡ	Click to see
⊙ Cited 7 items in WoS	Click to see citing articles in

Display Full Item Record

qr_code

트윗하기

KOASAS

Knowledge Service Development Team, KAIST 291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea. T. 82-42-350-4493 Email. koasas@kaist.ac.kr
Copyright © 2016. Korea Advanced Institute of Science and Technology. All Rights Reserved.

KOASAS

KOASAS

Browse

An efficient pre-processing method to identify logical components from PDF documents

This item is cited by other documents in WoS

KOASAS

Communities & Collections