Enhanced Graph Based Approach for Multi Document Summarization
Shanmugasundaram Hariharan1, Thirunavukarasu Ramkumar2, and Rengaramanujam Srinivasan3
1Department of Computer Science and Engineering, TRP Engineering College, India
2Department of Computer Application, A.V.C College of Engineering, India
3School of Computer Science, Bangladeshi Students Association University, India
1Department of Computer Science and Engineering, TRP Engineering College, India
2Department of Computer Application, A.V.C College of Engineering, India
3School of Computer Science, Bangladeshi Students Association University, India
Abstract: Summarizing documents catering the needs of an user is tricky and challenging. Though there are varieties of approaches, graphical methods have been quite popularly investigated for summarizing document contents. This paper focus its attention on two graphical methods namely-LexRank (threshold) and LexRank (Continuous) proposed by Erkan and Radev. This paper proposes two enhancements to the above work investigated earlier by adding two more features to the existing one. Firstly, discounting approach was introduced to form a summary which ensures less redundancy among sentences. Secondly, position weight mechanism has been adopted to preserve importance based on the position they occupy. Intrinsic evaluation has been done with two data sets. Data set 1 has been created manually from the news paper documents collected by us for experiments. Data set 2 is from DUC 2002 data which is commercially available and distributed or accessed through National Institute of Standards Technology (NIST). We have shown that the based upon precision and recall parameters were comprehensively better as compared to the earlier algorithms.
Keywords: Page rank, lexical rank, damping, threshold, summarization.
Keywords: Page rank, lexical rank, damping, threshold, summarization.
Received July 13, 2011; accepted December 29, 2012; published online August 5, 2012