Back to Search
Start Over
An Information Distillation Framework for Extractive Summarization
- Source :
- IEEE/ACM Transactions on Audio, Speech, and Language Processing. 26:161-170
- Publication Year :
- 2018
- Publisher :
- Institute of Electrical and Electronics Engineers (IEEE), 2018.
-
Abstract
- In the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some realistic tasks such as document summarization. Nevertheless, classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions in this paper are threefold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph of interest. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. Third, a new summarization framework, which can take both relevance and redundancy information into account simultaneously, is also introduced. We evaluate the proposed embedding methods (i.e., EV and D-EV) and the summarization framework on two benchmark summarization corpora. The experimental results demonstrate the effectiveness and applicability of the proposed framework in relation to several well-practiced and state-of-the-art summarization methods.
- Subjects :
- Context model
Acoustics and Ultrasonics
Computer science
business.industry
020206 networking & telecommunications
Context (language use)
02 engineering and technology
computer.software_genre
Automatic summarization
030507 speech-language pathology & audiology
03 medical and health sciences
Computational Mathematics
0202 electrical engineering, electronic engineering, information engineering
Computer Science (miscellaneous)
Relevance (information retrieval)
Artificial intelligence
Electrical and Electronic Engineering
Paragraph
0305 other medical science
Representation (mathematics)
business
computer
Feature learning
Natural language processing
Sentence
Subjects
Details
- ISSN :
- 23299304 and 23299290
- Volume :
- 26
- Database :
- OpenAIRE
- Journal :
- IEEE/ACM Transactions on Audio, Speech, and Language Processing
- Accession number :
- edsair.doi...........a0f35b3b728156bd5fec6bcc63efc7bb