Skip to main content
Defenses
Work deposited June 15, 2009 (approx.)

A methodology for the use of natural language processing in the search for information in digital documents

Brasília, June 15, 2009 (approximate date). Francisco Santiago do Carmo Pereira presented the Master's dissertation entitled A methodology for the use of natural language processing in the search for information in digital documents, at Master's Thesis in Informatics, Federal University of Espirito Santo (Brazil), advised by Prof. Sergio Antônio Andrade de Freitas.

Abstract

This dissertation presents a novel digital text search methodology grounded in the Nominal Structure of Discourse, expanding on anaphora resolution techniques proposed by Freitas [2005]. It employs these techniques to reveal the textual structure designed by authors, offering a unique approach to Information Retrieval (IR). Traditional IR models, like the vector space model and Latent Semantic Indexing, primarily use document terms for representation and search, often failing to account for the nuances of natural language, such as the use of anaphoras. These linguistic features can diminish the representational efficacy of classical models by complicating the identification of key entities within texts. To address these limitations, an alternative structural model that incorporates anaphora resolution into its computational representation of documents is introduced, based on Seibel Júnior's work [Seibel Júnior and Freitas, 2007]. This model, which focuses on the Discourse Nominal Structure for Searches (DNSS), aims to improve upon existing IR methodologies by ensuring a more nuanced and effective representation of documents. It does so by identifying and utilizing the central elements, or focuses, of text sentences, while also considering additional information provided by the Nominal Structure of Discourse (NSD). The dissertation elaborates on the development of this search structure, detailing the anaphora resolution process and its impact on enhancing the representation and search result quality of documents. Through the exploration of algorithms and experimentation, the study underscores the potential of this new methodology in advancing the field of Information Retrieval.

Deposited work

The full text is available at the Digital Library / Institutional Repository.

Open publication →