A methodology for the use of natural language processing in the search for information in digital documents
This dissertation presents a novel digital text search methodology grounded in the Nominal Structure of Discourse, expanding on anaphora resolution techniques proposed by Freitas [2005]. It employs these techniques to reveal the textual structure designed by authors, offering a unique approach to Information Retrieval (IR). Traditional IR models, like the vector space model and Latent Semantic Indexing, primarily use document terms for representation and search, often failing to account for the nuances of natural language, such as the use of anaphoras. These linguistic features can diminish the representational efficacy of classical models by complicating the identification of key entities within texts. To address these limitations, an alternative structural model that incorporates anaphora resolution into its computational representation of documents is introduced, based on Seibel Júnior's work [Seibel Júnior and Freitas, 2007]. This model, which focuses on the Discourse Nominal Structure for Searches (DNSS), aims to improve upon existing IR methodologies by ensuring a more nuanced and effective representation of documents. It does so by identifying and utilizing the central elements, or focuses, of text sentences, while also considering additional information provided by the Nominal Structure of Discourse (NSD). The dissertation elaborates on the development of this search structure, detailing the anaphora resolution process and its impact on enhancing the representation and search result quality of documents. Through the exploration of algorithms and experimentation, the study underscores the potential of this new methodology in advancing the field of Information Retrieval.
Abstract
This dissertation presents a novel digital text search methodology grounded in the Nominal Structure of Discourse, expanding on anaphora resolution techniques proposed by Freitas [2005]. It employs these techniques to reveal the textual structure designed by authors, offering a unique approach to Information Retrieval (IR). Traditional IR models, like the vector space model and Latent Semantic Indexing, primarily use document terms for representation and search, often failing to account for the nuances of natural language, such as the use of anaphoras. These linguistic features can diminish the representational efficacy of classical models by complicating the identification of key entities within texts. To address these limitations, an alternative structural model that incorporates anaphora resolution into its computational representation of documents is introduced, based on Seibel Júnior’s work [Seibel Júnior and Freitas, 2007]. This model, which focuses on the Discourse Nominal Structure for Searches (DNSS), aims to improve upon existing IR methodologies by ensuring a more nuanced and effective representation of documents. It does so by identifying and utilizing the central elements, or focuses, of text sentences, while also considering additional information provided by the Nominal Structure of Discourse (NSD). The dissertation elaborates on the development of this search structure, detailing the anaphora resolution process and its impact on enhancing the representation and search result quality of documents. Through the exploration of algorithms and experimentation, the study underscores the potential of this new methodology in advancing the field of Information Retrieval.
BibTeX
@mastersthesis{2009-francisco-santiago-do-carmo-pereira-a-methodology-for-the-use-of-natural-language-pro,
author = {Francisco Santiago do Carmo Pereira},
title = {A methodology for the use of natural language processing in the search for information in digital documents},
year = {2009},
publisher = {Biblioteca Central da Universidade Federal do Espirito Santo}
}