Skip to main content
Defenses
Work deposited June 15, 2010 (approx.)

Computational methodology for identification of nominal phrases in Portuguese

Brasília, June 15, 2010 (approximate date). Luana Vieira Morellato presented the Master's dissertation entitled Computational methodology for identification of nominal phrases in Portuguese, at Master's Thesis in Informatics, Federal University of Espirito Santo (Brazil), advised by Prof. Sergio Antônio Andrade de Freitas.

Abstract

Phrases, or syntagms, are units of meaning with syntactic functions within a sentence, as described by Nicola (2008). Generally, the sentences in any statement express content through the elements and combinations thereof provided by the language. This process forms sets and subsets that act as syntactic units within the larger unit of the sentence—syntagms, which can be categorized into nominal and verbal. Among these, nominal syntagms are of particular interest due to their higher semantic value. They are employed in Natural Language Processing (NLP) tasks such as anaphora resolution, automatic ontology building, parsing in medical texts for summary generation and vocabulary creation, or as an initial step in syntactic analysis processes. In Information Retrieval (IR), syntagms can be used to create terms in document indexing and search systems, improving results. This dissertation presents a computational methodology for identifying nominal syntagms in Portuguese language digital documents. It outlines the methodology for identifying and extracting nominal syntagms through the development of SISNOP—System for Identifying Nominal Syntagms in Portuguese. SISNOP, comprising a set of modules and programs, interprets unrestricted texts in natural language through morphological and syntactic analyses to retrieve nominal syntagms, also providing syntactic information such as gender, number, and degree of words within the extracted syntagms. Tested on corpora like CETENFolha and CETEMPúblico, SISNOP recognized 98.12% and 94.59% of sentences, identifying over 24 million syntagms. Its modules—Morphological Tagger, Nominal Syntagms Identifier, and Gender, Number, and Degree Identifier—were individually tested using a smaller dataset due to manual result analysis, achieving a precision of 82.45% and coverage of 69.20%.

Deposited work

The full text is available at the Digital Library / Institutional Repository.

Open publication →