Skip to main content
Publications
2010 dissertation

Text categorization methodology from unlabeled documents using an anaphora resolution process

With the ongoing expansion of electronic textual content, there's a need to organize all this information in a manageable way. Hence, the process of text categorization was developed to ease the manipulation and retrieval of information by dividing it into thematic categories. There are various approaches to achieving an automatic text categorizer, among which the supervised paradigm is the most traditional. Although supervised methodology shows precision comparable to human experts, the requirement for a pre-classified corpus can be a limiting factor in some applications. In these cases, a semi or unsupervised solution, which doesn't demand a complete and well-formed training set for categorizer construction, can be applied; instead, unlabeled documents are provided to the method. Both supervised learning paradigms and semi or unsupervised paradigms typically build a representation of texts based solely on term occurrence, not considering semantic factors. However, many intrinsic characteristics of natural language can make the process ambiguous, one of which is the use of various terms to refer to an entity already mentioned in the text, a linguistic phenomenon known as anaphora. This dissertation proposes a method for the conception of an unsupervised categorizer, utilizing the Nominal Structure of Discourse (NSD), developed by Freitas for anaphora resolution, as a foundation. For this purpose, a bootstrapping technique for categorization is implemented, aiming to obtain initial labeling for documents, which is used to generate a categorization model through the supervised paradigm. Besides being based on the NSD, the methodology of this work benefits directly from the anaphora resolution process, using the identified antecedents for anaphoras during the final categorization phase. This work presents details about the proposed methodology, explaining the developed algorithms, as well as the experiments conducted for method evaluation. Results show that the use of the anaphora resolution process is beneficial for an unsupervised categorization system.

Abstract

With the ongoing expansion of electronic textual content, there’s a need to organize all this information in a manageable way. Hence, the process of text categorization was developed to ease the manipulation and retrieval of information by dividing it into thematic categories. There are various approaches to achieving an automatic text categorizer, among which the supervised paradigm is the most traditional. Although supervised methodology shows precision comparable to human experts, the requirement for a pre-classified corpus can be a limiting factor in some applications. In these cases, a semi or unsupervised solution, which doesn’t demand a complete and well-formed training set for categorizer construction, can be applied; instead, unlabeled documents are provided to the method. Both supervised learning paradigms and semi or unsupervised paradigms typically build a representation of texts based solely on term occurrence, not considering semantic factors. However, many intrinsic characteristics of natural language can make the process ambiguous, one of which is the use of various terms to refer to an entity already mentioned in the text, a linguistic phenomenon known as anaphora. This dissertation proposes a method for the conception of an unsupervised categorizer, utilizing the Nominal Structure of Discourse (NSD), developed by Freitas for anaphora resolution, as a foundation. For this purpose, a bootstrapping technique for categorization is implemented, aiming to obtain initial labeling for documents, which is used to generate a categorization model through the supervised paradigm. Besides being based on the NSD, the methodology of this work benefits directly from the anaphora resolution process, using the identified antecedents for anaphoras during the final categorization phase. This work presents details about the proposed methodology, explaining the developed algorithms, as well as the experiments conducted for method evaluation. Results show that the use of the anaphora resolution process is beneficial for an unsupervised categorization system.

BibTeX

@mastersthesis{2010-debora-zupeli-bossois-metodologia-de-categorizacao-de-textos-a-partir-de-documentos-n,
  author = {Débora Zupeli Bossois},
  title = {Text categorization methodology from unlabeled documents using an anaphora resolution process},
  year = {2010},
  publisher = {Biblioteca Central da Universidade Federal do Espirito Santo}
}