Publication:
Improving self-attention based transformer performance for morphologically rich languages

Loading...
Thumbnail Image

Institution Authors

Department

Computer Engineering

Journal Title

Journal ISSN

Volume Title

Publisher

Graduate School

Research Projects

Organizational Units

Journal Issue

Abstract

This dissertation examines the transformative impact of the Bidirectional Encoder Representations from Transformers (BERT) algorithm on Natural Language Processing (NLP), with a particular focus on its application to Turkish, a morphologically rich language. The advent of BERT marked a significant shift from traditional recurrent neural networks to transformer-based models. This shift introduced self-attention mechanisms and the use of rich unsupervised data. The capacity to condition on both left and right contexts simultaneously enables a more nuanced and comprehensive understanding of linguistic structures. The versatility of BERT is evident in its application to a range of NLP tasks, including complex question answering and nuanced language inference. Its real-world impact is demonstrated by its implementation in Google's search engine, which has shifted from keyword-centric to semantic search paradigms. The research investigates the unique challenges presented by Turkish, a language with complex morphological structures that differ significantly from more commonly studied languages. Traditional tokenization methods and vocabulary sizes often prove inadequate for processing Turkish effectively. The dissertation traces the evolution of Turkish NLP from rule-based approaches to advanced neural models, critically assessing the progress and limitations in the field. This assessment highlights the need for nuanced approaches to tokenization and vocabulary size in Turkish NLP tasks. A key objective of the research is to ascertain the optimal vocabulary size for Turkish Named Entity Recognition (NER). The morphological richness of Turkish presents unique challenges, particularly the prevalence of out-of-vocabulary (OOV) words. The study examines the impact of tokenization granularity, influenced by vocabulary size, on the performance of BERT-based models on Turkish NER tasks. To address these issues, we developed Turkish-specific BERT language models, each trained and tuned with different vocabulary sizes and normalization settings. The results indicate that larger vocabulary sizes typically improve NER performance for languages like Turkish, thanks to their more comprehensive and nuanced representation of words as tokens. The dissertation discusses the factors influencing the performance of LLM, including the corpus used for vocabulary training, preprocessing operations, tokenization methods, and vocabulary size. It introduces a novel metric, the tokenization granularity rate, to quantify the level of granularity and explore its impact on various NLP tasks. These findings underscore the significance of meticulous vocabulary size selection for Turkish LLMs. Larger sizes tend to result in enhanced performance, largely due to reduced tokenization granularity and more efficacious capture of the language's morphological intricacies. A notable contribution of the research is the introduction of the BERT2D model, which employs two-dimensional positional embeddings to more effectively capture the morphological intricacies of agglutinative languages like Turkish. This innovative model addresses the limitations of linear, one-dimensional positional embeddings in standard BERT, particularly for languages that necessitate sophisticated positional encoding due to high tokenization granularity. The efficacy of BERT2D is validated through a series of rigorous experiments, including pretraining, fine-tuning, and evaluation. These experiments employ dual positional embeddings for whole words and subwords. The results demonstrate that BERT2D, particularly when combined with whole-word masking, consistently outperforms benchmark models in various NLP tasks, establishing new standards in Turkish NLP applications. The research establishes the superior performance of the Turkish-specific ITU-TurkBERT model over the multilingual BERT in various Turkish downstream tasks. This achievement highlights the model's enhanced ability to capture language-specific phenomena and reinforces the critical role of vocabulary size in dealing with the morphological complexity of Turkish. The study finds that optimal vocabulary sizes for different tasks often exceed 64K, and examines the mixed effectiveness of strategies such as normalization, morphological tokenization, and corpus size reduction compared to the base language model. The BERT2D models presented in this research demonstrate consistent performance, exceeding that of standard BERT-based models, in tasks such as token classification and text classification. This superior performance is achieved with minimal additions to the model parameters. The incorporation of whole-word masking has been demonstrated to be especially advantageous for text classification tasks. These findings contribute to the advancement of Turkish processing and offer insights that may be applicable to other morphologically rich languages. The dissertation concludes with an outline of several potential avenues for future research. Further research should be conducted to extend the tests to other named entity recognition datasets, enrich the attention calculation for named entity recognition, explore other downstream tasks with cased and newer versions of BERT, and experiment with different types of layers in BERT's output structure. Furthermore, the application of the proposed hierarchical recommendation network to cross-domain recommendation problems represents an important area for future exploration and improvement. Overall, this research significantly advances the field of Turkish NLP by addressing the unique challenges posed by the language's rich morphological structure, paving the way for more effective and nuanced NLP applications across diverse linguistic landscapes.

Description

Thesis (Ph.D.) -- Istanbul Technical University, Graduate School, 2024

Journal or Series

ISSN

ISBN

Rights

Keywords

Natural language processing, Doğal dil işleme

Citation

Endorsement

Review

Supplemented By

Referenced By

Related Patent

Related Goal

29
Görüntülenme
42
İndirme
Google Scholar
Scholar'da Ara ↗
Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.