Yayın: Türkçe İçin Metin Özetleme
Yükleniyor...
Dosyalar
Tarih
Yazarlar
Danışman
Akademik Birim
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Fen Bilimleri Enstitüsü
Institute of Science and Technology
Institute of Science and Technology
Türü
Özet
Bilgi erişimi genel şekliyle, depolanmış bilgi derleminden belirli bilgi gereksinimiyle ilgili bölümlere erişim yöntemine yönelik çalışma olarak tanımlanabilir. Bilgi erişiminin altkümelerinden biri olan metin özetleme, bir belgeyi girdi olarak alan ve çıktı olarak daha kısa, aslının yerine geçen ve onun en önemli içeriğini barındıran bir süreç olarak tanımlanabilir. Yüksek verimlilik, yüksek başarım ve düşük uygulama maliyeti bugünkü araştırmalarda ve pratik uygulamalarda genellikle istatistiksel yöntemlerin kullanılmasının sebebidir. Türkçe, sondan eklemeli ve kurallı yapısı, çok az miktarda kuralsız sözcük içermesi nedeniyle bilgi erişimi araştırmacılarının ilgisini çekmiştir. Türkçenin bu özellikleri, Türkçe için yapılan tüm bilgi erişimi sistemlerinde gövdeleme işlemine önem kazandırmıştır. Bu tezde, Türkçe için farklı metin özetleme yöntemleri tanıtılıp uygulanmıştır. Diğer tüm Türkçe bilgi erişimi sistemlerinde de gerekli olduğu gibi, Türkçenin sondan eklemeli yapısının gözetilmesi amacıyla farklı gövdeleme algoritmalarının özetleme başarımına etkisi incelenmiştir. Başarımlarının daha yüksek olması amacıyla, gerçeklenen gövdeleme algoritmalarında sözcüklerin olası kök ve ek birleşimlerini üreten biçimbirimsel çözümleyici kullanılmıştır. Gövdelenmiş bu sözcükler farklı özetleme yöntemleri aracılığıyla incelenip her yöntem için özette yer alacak cümleler belirlenmiştir. Daha sonra bu yöntemlerin ürettiği sonuçlar birleştirilerek son özet oluşturulmuştur.
Information retrieval can be broadly defined as the study of how to determine and retrieve the portions, which are relevant to particular information needs, from a corpus of stored information. One of the subsets of information retrieval is text summarization. Text summarization can be defined as the process which takes a document as input and outputs a shorter document which is condensed and can be used instead of the original. Today’s researches and practical applications about text summarization mostly use the early statistical methods because of high efficiency, high performance and low application cost of these approaches. In this study, different statistical methods for text summarization are described and developed for Turkish. The effect of different stemming algorithms on summarization efficiency has been studied for the aim of taking into consideration the agglutinative structure of Turkish, as it is necessary in all other information retrieval systems for this language. Morphological analyzer, which outputs the root and affix combinations of the input word, has been used in stemming algorithms to increase the efficiency of the text summarization. These stemmed words have been studied by different summarization methods and sentences which will be included in the summary have been chosen. In the end, the final summary has been created by combining the results of these methods.
Information retrieval can be broadly defined as the study of how to determine and retrieve the portions, which are relevant to particular information needs, from a corpus of stored information. One of the subsets of information retrieval is text summarization. Text summarization can be defined as the process which takes a document as input and outputs a shorter document which is condensed and can be used instead of the original. Today’s researches and practical applications about text summarization mostly use the early statistical methods because of high efficiency, high performance and low application cost of these approaches. In this study, different statistical methods for text summarization are described and developed for Turkish. The effect of different stemming algorithms on summarization efficiency has been studied for the aim of taking into consideration the agglutinative structure of Turkish, as it is necessary in all other information retrieval systems for this language. Morphological analyzer, which outputs the root and affix combinations of the input word, has been used in stemming algorithms to increase the efficiency of the text summarization. These stemmed words have been studied by different summarization methods and sentences which will be included in the summary have been chosen. In the end, the final summary has been created by combining the results of these methods.
Tanım
Tez (Yüksek Lisans) -- İstanbul Teknik Üniversitesi, Fen Bilimleri Enstitüsü, 2007
Thesis (M.Sc.) -- İstanbul Technical University, Institute of Science and Technology, 2007
Thesis (M.Sc.) -- İstanbul Technical University, Institute of Science and Technology, 2007
Dergi veya Seri
ISSN
ISBN
Haklar
İTÜ tezleri telif hakkı ile korunmaktadır. Bunlar, bu kaynak üzerinden herhangi bir amaçla görüntülenebilir, ancak yazılı izin alınmadan herhangi bir biçimde yeniden oluşturulması veya dağıtılması yasaklanmıştır.
İTÜ theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
İTÜ theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
Anahtar Kelimeler
Bilgi Erişimi, Metin Özetleme, Gövdeleme, Information Retrieval, Text Summarization, Stemming
Alıntı
Onay
Gözden geçir
Tamamlayıcı Bilgiler
Referans Gösteren
118
Görüntülenme
580
İndirme
Google Scholar
Scholar'da Ara ↗ Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.