Yayın: Veri Ambarlarında Verilerin Temizlenmesi
Yükleniyor...
Dosyalar
Tarih
Yazarlar
Danışman
Bölüm / Program
Bilgisayar Mühendisliği
Computer Engineering
Computer Engineering
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Fen Bilimleri Enstitüsü
Institute of Science and Technology
Institute of Science and Technology
Türü
Özet
Bu çalışmada veri ambarı sistemlerinde karşılaşılan çeşitli veri kalitesi problemleri sınıflandırılmış ve veri kalitesinin artırılması için kullanılan çeşitli teknikler üzerinde durulmuştur. Veri ambarlarında giderilmesi en güç olan ve üzerinde en fazla çalışılan konuların başında tekrarlı kayıtların tespit edilip ayıklanması gelmektedir. Diğer problemler ise nispeten daha basit bazı yöntemler kullanılarak giderilebilmektedir. Özellikle tekrarlı ve mükerrer kayıtların ortaya çıkma sebeplerinin başında yazım hataları gelmektedir. Tamamen aynı olan iki kayıt sisteme girilirken sezilip engellenebilir, fakat girilen kayıtlar arasında yazım farklılıkları olduğunda bunların veri girişi sırasında sezilmesi pek mümkün değildir. Bunun için bu çalışmada öncelikle yazım hatalarının belirlenebilmesi için sözlük kullanılmadan, Türkçe’ye özgü kurallardan yararlanılarak yazım hatalarının tespit edilmesi amaçlanmıştır. Bunun yanında n-gram metodu ile istatistiksel olarak da yazım denetiminin yapılması amaçlanmıştır. Daha sonraki adımda ise sistemde yer alan tekrarlı kayıtların belirlenip ayıklanması amacıyla sıralı komşu metodu olarak bilinen yöntemin bazı ek kurallar ile zenginleştirilerek uygulaması yapılmıştır.
Metin ÇINAR In this thesis, data quality problems in data warehouses are classified and the different methodologies applicable to the variety of data cleaning problems are presented. The main data quality problem in a data warehouse is duplicate records, so main consideration of this work is detection and elimination of the duplicate records. Other data quality problems can be solved rather easier methods. Incorrect or missing data values, inconsistent value naming conventions, and incomplete information cause “dirty” data files. Hence, it is not surprise to encounter multiple records referring to the same real world entity. Exactly same records can be detected easily during data entry but if there are slightly differencies, it is almost impossible to detect them at data entry step. Therefore, at first step it is aimed to detect spelling errors using Turkish grammer rules without using a Turkish dictionary. In addition to these grammer rules n-gram statistic techniques are used to increase succesfully detected misspelled words. For this goal, di-gram, tri-gram and four-gram statistic tables generated using some turkish corpus and Turkish spelling guide. At second step, it is aimed to detect and eliminate duplicate records in the system using enriched sorted neighbourhood method(SNM) with field weights.
Metin ÇINAR In this thesis, data quality problems in data warehouses are classified and the different methodologies applicable to the variety of data cleaning problems are presented. The main data quality problem in a data warehouse is duplicate records, so main consideration of this work is detection and elimination of the duplicate records. Other data quality problems can be solved rather easier methods. Incorrect or missing data values, inconsistent value naming conventions, and incomplete information cause “dirty” data files. Hence, it is not surprise to encounter multiple records referring to the same real world entity. Exactly same records can be detected easily during data entry but if there are slightly differencies, it is almost impossible to detect them at data entry step. Therefore, at first step it is aimed to detect spelling errors using Turkish grammer rules without using a Turkish dictionary. In addition to these grammer rules n-gram statistic techniques are used to increase succesfully detected misspelled words. For this goal, di-gram, tri-gram and four-gram statistic tables generated using some turkish corpus and Turkish spelling guide. At second step, it is aimed to detect and eliminate duplicate records in the system using enriched sorted neighbourhood method(SNM) with field weights.
Tanım
Tez (Yüksek Lisans) -- İstanbul Teknik Üniversitesi, Fen Bilimleri Enstitüsü, 2003
Thesis (M.Sc.) -- İstanbul Technical University, Institute of Science and Technology, 2003
Thesis (M.Sc.) -- İstanbul Technical University, Institute of Science and Technology, 2003
Dergi veya Seri
ISSN
ISBN
Haklar
İTÜ tezleri telif hakkı ile korunmaktadır. Bunlar, bu kaynak üzerinden herhangi bir amaçla görüntülenebilir, ancak yazılı izin alınmadan herhangi bir biçimde yeniden oluşturulması veya dağıtılması yasaklanmıştır.
İTÜ theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
İTÜ theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
Anahtar Kelimeler
Verilerin temizlenmesi, Veri tutarlılığı, Türkçe yazım denetimi, Tekrarlı kayıt tespiti, Katar benzerliği, Data cleaning, Data consistency, Turkish spell checking, Duplicate elimination, String similarity
Alıntı
Onay
Gözden geçir
Tamamlayıcı Bilgiler
Referans Gösteren
21
Görüntülenme
103
İndirme
Google Scholar
Scholar'da Ara ↗ Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.