Geri Dön

İTÜ NER - Türkçe metinlerde adlandırılmış varlık tespiti

ITU NER - named entity recognition on Turkish texts

  1. Tez No: 798113
  2. Yazar: GÖKHAN AKIN ŞEKER
  3. Danışmanlar: DOÇ. DR. GÜLŞEN ERYİĞİT
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Bilgi çıkarımı, Doğal dil işleme, Varlık isimleri, Yapay zeka, Information extraction, Natural language processing, Named entities, Artificial intelligence
  7. Yıl: 2015
  8. Dil: Türkçe
  9. Üniversite: İstanbul Teknik Üniversitesi
  10. Enstitü: Bilişim Enstitüsü
  11. Ana Bilim Dalı: Bilgisayar Bilimleri Ana Bilim Dalı
  12. Bilim Dalı: Bilgisayar Bilimleri Bilim Dalı
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Adlandırılmış Varlık Tespiti (NER – Named Entity Recognition) en basit şekilde; metin içinden ilgilenilen varlık türlerine ait sözcük adlarının belirlenip bunlara doğru sınıf etiketlerinin atanması olarak tanımlanabilir. Literatürde üzerinde en çok çalışılan türler MUC-6 konferansındaki ortak görevle tanımlanan ENAMEX (kişi, yer, kurum adları) tipleridir. Aynı görevde tanımlanan diğer tipler olan TIMEX (tarih ve saat ifadeleri) ve NUMEX (yüzde ve parasal ifadeler) tipleri de diğer yaygın çalışılan sınıflar olarak karşımıza çıkmakla birlikte aranacak varlık türleri için herhangi bir sınırlama yoktur, protein adları, gen adları, ilaç adları gibi çok çok farklı alanlarda çalışmalara da rastlanabilmektedir. Bu çalışma temel olarak üç aşamada yürütülmüştür. Birinci aşamada ENAMEX türleri üzerinde resmi dille yazılmış metinlerde çalışan bir sistem ortaya konmuş mevcut Türkçe NER sistemleri arasında en yüksek başarım raporlanmıştır; ikinci aşamada bu sisteme TIMEX ve NUMEX türleri eklenerek üzerinde çalışılan tür sayısı yediye çıkarılmıştır; üçüncü aşamada ise bu sistem günlük konuşma diline yakın olan Web 2.0 metinlerine uyarlanmıştır. Birinci aşamada literatürde mevcut çalışmalar incelenirken neredeyse hemen hemen tüm çalışmaların farklı veri kümeleri üzerinde test edildiği veya değerlendirmede farklı kıstaslar esas alındığı için karşılaştırılabilir olmadığı tespit edilmiş ve geçmiş çalışmalar için değerli sonuçlar ortaya koyduğu düşünülen bir çalışma ile konu üzerindeki önemli geçmiş yayınların detaylı değerlendirmesi yapılmıştır. Çalışma sonucunda ortaya konan model, makine öğrenmesi metodu olarak Şartlı Rastgele Alanlar (CRFs) kullanırken diğer yanda titiz bir çalışma ile derlenen alan atlaslarından (gazetteer) da faydalanıldığı için hibrit bir model olarak nitelenebilir. Bu aşamanın sonunda Türkçe gazete haber metinlerinde MUC kıstaslarıyla %95, CoNLL kıstaslarıyla %92 F-ölçütü başarımı ile literatürdeki en yüksek başarım raporlanmıştır. İkinci aşamada birinci aşamanın çıktısı olan modele TIMEX ve NUMEX türlerini de tespit edebilme yeteneği eklenmiştir. Bu aşamada yapılan temel iş birinci aşamada kullanılan verinin yedi tür için yeniden işaretlenmesi ve yeni eklenen türlerin tanınmasında başarımı artırmak için ilave alan atlasları ve CRFs özellikleri eklenmesidir. Sonuçta yedi tür için de benzer oranda yüksek başarım elde edilmiştir. Üçüncü aşamada resmi dille yazılmış metinlerde çalışan model, serbest biçimli dile uyarlanarak, Web 2.0 verisinde çalışmalar yapılmıştır. Bu aşamada iki ayrı sosyal medya veri kümesi işaretlenmiş ve kuralsız metinlerin kurallı metinlere benzetimini sağlamaya yönelik düzeltme adımları eklenmiştir. Twitter veri kümesi üzerinde %68 ile literatürdeki en yüksek başarım oranlarına ulaşılmıştır. Araç diğer güncel bilimsel çalışmalarda kullanılan veri kümeleri üzerinde de test edilerek sonuçlar karşılaştırmalı olarak verilmiştir. Bu çalışma ile hazırlanan üç adet işaretli veri kümesi ve geniş alan atlasları (kişi ad, kişi soyad, yer adları gibi) bu alanda yapılacak sonraki çalışmalarda faydalanılabilecek önemli kaynaklar olarak araştırmacıların hizmetine açıktır. Modelin kendisi de İTÜ Doğal Dil İşleme Araçları arasında çevrimiçi kullanıma açılmıştır.

Özet (Çeviri)

Named Entity Recognition(NER) is a crucial stage in many Natural Language Processing (NLP) tasks including information retrieval, machine translation and opinion mining. The task aims to identify and classify certain types of entities such as names (e.g. person, location, organization, protein, genes), numerical (e.g. percent, monetary values) and temporal expressions (e.g. date, time) in text. The NER research was firstly started in early 1990s for English. In 1995, with the high interest of the research community, the success rates for English achieved nearly the human annotation performance on news texts. MUC and CoNLL conferences define three basic types of named entities which are: 1- ENAMEX (person, location and organization names), 2- TIMEX (temporal expressions: date and time entities) and 3- NUMEX (numerical expressions: monetary expressions and percentages). Although these became almost a de facto standard to evaluate the systems' performances, NER is not limited to only these types and it is also applied to different application areas in the literature such as determining protein names, medicine names, book titles etc... This study reports the highest results(92% on formal news texts dataset, 68% in Twitter dataset and 65% in balanced Web 2.0 data set in CoNLL metrics) in the literature for Turkish named entity recognition; more spesifically for the task of detecting ENAMEX, TIMEX and NUMEX types. An in depth analysis of the previous reported results are given and comparisons with them are made whenever possible. Used statistical model is conditional random fields (CRFs). Presented model is a hybrid model which depends on the usage of rich morphological structure of the Turkish language as features to CRFs together with the use of some basic and generative gazetteers. In this study CRF++ an open source implementation of CRFs is used. This study was organized in three phases. In the first phase a state-of-the-art NER system for ENAMEX types in formal written Turkish texts have been revealed; at the second phase the system was extended to 7 entity types adding the NUMEX and TIMEX types; in the third phase system is adapted to informal Web 2.0 types. In the first phase a Turkish NER model using conditional random fields(CRFs) trained with morphological and lexical features had been presented. This model only classifies the ENAMEX types and reports F-Measures of 95% in MUC metrics, and 92% in CoNLL metrics. Also large scale person (First names gazetteer of 44,048 tokens and Surnames gazetteer of 138.844 tokens), and location names (33.551 tokens) gazetteers and relatively small location, organization and person name generator gazetteers (<100 tokens) had been compiled and made available for the future work for the researchers, in this stage. Revealed model consists of tokenization, morphological analysis and gazetteer lookup steps. Another outcome of this phase was a short survey which tried to compare the results and the evaluation metrics of recent NER work in Turkish. This is an important contribution since the results given in previous works were not comparable because of first, different evaluation metrics giving different credits to partial matches were used and second, the studies focused to different sets of named entity types and provided their results as the average of these. In the second phase, model revealed at the first phase was extended to 7 entity types including also the NUMEX and TIMEX types as addition to existing ENAMEX tags. At this stage Turkish newspaper corpus used in the first page has been reannotated for 7 entity types. Final dataset consists of 500K words 15,352 person names, 10,404 location names, 9,571 organization names, 1,486 date entities, 169 time entities, 638 monetary expressions and 710 percentage expressions. One tenth of the data (47,344 words) is reserved for testing and remaining is used the for training purposes. At this stage, two extra generator gazetteers, namely currency names and month names gazetteers, are added. This model also reports %92 F-Measure in CoNLL metrics. In the third phase, the model is adapted to informal texts on Web 2.0 domain. This stage included annotation of two datasets. First one named ITU Web Treebank is a balanced dataset collected from different web 2.0 datasources consisting of 43K tokens and the second one is a Twitter dataset with 50K tokens. Both datasets are annotated for ENAMEX, TIMEX, NUMEX data types in accordance with MUC-6 guidelines. Adaptation of the system to informal texts is composed of two kind of operations: 1- Normalization of the data a. An extra gazetteer is added for authomatic capitalization of some names (some city-country names and some frequent oraganization names) which are always proper names. Note that this gazetteer doesn't include city names like Ordu (it has a meaning of army in Turkish) or Aydın (it has meanings like intellectual and bright in Turkish). This gazetteer is different from other gazetteers (both base and generator gazetters) because it is not related to CRFs features, it is only used to correct capitalization errors. b. A one character deascification is used for gazetteer lookup features. Thus all“i”,“u”,“o”,“g”,“c”,“s”characters in the token replaced to“ı”,“ü”,“ö”,“ğ”,“ç”and“ş”characters respectively one at each and looked up in the gazetteer. Please not that this deascification is not always mean a minimum edit distance value of one. It will match Kutahya to Kütahya and Misir to Mısır but wont match Kadikoy to Kadıköy. This kind of deascification raised up the performance about 15%, while other kinds of deascification such as multicharacter replacement or an edit distance based lookup performed worse. 2- Adding an extra feature to CRF model for catching Twitter mention names. This feature only includes a format check (starts with a“@”character and continues with letters, digits and underscore character until next White space character), there is no online check if the token is a real mention name or not. Introduced system obtains the highest results in the literature for Turkish NER and also tested on some other datasets from other recent studies, and achieved better results for all. Annotated (gold standard) data is valuable for all NLP tasks. Gazetteers, or entity dictionaries, are important for improving the performance of named entity recognition. However, building and maintaining high-quality gazetteers is very time consuming. So the another important outcome of this study is the presented language resources open for the usage of all researchers. This language resources include: 1. Annotated dataset of Turkish newspaper texts. (500K tokens) 2. Annotated dataset of Turkish tweets (50K tokens) 3. Annotated dataset of ITU Web treebank 4. Person first names gazetteer (44K tokens) 5. Person surnames gazetteer (138K tokens) 6. Location names gazetteer (33K tokens) 7. 6 other relatively small gazetteers In conclusion a succesful model is proposed for Turkish NER problem and some important resources have been created and presented for future research. All results and the system itself is open for usage of all researchers.

Benzer Tezler

  1. Kısa metinlerde varlık ismi tanıma

    Named entity recognition on Turkish short texts

    BEYZA EKEN

    Yüksek Lisans

    Türkçe

    Türkçe

    2015

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    YRD. DOÇ. DR. AHMET CÜNEYD TANTUĞ

  2. Analysis of natural language processing techniques and development of Turkish named entity recognition tool for travel-tourism voice assistant

    Doğal dil işleme tekniklerinin incelenmesi ve seyahat-turizm sesli asistanı için Türkçe varlık ismi tanıma aracı geliştirilmesi

    DENİZ GÜL ÖZCAN

    Yüksek Lisans

    İngilizce

    İngilizce

    2020

    Mühendislik BilimleriAkdeniz Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DOÇ. DR. ÜMİT DENİZ ULUŞAR

  3. Tek yumurta ikizlerinin el yazılarının incelenmesi

    Başlık çevirisi yok

    BEHİCE ŞEYDA TÜREDİ

    Yüksek Lisans

    Türkçe

    Türkçe

    2016

    Adli TıpAnkara Üniversitesi

    Adli Bilimler Ana Bilim Dalı (disiplinlerarası)

    DOÇ. DR. NERGİS CANTÜRK

  4. Improving self-attention based transformer performance for morphologically rich languages

    Morfolojik açıdan zengin diller için öz dikkat tabanlı dönüştürücü performansının iyileştirilmesi

    YİĞİT BEKİR KAYA

    Doktora

    İngilizce

    İngilizce

    2024

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DOÇ. DR. AHMET CÜNEYD TANTUĞ

  5. Kıyıboyu katı madde modelleriyle Doğu Karadeniz bölgesinde kıyı çizgisi değişimlerinin incelenmesi

    Investigation of the shorline change at the black sea region with longshore sediment transport models

    MUSTAFA ŞAŞAL

    Doktora

    Türkçe

    Türkçe

    2000

    İnşaat Mühendisliğiİstanbul Teknik Üniversitesi

    PROF.DR. NECATİ AĞIRALİOĞLU