Geri Dön

Multimodal Machine Translation

Başlık çevirisi mevcut değil.

  1. Tez No: 622480
  2. Yazar: OZAN ÇAĞLAYAN
  3. Danışmanlar: PROF. DR. DANIŞMAN YOK
  4. Tez Türü: Doktora
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2019
  8. Dil: İngilizce
  9. Üniversite: Le Mans Unıversıte
  10. Enstitü: Yurtdışı Enstitü
  11. Ana Bilim Dalı: Belirtilmemiş.
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Machine translation aims at automatically translating documents from one language to another without human intervention. With the advent of deep neural networks (DNN), neural approaches to machine translation started to dominate the €eld, reaching stateof- the-art performance in many languages. Neural machine translation (NMT) also revived the interest in interlingual machine translation due to how it naturally €ts the task into an encoder-decoder framework which produces a translation by decoding a latent source representation. Combined with the architectural ƒexibility of DNNs, this framework paved the way for further research in multimodality with the objective of augmenting the latent representations with other modalities such as vision or speech, for example. ‘is thesis focuses on a multimodal machine translation (MMT) framework that integrates a secondary visual modality to achieve be‹er and visually grounded language understanding. I speci€cally worked with a dataset containing images and their translated descriptions, where visual context can be useful for word sense disambiguation, missing word imputation, or gender marking when translating from a language with gender-neutral nouns to one with grammatical gender system as is the case with English to French. I propose two main approaches to integrate the visual modality: (i) a multimodal a‹ention mechanism that learns to take into account both sentence and convolutional visual representations, (ii) a method that uses global visual feature vectors to prime the sentence encoders and the decoders. ‘rough automatic and human evaluation conducted on multiple language pairs, the proposed approaches were demonstrated to be bene€cial. Finally, I further show that by systematically removing certain linguistic information from the input sentences, the true strength of both methods emerges as they successfully impute missing nouns, colors and can even translate when parts of the source sentences are completely removed.

Özet (Çeviri)

Özet çevirisi mevcut değil.

Benzer Tezler

  1. Learning visually-grounded representationsusing cross-lingual multimodal pre-training

    Çok dilli çok kipli ön öğrenme ile görsel tabanlı temsillerin öğrenilmesi

    MENEKŞE KUYU

    Yüksek Lisans

    İngilizce

    İngilizce

    2020

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolHacettepe Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DOÇ. DR. MEHMET ERKUT ERDEM

  2. Müze iletişimi ve özel alan çevirisi

    Museum communication and specialised translation

    HANDE ÇİL TEYMOURI NAGHADEH

    Yüksek Lisans

    Türkçe

    Türkçe

    2024

    MüzecilikYıldız Teknik Üniversitesi

    Sanat ve Tasarım Ana Sanat Dalı

    PROF. DR. KADRİYE TEZCAN AKMEHMET

  3. Çok dilli ortamlarda gerçek zamanlı ses ve video konferans için ölçeklenebilir bir sistem tasarımı

    A scalable system design for real-time audio and video conferencing in multilingual environments

    EREN ÇAĞLAR

    Yüksek Lisans

    Türkçe

    Türkçe

    2026

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolYıldız Teknik Üniversitesi

    Veri Bilimi ve Büyük Veri Ana Bilim Dalı

    PROF. DR. MEHMET SIDDIK AKTAŞ

  4. Multilingual, multimodal and explainable approaches for automated fact-checking problem

    Otomatik doğrulama problemi için çok dilli, çok modlu ve açıklanabilir yaklaşımlar

    RECEP FIRAT ÇEKİNEL

    Doktora

    İngilizce

    İngilizce

    2025

    DilbilimOrta Doğu Teknik Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    PROF. DR. PINAR KARAGÖZ

  5. ArGemma: Gemma'nın Arapçaya uyarlanddırılması için ince ayar ve çok görevli öğrenme mimarisi

    ArGemma: A fine-tuning and multi-task learning architecture for adapting Gemma to Arabic

    TAHA AL-SELWI

    Yüksek Lisans

    Türkçe

    Türkçe

    2026

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolSakarya Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ SERAP ÇAKAR KAMAN