Geri Dön

ArGemma: Gemma'nın Arapçaya uyarlanddırılması için ince ayar ve çok görevli öğrenme mimarisi

ArGemma: A fine-tuning and multi-task learning architecture for adapting Gemma to Arabic

  1. Tez No: 1017295
  2. Yazar: TAHA AL-SELWI
  3. Danışmanlar: DR. ÖĞR. ÜYESİ SERAP ÇAKAR KAMAN
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2026
  8. Dil: Türkçe
  9. Üniversite: Sakarya Üniversitesi
  10. Enstitü: Fen Bilimleri Enstitüsü
  11. Ana Bilim Dalı: Bilgisayar Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Bilgisayar Mühendisliği Bilim Dalı
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Bu tez çalışması, düşük kaynaklı Arapça Doğal Dil İşleme alanında süregelen zorlukları ele almak amacıyla geliştirilen, Google'ın açık kaynaklı Gemma modelinin Arapça'ya uyarlanmış bir versiyonu olan ArGemma'yı tanıtmaktadır. Büyük Dil Modelleri İngilizce ve diğer yüksek kaynaklı dillerde son teknoloji düzeyinde sonuçlar elde etmiş olsalar da, Arapça'da çeşitli ve kapsamlı veri kümelerinin eksikliği ile tam ince ayarın yüksek hesaplama maliyeti nedeniyle performans sınırlı kalmıştır. Bu araştırma, denetimli ince ayar ile düşük dereceli uyum yöntemlerini birleştiren verimli bir ince ayar çerçevesi önermekte, ayrıca SFT'nin bilgi getirmeye dayalı üretim ve az örnekli ipucu verme yaklaşımlarıyla bütünleştirildiği hibrit bir yapı ile genişletilmektedir. Sistem, çeviri, özetleme, hikâye ve diyalog üretimi ile klasik Arapça ve modern standart Arapça arasındaki köprüyü kapsayan dokuz titizlikle hazırlanmış Arapça veri kümesi üzerine inşa edilmiştir. Her veri kümesi, denetimli eğitim süreciyle uyumlu olacak şekilde talimat–yanıt çiftleri biçiminde yeniden düzenlenmiştir. Alan temelli kaynaklar —örneğin tefsirli Kur'an metinleri, Arap atasözleri, sözlükler ve şiirler— ise getirmeye dayalı öğrenmeyi destekleyen anlamsal bir bilgi tabanı oluşturmak için kullanılmıştır. Bu mimari sayesinde ArGemma, dilsel zenginliği ve kültürel tutarlılığı korurken verimli bir uyarlama sağlamaktadır. LoRA tabanlı parametre-verimli ince ayar yöntemi, modelin ana ağırlıklarını sabit tutarak yalnızca %1'den az parametrenin güncellenmesini sağlar. Böylece bellek kullanımı ve eğitim süresi önemli ölçüde azalır. ArGemma'nın dört farklı sürümü geliştirilmiştir: LoRA-Translation, LoRA-Summarization, LoRA-Generation ve LoRA-ClassicMSA (Hibrit). Tüm sürümler, Kaggle platformundaki GPU/TPU altyapısı kullanılarak eğitilmiş ve küçük araştırma ekiplerinin erişebileceği biçimde tekrarlanabilirlik sağlanmıştır. Modelin performansı, nicel (BLEU, ROUGE) ve nitel (insan değerlendirmeleri) kriterlerini birleştiren kapsamlı bir çerçeveyle ölçülmüştür. Sonuçlar, çeviri görevinde BLEU puanının 2B model için 0.31'den 0.70'e, 9B model için 0.48'den 0.77'ye yükseldiğini göstermektedir. ROUGE metrikleri ise özetleme kalitesinde tutarlı artışlar ortaya koymuştur. Arapça ana dili olan değerlendiriciler ve dil uzmanları, özellikle klasik ve modern Arapça arasındaki anlam köprülemede modelin akıcılık, bağdaşıklık ve kültürel uygunluk açısından yüksek performans sergilediğini belirtmiştir. Bu bulgular, bilgi getirme ve az örnekli ipucu verme yöntemleriyle birleştirilen parametre-verimli ince ayarın, yüksek ve düşük kaynaklı diller arasındaki performans farkını önemli ölçüde kapatabileceğini ortaya koymaktadır. ArGemma, çeviri, özetleme, hikâye anlatımı ve diyalog üretimi gibi temel görevlerde güçlü sonuçlar elde ederken, aynı zamanda Arapça dil varyantları arasında bağlamsal anlayışı güçlendiren yenilikçi bir hibrit mimari sunmaktadır. Bunun ötesinde, çalışma açık kaynak yapay zekâ araştırmalarının demokratikleşmesine katkı sağlamaktadır. ArGemma, ileri seviye dil modeli ince ayarını büyük donanım kaynaklarına gerek kalmadan mümkün kılarak araştırmacılara yeni olanaklar sunmaktadır. Sistem modüler tasarımı sayesinde gelecek çalışmalar için ölçeklenebilir bir temel oluşturur: farklı Arap lehçelerinin dahil edilmesi, insan geri bildiriminden pekiştirmeli öğrenme ve metin, görsel ve ses gibi çoklu modal verilerin işlenmesi bu yönelimlerden bazılarıdır. Sonuç olarak, bu tez verimli, bütüncül ve kültürel açıdan farkında bir Arapça dil modeli uyarlaması sunmakta; morfolojik açıdan zengin ve düşük kaynaklı diller için yeni bir ince ayar standardı ortaya koymaktadır. ArGemma, Arapça NLP'nin yeteneklerini geliştirmekle kalmayıp, küresel yapay zekâ araştırmalarında kapsayıcılığı artıracak ölçeklenebilir bir araştırma çerçevesi de sağlamaktadır.

Özet (Çeviri)

This thesis presents ArGemma, an Arabic-adapted variant of Google's open-source Gemma model, developed to overcome the persistent and multifaceted challenges of Arabic natural language processing (NLP) in low-resource settings. While recent advancements in large language models (LLMs) have dramatically transformed NLP for high-resource languages—enabling human-like generation, reasoning, and understanding—the advancements have not translated equally across languages. Arabic, despite being one of the world's most widely spoken languages, continues to suffer from a scarcity of high-quality, diverse, and domain-rich datasets. Its complex morphology, root-based structure, rich derivational patterns, and wide spectrum of dialectal and stylistic variations present additional obstacles that limit the applicability of mainstream LLMs. The linguistic gap becomes even more evident when considering the computational burden associated with full fine-tuning of large models, which places a heavy barrier on researchers and institutions lacking advanced hardware resources. Motivated by these limitations, this thesis introduces a scalable, cost-efficient, and culturally aligned framework for adapting LLMs to Arabic through a combination of Supervised Fine-Tuning (SFT), Low-Rank Adaptation (LoRA), Retrieval-Augmented Generation (RAG), and Few-Shot Prompting (FSP). The proposed framework is built upon nine carefully curated Arabic datasets representing a wide variety of tasks: machine translation, abstractive summarization, story generation, dialogue construction, and semantic bridging between Classical Arabic and Modern Standard Arabic (MSA). These datasets were not only collected but meticulously refined, cleaned, and transformed into instruction–response format—a structure that has become the standard for modern alignment-tuned models. Additionally, domain-specific resources such as Quranic text with Tafsir, Arabic Proverbs, lexicons, and classical poetry were leveraged to build a knowledge-rich retrieval base supporting the RAG component. This hybrid formulation allows the model to integrate explicit linguistic and cultural knowledge while maintaining the generative fluency of LLMs. Through this hybridization, ArGemma achieves a balanced adaptation: it benefits from general-purpose language modeling while grounding its outputs in domain-relevant Arabic sources, thus significantly mitigating issues related to hallucination, semantic inconsistency, and cultural misalignment. Fine-tuning was carried out using LoRA, a parameter-efficient method that updates less than 1% of model parameters while keeping the base model frozen. This dramatically reduces memory requirements, training time, and computational overhead—making the adaptation of high-quality LLMs feasible even in modest environments such as Kaggle GPUs and TPUs. Four specialized versions of ArGemma were developed: LoRA-Translation for bilingual generation, LoRA-Summarization for Arabic text condensation, LoRA-Generation for creative story and dialogue tasks, and LoRA-ClassicMSA—a hybrid model integrating SFT with RAG and FSP for bridging stylistic and semantic gaps between Classical Arabic and MSA. Each version was trained using Kaggle's high-performance GPU/TPU infrastructure, ensuring reproducibility, accessibility, and practical scalability for researchers without access to enterprise-level hardware. To evaluate the performance of ArGemma, a comprehensive framework combining quantitative and qualitative assessments was employed. Quantitative evaluation for translation used BLEU scores, which improved significantly from 0.31 to 0.70 for the 2B model and from 0.48 to 0.77 for the 9B model—demonstrating the strong impact of domain-aligned, instruction-formatted data. Summarization performance was evaluated through ROUGE-1 and ROUGE-L, revealing consistent gains in recall, precision, and F1-score across all model variants. These improvements highlight the effectiveness of the LoRA fine-tuning process in enabling the model to capture key information while maintaining fluency and structural coherence. Qualitative evaluation was conducted through human assessments involving native speakers and professional Arabic linguists. They evaluated generated texts based on fluency, lexical appropriateness, contextual understanding, logical flow, and cultural alignment. Their analyses indicated a substantial improvement, especially in tasks involving narrative structure, dialogue naturalness, and Classical–MSA bridging. The hybrid LoRA-ClassicMSA variant, in particular, demonstrated high sensitivity to linguistic nuance, capturing not only lexical meaning but also stylistic subtleties and culturally embedded expressions. One of the central contributions of this thesis is demonstrating that parameter-efficient fine-tuning, when combined with retrieval-based and example-based augmentation strategies, can significantly close the performance gap between high- and low-resource languages. The hybrid use of RAG and FSP enhances the model's grounding, enabling it to generate culturally aware, semantically precise, and stylistically consistent text. This is particularly important in Arabic, where meaning is deeply intertwined with morphology, syntax, and literary tradition. By enabling the model to retrieve real examples and context-specific references during generation, the RAG component reduces hallucination and enhances factual accuracy. Meanwhile, Few-Shot Prompting introduces flexibility, allowing the model to generalize across a wider range of linguistic tasks even with limited training examples. Together, these elements form a unified framework that elevates the reliability and quality of Arabic language modeling. Beyond performance scores, the thesis emphasizes the broader implications of open-source AI democratization. Many groundbreaking LLMs remain accessible only through APIs, limiting research freedom and restricting adaptation, especially for underrepresented languages. ArGemma shows that with careful dataset design and efficient fine-tuning strategies, it is possible to create high-quality, domain-adapted Arabic models using entirely open-source tools and publicly available computational resources. This holds significant value for academic researchers, government agencies, cultural institutions, and developers working in Arabic-speaking regions. By lowering the entry barrier, the proposed framework enables wider participation in Arabic AI development and challenges the dominance of high-resource languages in shaping AI research trends. In addition to technological efficiency, this work contributes semantically and culturally to the field. Arabic is not a monolithic language but a continuum of registers—from Quranic Classical Arabic to formal MSA and diverse regional dialects. The inclusion of Classical–MSA bridging as a dedicated task represents one of the first attempts to operationalize stylistic transformation within Arabic LLMs. The model demonstrates the ability to modernize classical expressions while preserving their meaning and cultural tone, a capability highly valuable for digital humanities, education, and automated content modernization. Moreover, the curated datasets supporting this task form an important resource for future researchers aiming to investigate cross-register semantic mapping or stylistic adaptation in Arabic. This research also introduces nine Arabic datasets covering diverse linguistic domains—translation, summaries, dialogues, narratives, lexicons, poetry, religious text with interpretations, and proverbial expressions. These datasets, converted into instruction–response format, serve both as training resources and as benchmarks for future work. Their contribution lies not merely in their content but also in their structure: by aligning them with modern instruction-tuning practices, the thesis establishes a consistent foundation for future Arabic LLM adaptation projects. Given the historical scarcity of high-quality, open-access Arabic NLP datasets, the availability and organization of these resources represent a substantial contribution to the field. The thesis further proposes a methodological workflow that can be replicated and extended beyond Arabic. The combination of LoRA-based fine-tuning, data curation, retrieval augmentation, and task-specific model variants offers a generalizable pipeline for adapting LLMs to other low-resource languages with complex morphology, such as Amharic, Somali, Kurdish, or Pashto. The modular nature of the system makes it easy to incorporate additional components such as dialect adaptation, multimodal processing, or reinforcement learning from human feedback (RLHF). In fact, RLHF is identified as one of the most promising directions for future work, as it would allow ArGemma to refine its responses according to human preferences, ethical considerations, and contextual expectations specific to Arabic-speaking communities. From a computational perspective, LoRA proves to be an ideal method for adapting LLMs in environments with limited hardware. Updating small low-rank matrices instead of the entire parameter set not only reduces memory load but also prevents catastrophic forgetting—a problem common when models are over-fine-tuned on narrow datasets. This stability is crucial for ensuring that ArGemma maintains the broad general-language capabilities of Gemma while gaining deep expertise in Arabic tasks. Furthermore, the decision to train on Kaggle's infrastructure highlights the viability of an accessible, reproducible research pathway—free from the constraints of proprietary cloud services. In conclusion, this thesis delivers a comprehensive, technically robust, and culturally informed framework for adapting open-source LLMs to Arabic. ArGemma stands as a practical demonstration that low-resource languages can achieve competitive NLP performance through efficient fine-tuning, carefully designed datasets, and intelligent hybrid learning strategies. The model not only improves Arabic NLP capabilities in translation, summarization, dialogue, and storytelling but also pioneers the integration of Classical–MSA semantic bridging as a distinct task—opening doors to novel applications in education, digital heritage, and computational linguistics. More broadly, the thesis contributes to the democratization of AI research, showing that high-quality adaptation is achievable without large-scale proprietary infrastructures. By making Arabic LLM research more accessible, the study supports future advancements in dialect modeling, RLHF-based refinement, multimodal Arabic AI, and cross-cultural computational linguistics. ArGemma thus represents both an immediate contribution to Arabic NLP and a blueprint for inclusive, sustainable AI development in low-resource language contexts worldwide.

Benzer Tezler