A hierarchical framework for preventing robot manipulation failures
Robot etkileşimlerinde hataları önlemek için çok-katmanlı güvenli yürütme mimarisi
- Tez No: 1014686
- Danışmanlar: PROF. DR. SANEM SARIEL UZER, DOÇ. DR. EREN ERDAL AKSOY
- Tez Türü: Doktora
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Makine öğrenmesi, Otonom robotlar, Yapay zeka, Machine learning, Autonomous robots, Artificial intelligence
- Yıl: 2026
- Dil: İngilizce
- Üniversite: İstanbul Teknik Üniversitesi
- Enstitü: Lisansüstü Eğitim Enstitüsü
- Ana Bilim Dalı: Bilgisayar Mühendisliği Ana Bilim Dalı
- Bilim Dalı: Bilgisayar Mühendisliği Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
Günümüzde robotların yaşamımızdaki rolü giderek artmakta, yalnızca endüstriyel ortamlarda değil, aynı zamanda günlük yaşantımızda da daha görünür hale gelmektedir. Özellikle ev içi hizmet robotları, sağlık destek sistemleri, yaşlı bakımı ve çocuk eğitimi gibi alanlarda insanlar ile doğrudan temas hâlinde çalışan robotların sayısı her geçen gün çoğalmaktadır. Bu yakın fiziksel etkileşim, robotların yalnızca görevleri başarıyla yerine getirmesini değil, aynı zamanda güvenli bir şekilde çalışmasını da zorunlu kılmaktadır. İnsanlarla aynı ortamı paylaşan robotların güvenliği, olası fiziksel zararların önlenmesi, insan-robot etkileşiminde güvenin sağlanması ve toplum tarafından kabul edilebilirliğin artırılması açısından kritik bir öneme sahiptir. Bu bağlamda, robotların hem görev odaklı verimliliklerini hem de çevresel ve sosyal güvenliklerini bir arada sağlayacak yaklaşımlar geliştirmek, modern robotik sistemlerin geleceği açısından kaçınılmaz bir ihtiyaç haline gelmiştir. Bu nedenlerden dolayı bu tez çalışmasında, robot beceri öğrenimi ve güvenli robot etkileşimi konularına odaklanılmış ve bu alandaki mevcut zorluklara çözüm olarak iki temel yaklaşım önerilmiştir: SafeEx (Güvenli Yürütme Döngüsü) ve SafeIt (Güvenli İteratif Güncelleme Döngüsü). Bu mimariler, robotların yalnızca görev odaklı hareket etmesini değil, aynı zamanda karşılaştıkları hataları tanıyıp önleyebilmelerini ve bu süreçte yeni önleyici beceriler kazanarak güvenlik seviyelerini artırmalarını sağlamayı hedeflemektedir. Günümüzde robotlar birçok görevde başarılı olsa da, beklenmeyen çevresel değişiklikler, sensör hataları veya eksik ön bilgiler gibi nedenlerle başarısızlıklarla karşılaşabilmektedir. Bu başarısızlıklar, robot sistemlerinin yalnızca görev başarısını değil, aynı zamanda güvenliğini de tehdit etmektedir. Bu bağlamda tez, robotların beceri öğrenim sürecine güvenlik faktörlerini entegre eden modüler ve esnek bir çözüm sunmaktadır. SafeEx mimarisi, robotların görevlerini güvenli şekilde yürütebilmesini sağlamak için tasarlanmış bir yürütme mekanizmasıdır. Bu yapı, temel görev becerilerini (base skills) ve hata önleme becerilerini (failure prevention skills) içeren modüler bir beceri kitaplığına dayanmaktadır. Bu becerilerden hangisinin kullanılacağı, gerçek zamanlı olarak risk sınıflandırıcıları tarafından belirlenen risk seviyelerine göre seçilmektedir. Eğer potansiyel bir hata riski tespit edilirse, sistem otomatik olarak temel beceri yerine ilgili önleyici beceriyi devreye sokmaktadır. Bu sayede robot, görevi terk etmeden güvenliğini koruyarak hareket etmeye devam edebilmektedir. SafeEx'in en güçlü yönlerinden biri, modüler olması sayesinde farklı hata türlerine yönelik becerilerin bağımsız şekilde öğrenilmesine ve sisteme entegre edilmesine olanak tanımasıdır. SafeEx bilinen hatalar için etkili bir yapı sunarken, daha önce görülmemiş yeni hatalara karşı kendiliğinden uyum sağlayamamaktadır. Bu sorunu çözmek amacıyla tezde SafeIt mimarisi geliştirilmiştir. SafeIt, yürütme sırasında ortaya çıkan başarısızlıkları algılayarak yeni hataları keşfeder ve robotun bu hataları sınıflandırmasına, yeniden üretmesine ve önlemesine olanak tanıyan bir öğrenme döngüsünü başlatır. Bu süreç, insan uzman tarafından yalnızca bir kritik anın işaretlenmesiyle başlatılır. Sonrasında robot bu bilgiyi kullanarak bir risk sınıflandırıcı (failure classifier) öğrenir, ardından agresif bir beceri (aggressor skill) ile hatayı yeniden üretmeyi öğrenir ve son olarak hata önleme becerisini geliştirerek sistemi günceller. Bu öğrenilen yeni beceriler ve sınıflandırıcılar, SafeEx döngüsüne dahil edilerek robotun güvenlik kabiliyetleri artırılır. Tezde önerilen yöntemlerin etkinliği hem simülasyon ortamlarında hem de gerçek robotlarla yapılan deneylerde gösterilmiştir. Karıştırma (stir) ve itme (push) gibi görevlerde robotlara SafeEx ve SafeIt mimarileri entegre edildiğinde, başarı oranları ve güvenlik performansları önemli ölçüde artmıştır. Özellikle karmaşık çoklu görev senaryolarında SafeEx, görevleri güvenli şekilde yerine getirerek diğer geleneksel yöntemlere kıyasla daha esnek ve etkili çözümler sunmuştur. Bunun yanında SafeIt, zamanla karşılaşılan yeni hata türlerine adaptasyon sağlayarak sistemin öğrenme kapasitesini sürekli güncellemiş ve yaşam boyu öğrenme (lifelong learning) açısından önemli bir potansiyel sunmuştur. Bu çalışmanın önemli katkılarından biri de becerilerin modüler ve yeniden kullanılabilir şekilde tasarlanmasıdır. Tek bir bileşik politika (compound policy) yerine, her bir hata türü için bağımsız beceriler geliştirilmiş ve sistem bu becerileri gerektiğinde otomatik olarak devreye alacak şekilde yapılandırılmıştır. Böylece sistem, yeniden eğitim gerektirmeden yeni beceriler ekleyebilir ve esnek bir şekilde genişleyebilir. Ek olarak, SafeIt ile öğrenilen hata önleme becerilerinin farklı görevlerde de tekrar kullanılabildiği ve bu sayede öğrenme verimliliğinin artırıldığı gösterilmiştir. Gerçek dünyaya aktarım (sim-to-real) problemine yönelik olarak da, simülasyon ve fiziksel robotlar arasında durum ve hareket alanlarının hizalanması sağlanarak aktarım başarısı güvence altına alınmıştır. Bu çalışmada, yukarıda anlatılan problemlere, robot öğrenmesi alanında önerilen çözümlere odaklanabilmek amacıyla bazı sınırlamalar getirilmiştir. Özellikle hata algılama ve tanımlama konularında kullanılan yöntemler sınırlı kalmış ve daha çok öğrenmeye odaklanılmıştır. Hataların tanımlanması sürecinde insan uzman müdahalesine ihtiyaç duyulması, sistemin tam anlamıyla özerk olmadığını göstermektedir. Gelecekte bu sürecin daha da otomatik hale getirilmesi için kendi kendine gözetimli (self-supervised) veya denetimsiz (unsupervised) öğrenme tekniklerinin entegrasyonu faydalı olabilir. Ayrıca çoklu duyusal veri kaynaklarının (görsel, işitsel, dokunsal vb.) birleştirilmesiyle sistemin risk sınıflandırma kapasitesinin artırılması da potansiyel bir geliştirme alanıdır. Bu tezde sunulan SafeEx ve SafeIt mimarileri, görev başarımı ile güvenliği bir arada ele alan bütüncül bir yaklaşım sunmaktadır. Özellikle insan-robot etkileşimi, ev robotları, hizmet robotları ve endüstriyel otomasyon gibi uygulama alanlarında bu tür sistemlerin gerekliliği her geçen gün artmaktadır. Robotların yalnızca görevleri yerine getirmesi değil, aynı zamanda bunu güvenli bir biçimde yapabilmesi, bu teknolojilere duyulan güveni artırmakta ve daha geniş kullanım alanlarına imkan sağlamaktadır. Bu bağlamda tez, gelecekteki bilişsel ve güvenlik odaklı robot sistemleri için sağlam bir temel oluşturmakta ve özerk robotların hem görev hem de güvenlik performanslarını iyileştirme yönünde önemli katkılar sunmaktadır. Sonuç olarak, bu tez güvenli ve uyarlanabilir robot etkileşimi alanında önemli bir boşluğu doldurmakta ve hem akademik hem de endüstriyel uygulamalar açısından yüksek etki potansiyeline sahip bir yaklaşım ortaya koymaktadır. SafeEx ile güvenli yürütme sağlanırken, SafeIt ile yeni hatalara karşı adaptasyon kabiliyeti kazandırılmıştır. Böylece robotlar yalnızca başarılı değil, aynı zamanda güvenli şekilde görevlerini sürdürebilen ve değişen koşullara öğrenme yoluyla uyum sağlayabilen sistemler haline gelmiştir.
Özet (Çeviri)
The presence of robots in our daily lives is steadily increasing. While they have traditionally been confined to industrial settings, robots are now entering human-centered environments such as homes, hospitals, schools, and public spaces. In particular, service robots operating in domestic settings, providing assistance with chores, eldercare, and personal support, are becoming more common. As robots begin to physically share space with humans and interact with them directly, ensuring their safety becomes a critical concern. It is no longer sufficient for robots to simply perform tasks successfully; they must also do so without posing any harm to their human counterparts. The closer robots get to humans, the more important it becomes to develop systems that can guarantee safe operation, manage unpredictable interactions, and respond appropriately to unforeseen situations. Ensuring robotic safety not only prevents potential accidents or injuries but also fosters human trust and social acceptance. Therefore, designing robots that are both functionally competent and inherently safe is essential for the widespread and sustainable integration of robotics into our everyday lives. This thesis addresses the critical challenge of safe robot manipulation by integrating skill learning with safety assurance through two interconnected frameworks: SafeEx (Safe Execution Loop) and SafeIt (Safe Iterative Update Loop). The central motivation behind this work lies in the observation that robot skills, while effective in accomplishing specific tasks, often lack mechanisms to prevent or adapt to failures, especially when deployed in dynamic or unstructured environments. Failures in robotic manipulation, such as spills, collisions, or object drops, can result from environmental uncertainty, sensor noise, or incomplete prior knowledge. These failures pose risks not only to task success but also to the safety and reliability of robotic systems. As such, the thesis aims to enable robots to learn and apply failure prevention strategies in a scalable, interpretable, and reactive manner. The SafeEx framework serves as the core runtime system for managing safe robot operation. It relies on a modular skill library that includes both task-oriented base skills and risk-mitigating failure prevention skills. When a known risk is detected, through risk classifiers trained to recognize specific unsafe situations, the high-level skill selection mechanism within SafeEx dynamically switches from the base skill to the corresponding failure prevention skill. This reactive skill-switching mechanism ensures that robots can handle safety-critical situations without abandoning their task objectives. Importantly, the modularity of the approach allows each failure type to be handled independently through a corresponding failure prevention policy, which enables more effective learning, easier debugging, and higher interpretability. While SafeEx is effective for handling known failure types, it cannot, by itself, address failures that arise unexpectedly during execution. This limitation is overcome by the introduction of SafeIt, a complementary learning loop that is invoked when unknown failures occur. SafeIt begins with a failed trajectory, from which a human expert labels a single point marking the onset of failure. From this minimal supervision, the system constructs a one-shot classifier to estimate the failure risk in future executions. Using this classifier, an adversarial skill (termed the aggressor) is learned to actively reproduce risky states, thereby allowing reinforcement learning to sample data efficiently for training the corresponding failure prevention skill. Once the classifier and prevention skill are learned, they are integrated back into the SafeEx loop, incrementally expanding the robot's safety capabilities. The thesis provides extensive experimental evidence supporting the effectiveness of these frameworks. In a set of case studies involving both simulated and real-world tasks, such as stirring and pushing objects, robots augmented with SafeEx and SafeIt demonstrate significantly higher success rates and robustness compared to baseline methods. In single-task experiments, SafeIt-augmented hand-coded policies showed marked improvements in failure tolerance. In more complex multi-task scenarios involving multiple failure types, SafeEx was able to manage task execution with increased adaptability and precision. These results highlight the scalability of the modular approach, especially when compared to compound reinforcement learning models that attempt to encode both task objectives and safety constraints into a single policy, often resulting in conservative or ineffective behaviors. One of the key strengths of this work is its emphasis on modularity and reusability. Unlike monolithic models, which must be retrained entirely to handle new failures, the SafeEx/SafeIt system allows new skills and classifiers to be appended incrementally. This supports real-time adaptation and continual learning without catastrophic forgetting. Furthermore, the failure prevention skills are designed to be generic and reusable across different tasks and base skills. This increases learning efficiency and reduces the need for exhaustive retraining. The thesis also addresses the sim-to-real transfer problem by carefully aligning state and action spaces between simulation and physical environments. Real-world validations confirm the practical feasibility of the approach. Despite these contributions, the thesis also identifies several limitations. While the focus has been on integrating failure prevention into skill learning, perception-based safety mechanisms such as anomaly detection and general failure identification were not explored in depth. The reliance on a human expert to initialize SafeIt with labeled failure onset points also highlights the semi-autonomous nature of the current system. Future research could focus on minimizing this human involvement by incorporating self-supervised or unsupervised failure detection techniques. Another area for development is the integration of multi-modal sensory data, which could allow the risk classifiers to generalize more effectively to complex real-world conditions. The potential future directions for this line of research are rich and promising. One natural extension is the exploration of lifelong learning paradigms, where the robot continuously accumulates experience and autonomously refines its risk models and skill library over extended periods of deployment. Another avenue involves integrating natural language guidance or user preferences into the safety system, enabling more human-centered interaction and collaborative safety planning. Additionally, expanding the framework to cooperative multi-robot systems could yield new insights into distributed safety management and coordination in shared workspaces. The broader impact of this thesis lies in its advancement of safety as a first-class objective in robot learning. The proposed frameworks offer a pathway toward more reliable, interpretable, and user-aligned robotic systems capable of adapting to dynamic environments and evolving safety requirements. In industrial automation, service robotics, or human-robot interaction domains, these capabilities are essential not only for functional success but also for building trust and ensuring compliance with safety standards. By decoupling task performance from failure prevention and adopting a structured, modular, and learnable approach, this work makes a compelling case for a shift in how safety is conceptualized in reinforcement learning for robotics. In conclusion, this thesis presents a significant contribution to the field of safe and adaptive robot manipulation by proposing and validating a unified framework that integrates skill learning, safety assurance, and incremental adaptation. Through the SafeEx and SafeIt frameworks, it bridges the gap between theoretical reinforcement learning and the practical requirements of safety in real-world deployment. The work lays a strong foundation for future explorations into autonomous systems that can not only perform complex tasks but do so in a manner that is safe, resilient, and trustworthy.
Benzer Tezler
- 1985-1986 yıllarında Polatlı Devlet Hastanesine kaza nedeniyle başvuranların incelenmesi
Başlık çevirisi yok
FERHAN ŞENOL
Yüksek Lisans
Türkçe
1986
İlk ve Acil YardımGazi ÜniversitesiKazaların Çevresel ve Teknik Araştırması Ana Bilim Dalı (disiplinlerarası)
DOÇ. DR. HİKMET PEKCAN
- Devrelerin büyük işaret cevaplarının bilgisayar yardımıyla bulunması
Başlık çevirisi yok
NÜKHET GÜNEYİ
Yüksek Lisans
Türkçe
1985
Elektrik ve Elektronik MühendisliğiUludağ ÜniversitesiElektrik-Elektronik Mühendisliği Ana Bilim Dalı
PROF. DR. ERGÜR TÜTÜNCÜOĞLU
- Ulaş Sağlık Ocağı merkezinde nüfusun bazı niteliklerine ve konutların durumuna ilişkin bir çalışma
An Investigation carried out at the Ulaş Public Health Centre concerning some characteristics of the population and housing conditions in the district of Ulaş
EROL ŞANLI
Doktora
Türkçe
1985
Halk SağlığıCumhuriyet ÜniversitesiHalk Sağlığı Ana Bilim Dalı
DOÇ. DR. SERVET ÖZGÜR
- Çimentonun sertleşmesi üzerinde kimyasal komponentlerin etkisi
Başlık çevirisi yok
NACİYE TÜRKEL
Yüksek Lisans
Türkçe
1986
Kimya MühendisliğiUludağ ÜniversitesiKimya Ana Bilim Dalı
PROF. DR. MUSTAFA CEBE
- Antakya tarihi ticaret merkezi mekansal yapı değişim ve gelişim sürecinin kent ticaret merkezi planlamasına etkinliği
The Effectiveness of spatial structural chance and development period of the historical trade center of Antakya in the urban trade planning
NEVİN TURGUT
Yüksek Lisans
Türkçe
1986
Şehircilik ve Bölge PlanlamaGazi ÜniversitesiŞehir ve Bölge Planlama Ana Bilim Dalı
YRD. DOÇ. DR. NURCAN UYDAŞ