Geri Dön

Acil yardım çağrısı jestlerinin bilgisayarla görü ve derin öğrenme yöntemleriyle gerçek zamanlı tespiti

Real-time detection of emergency call gestures using computer vision and deep learning methods

  1. Tez No: 1018159
  2. Yazar: FERHAT ERDOĞAN
  3. Danışmanlar: DR. ÖĞR. ÜYESİ MUHAMMED KOTAN
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Bilim ve Teknoloji, Computer Engineering and Computer Science and Control, Science and Technology
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2026
  8. Dil: Türkçe
  9. Üniversite: Sakarya Üniversitesi
  10. Enstitü: Fen Bilimleri Enstitüsü
  11. Ana Bilim Dalı: Bilişim Sistemleri Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Günümüzde şiddet, zorlayıcı kontrol, kaçırılma ve ani sağlık problemleri gibi acil durumlar, bireylerin sözlü iletişim kurmasının mümkün olmadığı ya da ciddi risk oluşturduğu senaryoları beraberinde getirmektedir. Bu tür durumlarda bireylerin yardım taleplerini sessiz, hızlı ve dikkat çekmeden iletebilmeleri hem bireysel güvenlik hem de toplumsal müdahale mekanizmalarının etkinliği açısından kritik öneme sahiptir. Son yıllarda bu gereksinim doğrultusunda, el jestlerine dayalı acil yardım çağrıları uluslararası ölçekte yaygınlaşmış; özellikle Tehlikedeyim (Signal for Help) jesti, aile içi şiddet ve zorlayıcı kontrol bağlamlarında sessiz yardım çağrısı olarak kabul görmüştür. Ancak bu jestlerin insanlar tarafından algılanması görsel dikkat ve farkındalık düzeyine bağlı olduğundan, yardım çağrılarının gözden kaçması gibi ciddi riskler ortaya çıkabilmektedir. Bu tez çalışmasında, acil yardım çağrısı el jestlerinin bilgisayarla görü ve derin öğrenme yöntemleri kullanılarak gerçek zamanlı biçimde otomatik olarak tespit edilmesi problemi ele alınmıştır. Çalışma, hata maliyetinin yüksek olduğu güvenlik-kritik bir senaryoya odaklanmış ve Tehlikedeyim (Signal for Help) ile Yardıma İhtiyacım Var (Need Help) olmak üzere iki sınıflı bir algılama problemi olarak kurgulanmıştır. Böylece özellikle yanlış negatif (false negative) hataların ayrıntılı biçimde analiz edilmesi ve bu hataların güvenlik açısından doğurabileceği risklerin görünür kılınması amaçlanmıştır. Yöntemsel olarak çalışmada, tek aşamalı nesne tespiti mimarileri arasında yer alan YOLO (You Only Look Once) ailesinin güncel sürümleri kullanılmıştır. YOLOv8m, YOLOv9m, YOLOv10m, YOLOv11m ve YOLOv12m modelleri, aynı eğitim parametreleri ve deneysel koşullar altında karşılaştırmalı olarak değerlendirilmiştir. Modeller, SFH-Dataset üzerinde eğitilmiş; ayrıca genellenebilirliği ölçmek amacıyla bu çalışma kapsamında oluşturulan Özel Veri Kümesi üzerinde test edilmiştir. Tüm deneyler Google Colab ortamında NVIDIA A100 GPU kullanılarak gerçekleştirilmiş ve 50 ile 100 epoch senaryoları altında performans analizleri yapılmıştır. Bu çalışmada, tekil kare bazlı jest tespit çıktıları üzerinden doğrudan alarm üretmek yerine,“Need Help”jestini izleyen belirli bir zaman penceresi içerisinde“Signal for Help”jestinin doğrulanmasını esas alan zamansal bir karar mantığı benimsenmiştir. Sliding window temelli bu zamansal doğrulama yaklaşımı, alarm üretim sürecinde ardışık karelerdeki tutarlılığı dikkate alarak yanlış alarm üretimini azaltmayı ve gerçek yardım çağrılarının korunmasını sağlayan daha dengeli bir alarm üretim stratejisi oluşturmayı hedeflemektedir. Deneysel sonuçlar, tek kare bazlı alarm üretimine kıyasla yanlış alarm sayısının belirgin biçimde azaldığını, buna karşılık gerçek yardım çağrılarının büyük ölçüde korunabildiğini ve sistem genelinde F1-Skor değerinde kayda değer bir iyileşme elde edildiğini göstermiştir. Bu bulgular, alarm kararında zamansal bağlamın kullanılmasının sistem güvenilirliğini önemli ölçüde artırdığını ortaya koymaktadır. Değerlendirme sürecinde Doğruluk (Accuracy), Hassasiyet (Precision), Duyarlılık (Recall) ve F1-Skor (F1-Score) metrikleri kullanılmıştır. Elde edilen bulgular, SFH-Dataset üzerinde tüm modellerin birbirine yakın ve yüksek performans düzeylerine ulaştığını; ancak gerçek dünya koşullarını daha iyi yansıtan Özel Veri Kümesi üzerinde modeller arasında belirgin performans farkları oluştuğunu göstermiştir. Özel Veri Kümesi sonuçlarında, yanlış negatiflerin (Miss) güvenlik açısından baskın risk faktörü olduğu görülmüş; bu bağlamda YOLOv11m (100 epoch) modeli, en düşük yanlış negatif sayısı (Signal for Help için 130, Need Help için 197) ve en dengeli hata dağılımı ile acil yardım çağrısı jestlerinin tespiti için en güvenilir çözüm olarak öne çıkmıştır. YOLOv11m modelinin ürettiği daha dengeli kare bazlı tespit çıktıları, önerilen zamansal karar mekanizması ile birlikte kullanıldığında sistem düzeyinde daha güvenilir alarm davranışı elde edilmesini mümkün kılmıştır. Buna karşılık YOLOv12m (100 epoch) modeli yanlış pozitif üretmemesine rağmen, özellikle Signal for Help sınıfında yüksek kaçırma oranları nedeniyle güvenlik-kritik kullanım açısından temkinli değerlendirilmesi gereken bir davranış sergilemiştir. Sonuç olarak bu tez, acil yardım çağrısı jestlerinin tespiti probleminde yalnızca yüksek doğruluk değerlerinin yeterli olmadığını; hata türleri, genellenebilirlik ve sistem düzeyinde karar mantığının güvenlik-kritik uygulamalar açısından belirleyici olduğunu ortaya koymaktadır. Çalışmanın bulguları, gerçek zamanlı güvenlik sistemlerinin tasarımına ve akademik literatüre, özellikle yanlış negatif odaklı değerlendirme perspektifiyle ve zamansal doğrulama tabanlı karar mekanizmalarının önemini ortaya koyarak önemli katkılar sunmaktadır.

Özet (Çeviri)

In modern societies, emergency situations such as domestic violence, coercive control, abduction, and sudden medical crises increasingly arise in contexts where verbal communication is either impossible or may directly endanger the individual. Victims in such circumstances are often unable to speak freely, attract attention, or explicitly request assistance without escalating the risk they face. Consequently, the ability to convey a request for help silently, rapidly, and discreetly has become a critical requirement for both individual safety and the effectiveness of societal intervention mechanisms. In response to this need, visual hand gesture based emergency signaling methods have gained international recognition in recent years. Among these, the“Signal for Help”gesture has emerged as one of the most widely accepted silent distress signals, particularly in the context of domestic violence and coercive control. The gesture is intentionally simple, culturally neutral, and executable with a single hand, allowing it to be performed under constrained physical and social conditions. However, despite its conceptual effectiveness, the practical success of such gestures remains heavily dependent on human visual attention, situational awareness, and contextual interpretation. In crowded environments, low-light conditions, or surveillance scenarios involving multiple camera feeds, these gestures may easily be overlooked, misinterpreted, or entirely missed. This reliance on human perception introduces significant vulnerabilities into emergency response systems. Even trained observers may fail to notice brief or partially occluded gestures, while bystanders may lack awareness of their meaning altogether. As a result, missed detections can lead to delayed intervention or complete absence of assistance, potentially resulting in severe harm. These limitations highlight the need for automated, real-time, and reliable visual systems capable of detecting emergency hand gestures without relying solely on human observation. Recent advances in computer vision and deep learning have made it increasingly feasible to analyze human actions and gestures automatically from visual data. In particular, convolutional neural network based object detection models have demonstrated strong performance in real-time recognition tasks across a wide range of applications. However, much of the existing literature on hand gesture recognition focuses on human computer interaction, sign language translation, or assistive communication systems, where the cost of errors is relatively low and misclassifications are often tolerable. In contrast, emergency gesture detection represents a safety-critical application in which different types of errors have drastically different consequences This thesis addresses the problem of real-time automatic detection of emergency hand gestures using computer vision and deep learning techniques, explicitly framing the task as a safety-critical detection problem. Unlike accuracy-oriented gesture recognition studies, this work emphasizes the importance of error type analysis, particularly false negative errors. In the context of emergency signaling, a false negative failing to detect a genuine request for help may lead to intervention delays or life-threatening outcomes. Therefore, minimizing false negatives is considered more critical than optimizing overall accuracy alone. In this study, false negatives are also referred to as miss errors, explicitly representing cases in which genuine emergency gestures remain undetected by the system, thereby posing direct safety risks. The detection problem is formulated as a binary classification task involving two semantically related gesture classes:“Signal for Help”and“Need Help.”This formulation reflects real-world emergency scenarios in which individuals may express distress progressively rather than through a single isolated gesture. By focusing on two critical classes, the study enables a deeper and more interpretable analysis of error distributions while avoiding unnecessary model complexity. Furthermore, the gestures are not treated as independent visual events but as complementary indicators that may occur within a meaningful temporal sequence. Methodologically, the study employs recent versions of the YOLO (You Only Look Once) family of single-stage object detection architectures. YOLO models are particularly well suited for real-time applications due to their unified detection pipeline, which predicts bounding boxes and class probabilities in a single forward pass. Compared to two-stage detectors, YOLO architectures offer a favorable balance between inference speed and detection accuracy, making them suitable for continuous video-based monitoring systems in safety-critical contexts where low latency is essential. In this thesis, YOLOv8m, YOLOv9m, YOLOv10m, YOLOv11m, and YOLOv12m models are evaluated within a systematic and comparative framework. The medium-scale variants are selected to balance computational efficiency and representational capacity, reflecting realistic deployment constraints for real-time systems. All models are trained and tested under identical experimental conditions, including consistent hyperparameters, data splits, and evaluation protocols. This controlled setup ensures that observed performance differences can be attributed primarily to architectural characteristics rather than experimental variability. Two datasets are used to evaluate model performance and generalizability. The first is the publicly available SFH-Dataset, which provides a standardized benchmark for emergency hand gesture detection and has been adopted in prior research. The second is a custom dataset constructed within the scope of this thesis to better reflect real-world conditions. The custom dataset incorporates variations in lighting, camera angles, background complexity, and user appearance, introducing distributional shifts that challenge model robustness and generalization capability beyond controlled laboratory settings. All experiments are conducted in the Google Colab environment using an NVIDIA A100 GPU. Training is performed under two scenarios consisting of fifty and one hundred epochs in order to analyze convergence behavior, stability, and sensitivity to training duration. Standard evaluation metrics, including Accuracy, Precision, Recall, and F1-Score, are reported to maintain comparability with existing literature. However, the primary analytical emphasis is placed on confusion matrix–based error analysis, enabling a detailed examination of false negative and false positive distributions for each gesture class. Beyond frame-level detection, the thesis introduces a safety-oriented temporal decision mechanism designed to improve system reliability in real-world deployments. Instead of triggering an alarm based on a single-frame prediction, the system requires the detection of a“Signal for Help”gesture to be temporally confirmed following a“Need Help”gesture within a predefined time window. This strategy reflects real-world behavioral patterns and aims to reduce the impact of sporadic misdetections while prioritizing reliable identification of genuine emergency situations. Experimental analyses demonstrate that this temporal verification mechanism significantly reduces false alarms while preserving the detection of real emergency gestures, thereby producing a more stable and safety-oriented alarm generation behavior compared to single-frame detection strategies. Experimental results demonstrate that all evaluated YOLO models achieve similarly high performance on the SFH-Dataset, indicating that controlled benchmark datasets may mask meaningful architectural differences. In contrast, substantial performance divergences emerge on the custom dataset, which more closely approximates real-world variability. In this setting, false negative errors are identified as the dominant risk factor, particularly for the“Signal for Help”class. Among all evaluated models, YOLOv11m trained for one hundred epochs demonstrates the most balanced and reliable performance. It achieves the lowest false negative counts on the custom dataset, with one hundred thirty missed detections for the“Signal for Help”class and one hundred ninety-seven for the“Need Help”class, while maintaining stable precision and recall values. This balanced error profile positions YOLOv11m as the most suitable model for safety-critical emergency hand gesture detection. By contrast, YOLOv12m trained for one hundred epochs exhibits a conservative detection behavior characterized by zero false positives but substantially higher miss rates, particularly for the“Signal for Help”gesture, which limits its suitability for high-risk applications. In conclusion, this thesis demonstrates that high overall accuracy alone is insufficient for evaluating emergency hand gesture detection systems. Instead, the distribution of error types, model generalizability across datasets, and system-level decision strategies play decisive roles in determining real-world reliability. By emphasizing false-negative-oriented evaluation, cross-dataset validation, and temporal decision logic, this work contributes a safety-centric perspective to the literature on hand gesture recognition. The findings provide practical guidance for the design of real-time, safety-critical visual detection systems and establish a robust foundation for future research on automated emergency response technologies and intelligent surveillance infrastructures. Moreover, the integration of temporal decision logic with deep learning–based detection models demonstrates that system-level reasoning mechanisms are as crucial as model accuracy itself in ensuring reliable emergency detection performance.

Benzer Tezler

  1. Dokuz Eylül Üniversitesi Tıp Fakültesi hastanesinde mavi kod uygulamalarının değerlendirilmesi

    Evaluation of blue code applications in dokuz Eylul University Medical Faculty Hospital

    NESLİHAN TÜTÜNCÜ KILIÇ

    Tıpta Uzmanlık

    Türkçe

    Türkçe

    2021

    Anestezi ve ReanimasyonDokuz Eylül Üniversitesi

    Anesteziyoloji ve Reanimasyon Ana Bilim Dalı

    PROF. DR. BAHAR KUVAKİ BALKAN

    DOÇ. DR. ŞULE ÖZBİLGİN

  2. Ticaret gemilerinde küresel tehlike haberleşme sistemleri G.M.D.S.S.

    Global distress communication systems on merchant vessels

    NUMAN ÇOKGÖRMÜŞLER

    Yüksek Lisans

    Türkçe

    Türkçe

    1997

    İletişim BilimleriDokuz Eylül Üniversitesi

    Deniz İşletmeleri Yönetimi Ana Bilim Dalı

    PROF. DR. Ö. BAYBARS TEK

  3. Kablosuz algılayıcının modül tasarımı

    Design of a wireless sensor module

    REGAİP MUTLU BİÇER

    Yüksek Lisans

    Türkçe

    Türkçe

    2008

    Elektrik ve Elektronik MühendisliğiKocaeli Üniversitesi

    Elektronik ve Haberleşme Mühendisliği Ana Bilim Dalı

    DOÇ. DR. ADNAN KAVAK

  4. Hacettepe Üniversitesi Sıhhiye Yerleşkesinde kardiyopulmoner arreste yönelik oluşturulan mavi kod uygulamasının süreç ve sonuçlarının değerlendirilmesi

    Evaluation of the process and outcome of blue code procedure developed for cardiopulmonary arrest in Hacettepe University Sihhiye campus

    ARZU TOPELİ İSKİT

    Yüksek Lisans

    Türkçe

    Türkçe

    2016

    HastanelerHacettepe Üniversitesi

    Epidemiyoloji Ana Bilim Dalı

    PROF. DR. BANU ÇAKIR

  5. Acil servis hekimlerinin radyasyondan korunma farkındalığı

    Radiation protection awareness of emergency physicians

    BAŞAK YILMAZ

    Tıpta Uzmanlık

    Türkçe

    Türkçe

    2018

    İlk ve Acil YardımUfuk Üniversitesi

    Acil Tıp Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ TOGAY EVRİN

    DR. ÖĞR. ÜYESİ EBRU ÖZAN SANHAL