Anticipating robot manipulation failures using knowledge distillation

Yükleniyor...
Küçük Resim

Bölüm / Program

Computer Engineering

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Graduate School

Özet

In this thesis, a novel framework is proposed to anticipate robot-object manipulation failures before their occurrance by using knowledge distillation and transformer-based multimodal learning. The motivation behind this study is to move beyond traditional failure detection methods, which identify errors after they happen, and instead focus on predicting them in advance to ensure robot safety and autonomy in real-world environments. The framework adopts a teacher-student architecture. The teacher model is trained on complete video sequences on manipulation task execution, including the failure moment, while the student model observes only the initial part of the sequence, without seeing the failure itself. The knowledge gained by the teacher is distilled into the student, enabling it to anticipate failures with limited temporal input. Both teacher and student models share the same architecture, which is based on ViViT (Video Vision Transformer) [2]. To enhance temporal reasoning, the framework incorporates RGB, depth, and optical flow modalities. Each modality is processed through a separate transformer stream, and its features are fused using late fusion. The optical flow is extracted from RGB frames using FlowNet2 [3], and added to the FAILURE dataset, which consists of 324 different manipulation episodes recorded with the Baxter robot. Tasks include pouring, pushing, placing, and stacking. The student model is trained with a hybrid loss that includes cross-entropy for classification and Maximum Mean Discrepancy (MMD) for distillation. The performance of different loss strategies (MMD, MSE, Jaccard) and fusion methods is analysed. The results show that the student model, when trained with RGB-D-F modalities and MMD-based modality-wise distillation, achieves an F1-score of 82.12% for 1-second anticipation and 79.79% for 2-second anticipation. These scores outperform unimodal setups and non-distilled baselines. In addition to offline experiments, the system is evaluated under real-time conditions using a sliding window over streaming video frames. The RGB-D-F student model performs consistently during pouring manipulation execution and achieves a 97.8% F1-score in real-time failure anticipation, and the duration of two successive predictions is 0.11 seconds. To measure the model's robustness, generalization experiments are conducted on manipulation actions that were excluded from training. The student model achieves an F1-score of 73.4% on unseen actions, confirming that the proposed system can generalize learned failure patterns to novel manipulation scenarios. In conclusion, this thesis presents a transformer-based multimodal anticipation system that can predict manipulation failures early, operate in real time, and generalize across tasks, contributing to safer and more adaptive robotic systems.

Tanım

Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2025

Dergi veya Seri

ISSN

ISBN

Haklar

Anahtar Kelimeler

robotic systems, robotik sistemler, learning models, öğrenme modelleri, robots, robotlar, autonomous robots, otonom robotlar

Alıntı

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

0

Views

0

Downloads