Yayın: Balancing mixed adversarial perturbations: Towards robust reinforcement learning
Yükleniyor...
Dosyalar
Tarih
Yazarlar
Danışman
Bölüm / Program
Mechatronics Engineering
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
ITU Graduate School
Türü
Özet
Deep Reinforcement Learning (DRL) has proven itself as an effective method to train autonomous agents which learn sequential decision-making through complex environment interactions. The combination of deep neural networks with reinforcement learning decision-making framework in DRL has led to outstanding achievements in game playing, robotics and autonomous systems. The agents discover their best policies through reward maximization which allows them to operate in complex state and action spaces that were previously impossible to handle. The deployment of DRL systems in safety-critical real-world applications requires researchers to focus on developing methods which protect these systems from adversarial attacks and unpredictable environmental changes. The impressive capabilities of DRL agents do not protect them from adversarial perturbations which are small input modifications that result in complete breakdowns of decision-making systems. The research community has thoroughly investigated these system weaknesses in supervised learning before expanding their analysis to reinforcement learning. The methods of adversarial training, domain randomization, contextual RL and robust MDPs serve as solutions to protect DRL systems from vulnerabilities. Conversely, other methods that strive to achieve strong and generalizable policies have certain limitations compared to adversarial training. Domain randomization needs human intervention to determine environmental parameter ranges which might fail to detect essential boundary situations and faces challenges when handling complex high-dimensional spaces while requiring expensive computational resources. On the other hand, agents trained with adversarial training detect policy weaknesses through automatic challenge generation without needing pre-defined context information. Contextual RL depends on well-defined, observable context variables, limiting its ability to generalize to unseen or ambiguous contexts, whereas adversarial training actively probes policy weaknesses without needing predefined contexts. Robust RL, while designed for worst-case scenarios, often produces overly conservative policies that underperform in typical conditions and relies on predefined uncertainty sets that may not capture real-world variability. Adversarial training, by contrast, dynamically exposes vulnerabilities through optimized adversarial examples, though it risks overfitting to specific attacks if not carefully balanced. Modern adversarial training methods used to improve adversarial robustness include feeding adversarial inputs to agents in the training process. These approaches focus on changes in one area-whether in state observations, actions, rewards, or environmental processes. The single-type defense approaches prove successful for their designated domains yet they do not handle the multiple perturbation sources which appear in actual environments. The primary limitation of current adversarial training methods is their narrow focus on single-modal perturbations. Real-world systems, such as autonomous drones or industrial robots face multiple disturbances that affect their perception systems through sensor noise and their control systems through actuation errors. Agents who learn to defend against observation attacks become exposed to action-space disruptions and agents who defend against action-space attacks become exposed to observation attacks. The inability of models to generalize between different types of perturbations creates an illusion of security. The existing restrictions in robustness transfer between different types of attacks under mixed adversarial conditions remain unexplained. This thesis addresses the aforementioned limitations by developing a complete system to train RL agents which protects them against state and action space mixed adversarial attacks. The main theoretical advancement brings forth the Action and State-Adversarial Markov Decision Process (ASA-MDP) which serves as a generalized game-theoretic framework for robust MDP models. On the implementation side, the proposed ASA-PPO algorithm serves as the implementation solution which uses a multi-head adversarial policy structure to achieve balanced and adaptive attack methods. The developed methods create an enhanced adversarial training system which trains agents to handle sophisticated threats that appear through multiple domains. The proposed ASA-PPO method undergoes evaluation through multiple continuous control tasks which run in the MuJoCo physics engine, including PendulumSwingup, FingerSpin, and CheetahRun. The optimization process for training the adversary and protagonist agents uses an alternating scheme based on Proximal Policy Optimization (PPO) algorithm. The adversary uses a multi-head neural network structure to produce state and action disturbances which share common features to represent their relationship patterns. The system uses fixed perturbation limits which stay within realistic boundaries. The training process is computationally intensive but yields agents that demonstrate strong robustness across a variety of adversarial scenarios. The ASA-PPO provides multiple essential benefits which surpass current methods. The framework provides complete protection against mixed adversarial attacks through its training process which generates a detailed robustness assessment. The multi-head adversarial architecture distributes perturbation budgets evenly between state and action spaces to stop the adversary from specializing in one particular attack type. The agents trained through ASA-PPO show better performance when they encounter new perturbation patterns and intensity levels. The method achieves high performance in normal operating conditions which proves that robustness systems do not compromise standard task execution. Traditional adversarial training methods are limited by their single-type perturbation focus, which fails to reflect the complexity of real-world deployment conditions. These methods often result in overfitting to specific attack patterns, leaving agents vulnerable to novel or hybrid threats. Additionally, naive combinations of different attack types (e.g., simply adding state and action perturbations) lead to imbalanced training regimes, where one perturbation type dominates and the agent remains weak against the other. These shortcomings underscore the need for more sophisticated, adaptive adversarial training frameworks like ASA-PPO. The experimental findings show that agents trained through ASA-PPO achieve better performance than baseline methods when facing different types of adversarial attacks. The results show that training agents against one type of perturbation (action-space) does not provide protection against the other type (observation-space) which demonstrates the need for hybrid training approaches. The ASA-PPO agents produce superior returns during both single-type and hybrid attacks and they perform better when encountering new perturbation levels. The research shows that adversarial resistance does not guarantee resistance against changes made to environmental parameters through mass or damping coefficient modifications. The results indicate that adversarial robustness and parametric robustness present separate challenges which need distinct solution approaches. This thesis establishes the groundwork for a new generation of robust RL systems capable of withstanding realistic, multi-modal adversarial threats. Future research could investigate adaptive perturbation balancing, extensions to partially observable environments (POMDPs) and integration with domain randomization to handle both adversarial and parametric uncertainties. The ASA-MDP framework also opens the door to theoretical analyses of equilibrium properties in mixed-adversarial settings. Ultimately, this work contributes to the broader goal of deploying safe, reliable, and trustworthy autonomous systems in complex, unpredictable environments.
Tanım
Thesis (Ph.D.) -- Istanbul Technical University, Graduate School, 2025
Dergi veya Seri
ISSN
ISBN
Haklar
Anahtar Kelimeler
mekatronik mühendisliği, mechatronics engineering, makine öğrenmesi, makine öğrenmesi, pekiştirmeli öğrenme, reinforcement learning
Alıntı
Koleksiyonlar
Onay
Gözden geçir
Tamamlayıcı Bilgiler
Referans Gösteren
5
Görüntülenme
93
İndirme
Google Scholar
Scholar'da Ara ↗ Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.