Adversarial attacks against machine learning algorithms at training stage

Yükleniyor...
Küçük Resim

Tarih

Bölüm / Program

Computer Engineering Programme

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Graduate School

Özet

Learning from data and processing them is one of the most important achievements that lead to improving ourselves in the past decade. We can apply that learned information to various domains like detecting interested objects via recorded video or live footage, deciding if a message is spam or not, translating paragraphs into other languages, auto-completing texts from basic words to complex sentences, moving agents in real-time processes, and so many others. With the increase in our process capability and improvements in hardware-related resources, we began to use machine learning algorithms to learn from data. We saw that we can improve our decision capability by applying multi-layer perceptrons and after that (deep) neural networks. Today, many researchers still develop new ways or improve existing ones to increase the state of the arts further. To accomplish a successful machine learning model, collecting huge and variant data samples and using a well-designed network are crucial steps. But, data collection can be a harsh process and relies on various third parties. Also, using others' networks can lead to unexpected model behaviors on specific inputs which are actually encoded into the inner layers of the network. Machine learning (or deep neural networks) networks are used in real-world systems to accomplish high-performance solutions for various problems like those mentioned above. But, it is known that networks are carried off their training side parameter and training data. From this characteristic, it is possible to extract similar data to the ones that the model trained on by the similarity to its distribution. Supplying the security of a system helps it to sustain its reliability and sustainability. However, reasons like those mentioned above can introduce flaws in systems. Attackers can use these flaws to decrease system performance and introduce vulnerabilities into them with help of those networks. During this work, we focused on three important aspects of the security of machine learning algorithms. First of all, we summarize adversarial attacks against machine learning and categorize them using the attacker and target model-based criteria. In many works adversarial attacks and their countermeasures are investigated, but the main difference of our is we focus on attacks and performance from both the attacker and target system sides. While exploring the design and performance of attacks, we analyze the response of the target model against them and the results that the attack achieves on them. Second, we investigate the robustness of classical machine learning algorithms against adversarial attacks. For this specific task, six machine-learning algorithms and 4 different datasets from different domains are used. Specifically, Support Vector Machine, Stochastic Gradient Descend, Logistic Regression, Random Forest, Gaussian Naive Bayes, and K-Nearest Neighbor algorithms are selected to create learning models and spam, botnet, malware, and breast cancer datasets are selected for training data samples. The label-flipping attack strategy, which is a data poisoning attack, is used to apply adversarial attacks on the targeted models and evaluate the robustness of those models. For evaluation, the change rates in their accuracy rate, f1-score, and AUC score are analyzed. For determining the algorithms with the highest performance, each algorithm's performance is compared with other algorithms in each learning environment. Last but not least, we further analyze the performance of models in a more complex attack scenario, more specifically backdoor attacks. Backdoor attacks are a specific attack strategy that happens in the training stage. Attackers aim to let the target model learn from a specific trigger pattern as the target class. In this work, a clean-label backdoor attack strategy is designed to generate poisoned data samples correctly labeled as they must be. An LSTM-based spam detector is trained for the main target model. Results show that the attack can successfully deceive the performance of the target model. While improvements in the performance of machine learning models still continue in various domains, the security aspect of machine learning is still a new research area. In this work, we focus on the flaws of machine learning networks by analyzing their performance change when they interact with adversaries in their training. Currently, we are still far from even a standard in the robustness evaluation, but with further research, we can accomplish great achievements in the stability of machine learning.

Tanım

Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2023

Dergi veya Seri

ISSN

ISBN

Haklar

Anahtar Kelimeler

Cyber security, Machine learning, Algorithms

Alıntı

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

8

Views

11

Downloads