Unveiling the performance of pre-processing approaches in machine learning based flood susceptibility mapping
Yükleniyor...
Dosyalar
Tarih
Yazarlar
Bölüm / Program
Hydraulics and Water Resources Engineering
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Graduate School
Türü
Özet
Floods represent one of the most catastrophic natural disasters, distinguished by their capacity to inflict substantial loss of life, extensive property damage, and considerable economic difficulties. The severity of these events is often intensified by a confluence of factors, including heavy rainfall, increasing population densities, rapid urbanization, and the overarching climate change effects. As urban areas expand and more individuals settle in flood-prone regions, the associated risks of flooding escalate, necessitating the development of effective strategies for flood risk assessment and management. In recent years, the incorporation of machine learning methodologies has emerged as a promising approach to enhance the precision and reliability of flood susceptibility models, maps, and early warning systems. This study concentrates on the San Joaquin River basin in California, a region that has faced significant flooding challenges. The primary aim of this research is to explore various pre-processing techniques that can be utilized to effectively assess flood susceptibility, employing the eXtreme Gradient Boosting (XGBoost) algorithm as the predictive framework. To accomplish this, the study identifies 22 critical flood conditioning factors relevant to the San Joaquin River basin. These factors encompass a diverse array of environmental and anthropogenic variables that influence flood risk, including land use, soil type, topography, and hydrological characteristics. The methodology employed in this research involves a comprehensive two-stage pre-processing analysis, which examines 18 distinct scenarios to ascertain the most effective approach for modeling flood susceptibility. The research findings indicate that the XGBoost model, when applied with robust scaling techniques and a 70/30 train-test split, achieved optimal performance, attaining an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.851. This metric reflects a high degree of accuracy in predicting flood susceptibility. Additionally, the study revealed that utilizing a 10x class imbalance ratio with random under sampling (RUS) during the training phase yielded the most precise results in the testing phases, with an AUROC of 0.835. The flood susceptibility maps produced from this analysis indicate that over 20 percent of the San Joaquin River basin is classified as being at high to very high risk of flooding. This critical information is vital for local authorities and stakeholders, as it underscores areas that necessitate immediate attention and intervention to mitigate potential flood impacts. Furthermore, the research employed SHapley Additive exPlanation (SHAP) values to interpret the model's predictions and identify the most significant factors contributing to flood susceptibility. The analysis highlighted the substantial influence of alluvial presence, proximity to geological faults, and transportation infrastructure. Collectively, these findings will enhance the existing literature on flood susceptibility mapping and inform the necessary precautions to be undertaken prior to the occurrence of flood events in the region.
Tanım
Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2025
Dergi veya Seri
ISSN
ISBN
Haklar
Anahtar Kelimeler
machine learning, makine öğrenmesi, flood, sel, mapping, haritalama