Graph-aware fraud detection in review platforms via layered feature engineering
| dc.contributor.advisor | Bahtiyar, Şerif | |
| dc.contributor.author | Tutuk, Mehmet Ali Han | |
| dc.contributor.authorID | 504221519 | |
| dc.contributor.department | Computer Engineering | |
| dc.date.accessioned | 2026-09-29T11:21:09Z | |
| dc.date.issued | 2026-06-09 | |
| dc.description | Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2026 | |
| dc.description.abstract | Online review platforms shape how users choose products, and the same review and rating signals feed the recommender systems that build on top of them. That same influence makes review platforms a target. Attackers write fake reviews, pay for positive ones, and coordinate rating activity to push specific products up or down; this pulls user decisions and the recommenders trained on these signals in the same direction. This thesis treats the detection of such reviews as a learning problem on review graphs. The difficulty is mostly structural. A fraudulent review almost never appears alone. Fraud accounts target the same products, post within narrow time windows, and produce similar rating patterns, and some imitate legitimate reviewer behavior closely enough to mask single-review evidence. These patterns are only visible when reviews are compared against one another. A natural representation for this kind of signal is a graph that links reviews, users, and products through shared targets, ratings, and timestamps. A review's position in this graph encodes information that no review-level feature can express on its own. Most recent work on this problem uses graph neural networks that operate on the review graph directly. Designs based on graph convolution, attention, and relation-specific message passing have produced consistent improvements on the standard review-fraud benchmarks. There is a cost to these improvements. End-to-end GNN training scales poorly in both nodes and relations, and the techniques used to reduce training cost discard part of the same relational signal that motivated the use of graphs. The trade-off becomes acute on the larger review datasets, where running the published GNN baselines under their original protocol is no longer practical at full scale. A different design is followed in this thesis. The relational signal is extracted from the graph and added to the feature space of a tabular classifier, instead of being processed inside an end-to-end graph model. The feature space is built in three levels: a base level that encodes review, text, and user attributes; a neighborhood level that aggregates the same attributes over local graph neighborhoods defined by the relation layers available in each dataset; and a full level that further appends relation-wise edge counts. Classification is then performed by gradient-boosted tree learners (XGBoost, LightGBM) over the resulting feature matrix. The three levels are designed for layer-by-layer comparison. Two graph-learning benchmark datasets, YelpChi (45,954 review nodes) and Amazon (11,944 user nodes), are used together with two larger datasets constructed from raw Yelp metadata, YelpZip (608,598 review nodes) and YelpNYC (359,052 review nodes). The two constructed datasets do not come with a graph; their relational structure is built in this thesis from the user, target, timestamp, and rating fields of the raw review records. A fixed 40/20/40 train/validation/test split and a 16-regime preprocessing sweep are applied uniformly across all four datasets, so the contribution of the layered representation can be read separately from the contribution of preprocessing. On the benchmarks, the Full layer reaches F1-macro 0.8989 on YelpChi and 0.9347 on Amazon, against the SplitGNN reference of 0.7403 and 0.7352 on the same metric, with consistent gains on AUC and GMean as well. Recall on YelpChi is 0.8016, below the 0.8256–0.8577 range reported by BWGNN, H2-FDetector, and SplitGNN; the corresponding false-alarm rate on the genuine class is 2.36%. The layered improvement is not driven by a single favorable configuration: the Base < Nbr < Full F1 ordering holds in 160/160 rows on YelpChi and 111/160 on Amazon, and Wilcoxon signed-rank tests confirm the same direction at p ≤ 0.010 across all six dataset-transition pairs in the primary 10-seed test. The location of the gain differs across datasets. On Amazon, almost all of the improvement comes from the neighborhood layer; on YelpChi, the structural layer adds a further Recall increase above the neighborhood step. The constructed datasets split this overall pattern by structural form. Eight attack-aware proxy subsets are defined on top of the constructed graphs, derived from structural attack patterns in the review-fraud literature: coordination-driven families on shared-target and burst activity, connectivity- and density-driven families on repeated-user relations and locally dense fraud blocks, and a camouflage family on fraud nodes embedded in genuine-dominant neighborhoods. The Base → Full F1 gain reaches several percentage points on the coordination-driven families, stays smaller but consistent on the connectivity- and density-driven ones, and is weakest on the camouflage regime, where absolute scores stay well below the rest of the families. The RSR hybrid family shows a Base → Nbr step that does not clear a one-sided Wilcoxon test on either constructed dataset, while still benefiting from the structural layer. The dominant layer therefore changes with the structural form of the underlying fraud, and the layered hierarchy makes that dependence readable at the metric level. Some limitations follow from the scope of the thesis. Detection metrics are reported, but a matched-hardware runtime and memory benchmark against the GNN baselines is not, so the lightweight side of the framework is supported by architectural choice and not yet by direct measurement. GNN comparisons on the constructed graphs are also outside the scope of the thesis, since there are no published GNN results on these constructed datasets and re-implementing existing GNN methods at this scale was not part of the present work. The four datasets used here also represent a single temporal snapshot, while real review platforms produce continuously arriving reviews. A fourth limitation concerns the constructed-side analysis: the eight proxy subsets are evaluated under a single seed with fixed predicate thresholds, fixed subset-size caps, and the dataset-native fraud-class proportion, so the family-level numbers should be read as a structural diagnosis rather than as precise comparative measurements. Four follow-up studies grow naturally out of these limitations: a matched-hardware runtime and memory benchmark against the GNN baselines under identical splits, running those same GNN baselines on the constructed graphs, a broader subset-level ablation along seed, threshold, cap, and class-balance axes, and a streaming or dynamic-graph extension in which the relation views are updated as new reviews arrive. | |
| dc.description.degree | M.Sc. | |
| dc.identifier.uri | https://hdl.handle.net/11527/81244 | |
| dc.language.iso | en | |
| dc.publisher | Graduate School | |
| dc.sdg.type | none | |
| dc.subject | Review Fraud Detection | |
| dc.subject | Yorum Sahtekarlığı Tespiti | |
| dc.subject | Graph Neural Networks | |
| dc.subject | Grafik Sinir Ağları | |
| dc.subject | Scalable Machine Learning | |
| dc.subject | Ölçeklenebilir Makine Öğrenmesi | |
| dc.title | Graph-aware fraud detection in review platforms via layered feature engineering | |
| dc.title.alternative | Katmanlı graf-tabanlı özellik türetme ile inceleme platformlarında sahtekarlık tespiti | |
| dc.type | Master Thesis |