Chapter Four · failure evidence

What Support Vector Machines got wrong, from 65 dissertations

The records evaluate Support Vector Machines across diverse scientific and engineering applications, highlighting recurring difficulties in real-world deployment. Practitioners frequently find that Support Vector Machines suffer from severe computational bottlenecks, sensitivity to class imbalance, and underperformance relative to both simpler statistical baselines and advanced ensemble or deep learning methods. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Support vector machines underperform tree ensembles and neural networks on complex data

14 theses · 12 institutions

Support vector models achieve lower accuracy and higher prediction errors than Random Forests, gradient boosting, and deep neural networks on tabular, image, and temporal datasets. These alternatives handle high feature dimensionality, complex sequence dynamics, and noise significantly better than support vector approaches.

Tried and failed

support vector machine classification applied to tabular delay classification. Outcome: worse than baseline. Reason: non-parametric model struggled compared to parametric deep neural network architectures

Novel deep learning approaches for flight delays prediction · Cranfield

Tried and failed

Support vector machine classification applied to metabolomics-based disease risk prediction. Outcome: worse than baseline. Reason: SVM had substantially lower predictive AUC compared to tree-based ensemble models

Eating Behaviors in Adolescence and Young Adulthood and Adult Cardiometabolic Disease Risk · Harvard

Lost to a baseline

Support vector machines (SVM) and logistic regression were outperformed by Random Forest when trained on human protein target abundances to separate case and control patients across metagenomic disease cohorts

PREDICTION OF BACTERIAL HORIZONTAL GENE TRANSFER AND DISEASE-RELEVANT HOST-MICROBIOME PROTEIN- PROTEIN INTERACTIONS · Cornell

Considered and rejected

Considered and rejected: Rejected Support Vector Machines and Logistic Regression in favor of Random Forest for binary visibility classification

Development of a machine vision system to estimate the physical attributes of potato tubers on-the-go at the post-harvest stage · DalSpace

Lost to a baseline

Support Vector Machines (SVMs) underperformed on training cross-validation (accuracy 0.678) compared to Random Forests (0.932) and Naive Bayes (0.792).

Justification Mining: Developing a novel machine learning method for identifying representative sentences and summarising sentiment in financial text · Oxford

Lost to a baseline

Support Vector Classifier (65.77% raw, 55.28% summary/Fourier/Wavelet) lost to untuned Random Forest (80.28% raw, 76.91% summary, 79.63% Fourier, 79.80% Wavelet) across all malware dataset representations.

Malicious Network Traffic Detection via Deep Learning: An Information Theoretic View · JScholarship

Considered and rejected

Considered and rejected: Support vector machines (SVM) with Generative Adversarial Active Learning (GAAL) were rejected because SVM performs poorly on complex image classification tasks.

Engineering-Driven Learning Approaches for Bio-Manufacturing and Personalized Medicine · Georgia Tech

Lost to a baseline

Support Vector Machine (SVM) with PCA feature reduction achieved substantially lower intent classification accuracy across all time steps compared to the Long Short-Term Memory (LSTM) deep learning network.

Intent sensing for assistive technology · Oxford

Tried and failed

Support Vector Regression applied to tabular spatial event regression. Outcome: worse than baseline. Reason: yielded high inference execution times and worse RMSE metrics than ensemble and neural baselines

Human sensor networks for natural disasters · Imperial

Tried and failed

Linear support vector classification applied to time-series image encodings. Outcome: worse than baseline. Reason: Failed to capture non-linear relationships across multi-channel image representations compared to convolutional models

Optimizing High-Performance Computing Systems: Insights from System Monitoring, Workload Management, and Scheduling Strategies · Texas Tech

Considered and rejected

Considered and rejected: Traditional machine learning models (Support Vector Regression and Random Forests) were rejected as the primary architecture for short-term arrival traffic forecasting in favor of recurrent neural networks (LSTMs) due to their inability to directly capture multi-scale sequential dependencies without manual lag feature engineering.

Predictive Modeling of Aircraft Arrival Times in the Terminal Maneuvering Area Through Data-Driven Techniques · Georgia Tech

Tried and failed

Support vector machines with high-dimensional descriptors applied to molecular formulation and glass forming classification. Outcome: worse than baseline. Reason: High feature dimensionality and noise sensitivity degraded classification performance relative to tree ensembles

Overcoming challenges of solid dosage formulation development by using emerging technologies · UT Austin

Lost to a baseline

Support Vector Machine (SVM), Random Forests, and Gaussian Processes achieved lower test dataset R2 compared to the 2-hidden-layer ANN model for predicting HFRC tensile stress-strain behavior.

Condition Assessment of Civil Infrastructure and Materials Using Deep Learning · Virginia Tech

Lost to a baseline

Partial least-squares regression (test R² = 0.81) and support vector regression (test R² = 0.86) were beaten by Multilayer Perceptron Regression (test R² = 0.96) when predicting 1-cm SOC from 10-cm trained hyperspectral models.

High-resolution & high-throughput methods for soil carbon research · Research Repository UCD

Class imbalance causes severe prediction bias toward the majority class

10 theses · 10 institutions

When trained on imbalanced datasets, support vector classifiers bias predictions heavily toward majority categories and yield low recall, low sensitivity, or zero true positives for minority classes. In several evaluations, models fail to exceed random baseline accuracy or predict all instances as negative.

Tried and failed

support vector machine with log-transformed count data applied to high-dimensional phenotype classification. Reason: severely biased predictions toward negative class, resulting in very low specificity

Gut Microbiome Composition and Attention Deficit Hyperactivity Disorder · Harvard

Tried and failed

support vector machine for binary classification applied to imbalanced time-series trace classification. Outcome: worse than baseline. Reason: failed to handle class imbalance compared to boosting methods with random undersampling

Hardware-level Vulnerabilities and Support for Secure and Safe Cyber-Physical Systems · Cornell

Tried and failed

support vector machine with metaheuristic feature selection applied to imbalanced genomic biomarker classification. Outcome: did not generalise. Reason: severe class imbalance degraded classification performance on minority classes

UNCERTAINTY MITIGATION IN IMAGE-BASED MACHINE LEARNING MODELS FOR PRECISION MEDICINE · Georgia Tech

Lost to a baseline

Logistic Regression ML model achieved a higher true positive rate (61.11%) than Support Vector Machine (27.78%) and Random Forest (33.33%) on the ALK mutation dataset.

Protein Dynamics in Mutated Kinases · Penn

Lost to a baseline

Support Vector Machines with RBF kernels failed to exceed the random baseline (87.1%) on the imbalanced healthy vs. gait disorder (N/GD) binary classification task (-0.8% and -3.8% below baseline for z-score and min-max normalization, respectively).

Human gait analysis : machine learning-based classification of gait disorders · DSpace-CRIS at TU Wien

Considered and rejected

Considered and rejected: Rejected Support Vector Machines and K-Nearest Neighbor algorithms because they overfit to the majority upland class and severely underestimated peatland extent

The Classification and Characterization of Canadian Boreal Peatland Sub-classes · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Rejected Support Vector Machine (SVM) for Stage-1 due to poor sensitivity (70.30%), lack of interpretability, and feature scaling requirements

Decentralized Processing Techniques for Biomedical Signal Classification · Research Repository UCD

Tried and failed

support vector machine on minimal feature subsets applied to multi-class mechanical fault classification. Outcome: did not generalise. Reason: insufficient feature representation caused severe class confusion and complete false negative rate on high-severity faults

PROGNOSTICS AND HEALTH MANAGEMENT OF WIND TURBINE GEARBOXES: DETECTION AND DIAGNOSTICS USING MACHINE LEARNING · WTAMU Repository

Tried and failed

multiple instance learning support vector machine applied to instance-level credibility prediction. Outcome: worse than baseline. Reason: predicted all instances as negative, yielding zero precision and recall

Anomalous Information Detection in Social Media · Virginia Tech

Tried and failed

Support Vector Classifier for event detection applied to traffic delay occurrence classification. Outcome: worse than baseline. Reason: SVC suffered from very low recall and detection rates for positive instances

A comprehensive analysis framework for addressing non-recurrent traffic congestion: Focusing on incident-induced delays employing connected vehicle data in transportation systems m · Iowa State

Support vector machines are beaten by simpler linear and statistical baselines

9 theses · 8 institutions

Standard logistic regression, multilinear regression, and discriminant analysis frequently outperform support vector machines on tabular and biometric prediction tasks. Adding kernel complexity or interaction terms increases prediction error and overfitting risks without delivering accuracy improvements over simpler linear models.

Tried and failed

support vector machines instead of logistic regression applied to human task performance prediction. Outcome: worse than baseline. Reason: limited sample size led to risk of overfitting without performance gains

Measuring and predicting detection performance on security images as a function of image quality · UT Austin

Lost to a baseline

Support Vector Machine (mean EER 0.1001) lost to baseline Logistic Regression (mean EER 0.0739)

Behavioral authentication in virtual reality environments · Texas Tech

Lost to a baseline

Support Vector Machine (SVM) and machine learning models were noted in cited literature (Beutel et al., 2019) to be outperformed by standard logit models in systemic crisis early-warning systems.

Išvestinių finansinių priemonių poveikio Euro zonos šalių sisteminei rizikai vertinimas Evaluation of Financial Derivatives Impact on Systemic Risk of Euro Area Countries · Mykolas Romeris University / Mykolo Romerio universitetas

Lost to a baseline

Support Vector Machine (SVM, 74.40% emergence, 84.34% stem elongation) and Random Forest (70.19% emergence, 82.48% stem elongation) were beaten by Logistic Regression (74.59% emergence, 84.40% stem elongation) on central RGB features.

Spatial Data Analysis in Digital Agriculture Application on Crops Growth Monitoring · Research Repository UCD

Lost to a baseline

Support Vector Machine (SVM) on original ORL data achieved 83.9% accuracy, which was outperformed by Quadratic Discriminant (QD) on original ORL data (91.5%).

Higher order tensor decompositions: from intuition to implementation and application · Imperial

Lost to a baseline

Support vector machine predicting gene expression from regulatory components achieved R2 = 0.073, performing worse than multiple linear regression (R2 = 0.08985).

Transcript cleavage and polyadenylation in plants · ResearchWorks

Considered and rejected

Considered and rejected: Rejected support vector machines for dictionary feature selection after cross-validation showed multinomial lasso performed better.

Value Representation · Harvard

Considered and rejected

Considered and rejected: Rejected the SVM (Support Vector Machine) model for the empirical evaluation of derivatives' impact on systemic risk because its parameters are difficult to interpret, it requires constraints/assumptions similar to simulation models, and it is less effective than logit regression.

Išvestinių finansinių priemonių poveikio Euro zonos šalių sisteminei rizikai vertinimas Evaluation of Financial Derivatives Impact on Systemic Risk of Euro Area Countries · Mykolas Romeris University / Mykolo Romerio universitetas

Tried and failed

support vector regression with polynomial feature augmentation applied to tabular prediction with non-linear features. Outcome: worse than baseline. Reason: adding polynomial and interaction terms increased prediction error compared to simple linear baselines

Unifying Strategic Military Force Design and Operational Warfighting: A Stochastic Game Approach · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Ridge, Lasso, Support Vector Regression (SVR), and Logistic Regression for area-volume curve fitting in favor of multilinear regression to preserve simplicity, interpretability, and physical consistency.

From Sources to Sinks: Advancing Surface Water Management Through Satellite Remote Sensing · ResearchWorks

Inappropriate kernel selection prevents models from capturing boundary structures

10 theses · 8 institutions

Linear kernels fail to capture nonlinear interactions and complex boundaries, while sigmoid kernels generate errors an order of magnitude higher than alternative baselines. Using high-degree polynomial kernels or unaligned nonlinear mappings leads to high cross-validation variance and severe overfitting.

Tried and failed

support vector machine with sigmoid kernel applied to spectroscopic regression and classification. Outcome: did not generalise. Reason: poor predictive performance and failure to validate on microbial count and classification tasks

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Tried and failed

linear support vector machines applied to tabular loan default classification. Outcome: worse than baseline. Reason: linear decision boundaries failed to capture non-linear feature interactions across retail banking datasets

Modelling the probability of household default for conventional and Islamic banking. · Cranfield

Tried and failed

Support Vector Regression with sigmoid kernel applied to stride length estimation from inertial sensors. Outcome: worse than baseline. Reason: The sigmoid kernel produced errors an order of magnitude higher than other regression baselines.

Wearable Devices: A Tool for the Malicious Reconnaissance of Private Spaces · Texas Tech

Tried and failed

support vector regression surrogate modeling applied to reactor multiphysics simulation dataset. Outcome: worse than baseline. Reason: systematically underperformed and unable to accurately capture complex variations in the dataset

Development of Coupled Machine Learning and Optimization Framework for Comprehensive Molten Salt Reactor Design · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Support Vector Machines and logistic regression for frequency security boundary extraction due to poor accuracy from simple hyperplanes on nonlinear high-dimensional boundaries.

Data-driven stability-constrained optimisation for software-defined power systems with high IBR penetration · Imperial

Tried and failed

support vector regression with linear kernel applied to photovoltaic power generation forecasting. Outcome: worse than baseline. Reason: linear kernel cannot capture nonlinear relationships in the data

A Coordinated Voltage Management Method Utilizing Battery Energy Storage Systems and Smart PV Inverters in Distribution Networks with High PV and Wind Penetrations · Virginia Tech

Tried and failed

non-linear kernel support vector machines applied to driver maneuver classification. Outcome: worse than baseline. Reason: None

Motion control considering human driver characteristics for driving safety enhancement of connected and automated vehicles · UT Austin

Considered and rejected

Considered and rejected: Rejected nonlinear kernels for Support Vector Machines (SVMs) in favor of a linear kernel to avoid overfitting models and poor performance.

Resolution and Reliability in Functional Connectivity Analysis · JScholarship

Tried and failed

support vector machine with cubic kernel applied to Raman spectral classification. Outcome: overfit. Reason: High-degree polynomial decision boundary led to severe overfitting and high cross-validation variance on spectral data

Nanobiotechnology Enabled Environmental Sensing of Water and Wastewater · Virginia Tech

Tried and failed

support vector regression applied to optical sensor pressure estimation. Outcome: worse than baseline. Reason: struggled to model complex non-linear mappings from unaligned multi-trace sensor interactions

PHYSICAL INTELLIGENCE ENABLED BY SOFT ROBOTIC SYSTEMS · Cornell

Computational scaling bottlenecks cause excessive runtimes and memory exhaustion

9 theses · 8 institutions

Large sample sizes and high feature dimensions cause multi-hour runtimes, high inference latency, and out-of-memory crashes during kernel matrix computation. Practitioners frequently abandon support vector algorithms because their computational costs scale poorly compared to alternative approaches.

Tried and failed

Support Vector Machine classification applied to multivariate time series telemetry anomaly detection. Outcome: worse than baseline. Reason: exhibited lowest classification accuracy and highest training time compared to ensemble and linear baselines

Machine Learning Based Spectrum Fingerprinting of Drones for Defensive Cyber Operations · Harvard

Tried and failed

Support vector regression applied to traffic signal preemption optimization. Outcome: did not generalise. Reason: High feature dimensionality relative to scenario sample count alongside high training times

Emergency Vehicle Preemption Strategies using Machine Learning to Optimize Traffic Operations · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Support Vector Machines (SVM) for core image classification due to exponential computational scaling bottlenecks on large datasets.

Application of Computational Methods to Data Integration and Geoscientific Problems in Mineral Exploration and Mining · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected Support Vector Machine (SVM) for subsequent trials due to multi-hour runtimes, high compute requirements, and negligible accuracy gains over RF

Evaluating Multifrequency SAR for Crop Classification in Saskatchewan, Canada and Georgia, USA · Carleton University Institutional Repository

Tried and failed

support vector machine on minimal frequency features applied to multi-class vibration fault diagnostics. Outcome: worse than baseline. Reason: insufficient feature dimensionality caused poor class separation and extreme training times

PROGNOSTICS AND HEALTH MANAGEMENT OF WIND TURBINE GEARBOXES: DETECTION AND DIAGNOSTICS USING MACHINE LEARNING · WTAMU Repository

Considered and rejected

Considered and rejected: Support Vector Machines (SVM) for load acceptance classification due to intractable run times on large transaction datasets.

Alternative Freight Contracts: Data-driven Design Under Uncertainty · MIT

Tried and failed

kernel support vector regression applied to high-dimensional genomic drug response prediction. Outcome: infeasible cost. Reason: distance kernel matrix computation caused out-of-memory crashes as feature dimensionality increased

Representation of features as images with neighborhood dependencies for improvement of anti-cancer drug sensitivity predictive modeling using machine learning · Texas Tech

Tried and failed

Support Vector Regression applied to pixelwise quantitative parameter mapping. Outcome: too slow. Reason: Poor computational scaling with large dataset sizes led to excessive training and inference times

Simultaneous morphological and quantitative imaging of cartilage using phase-cycled bSSFP · Imperial

Tried and failed

KNN and support vector regression applied to high-dimensional spatiotemporal inference. Outcome: worse than baseline. Reason: poor scaling and performance when handling high-dimensional multi-scale spatial and temporal features

Urban air pollution modelling with machine learning using fixed and mobile sensors · Imperial

Models fail to generalize across domain shifts and operational variations

7 theses · 5 institutions

Support vector models experience sharp performance drops when transferred across different system configurations, real-world optical measurements, and open-set conditions. Overfitting to training distributions and a lack of ensemble diversity prevent effective generalization to unseen experimental environments.

Tried and failed

support vector machine for pattern classification applied to fault classification across system configurations. Outcome: did not generalise. Reason: Model failed to generalize to different system operational configurations without larger training sets.

Protection and Cybersecurity in Inverter-Based Microgrids · Virginia Tech

Tried and failed

support vector regression on multiscale genomic windows applied to predicting repressive chromatin mark signals. Outcome: did not generalise. Reason: models failed to predict H3K9me3 distribution across validation ChIP-seq and CUT&RUN datasets

THE MOLECULAR INTERPLAY AND EVOLUTION OF PROMOTER-PROXIMAL PAUSING AND ITS RELATIONSHIP TO HISTONE MODIFICATIONS · Cornell

Tried and failed

high upsampling rates in support vector regression applied to digital predistortion for power amplifiers. Outcome: overfit. Reason: Excessive sampling redundancy led to overfitting and degraded worst-case error vector magnitude performance.

Traversing the Volterra series for digital predistortion applications · Iowa State

Tried and failed

bagging SVM ensembles for uncertainty estimation applied to out-of-distribution workload detection. Outcome: did not generalise. Reason: insufficient base-classifier diversity due to convex optimization of support vector machines

Trustworthy Binary Classifications in Dynamic Systems Under Uncertainty · Georgia Tech

Tried and failed

support vector regression for spectral super-resolution applied to hyperspectral image reconstruction. Outcome: did not generalise. Reason: introduced high-frequency noise and failed to generalize from synthetic shifts to real-world optical measurements

Computational approaches for sub-meter ocean color remote sensing · MIT

Tried and failed

deep support vector data description applied to cross-domain open-set fault diagnosis. Outcome: did not generalise. Reason: struggled to delineate effective decision boundaries across domain shifts and open-set conditions

Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech

Considered and rejected

Considered and rejected: Rejected classic radial basis function (RBF) and support vector machine (SVM) base models in CTS meta-modeling due to dataset overfitting/bias and poor generalizability to unseen designs.

Machine Learning in Physical Design for 2D and 3D Integrated Circuits · Georgia Tech

Tried and failed

support vector regression for spectral super-resolution applied to hyperspectral reconstruction from spatially oversampled imagery. Outcome: did not generalise. Reason: failed to reconstruct mid-to-high frequency spectral features under optical tilt and real-world filter variations

Computational approaches for sub-meter ocean color remote sensing · MIT

Poor probability calibration and lack of interpretability hinder adoption

3 theses · 2 institutions

Support vector machines do not naturally output calibrated probabilities, leading to poor out-of-sample risk scores and prior inference failures. Practitioners also reject these models because their parameters and hyperplanes are difficult to interpret compared to standard statistical models.

Tried and failed

support vector machines for classification applied to drug approval prediction. Outcome: worse than baseline. Reason: poor probability calibration led to the worst out-of-sample performance

Analytics for accelerating biomedical innovation · MIT

Considered and rejected

Considered and rejected: Support Vector Machines (SVMs) and K-Nearest Neighbors (KNN) were considered but rejected due to lack of interpretability and poor handling of high-dimensional tabular feature importance.

Using Predictive Models to Identify Trends Among Successful Dual-Use Startups · MIT

Considered and rejected

Considered and rejected: Rejected using Support Vector Machines (SVM) to learn prior probabilities in AttriInfer because SVM does not directly produce calibrated probabilities.

Privacy Protection via Adversarial Examples · DukeSpace

Left open by the authors

Problems the authors named and did not get to.

Left open

Develop multi-class support vector machine methods using surrogate loss to simultaneously estimate minimal clinically important differences for improvement and worsening. Blocker: None

An Estimation and Inferential Framework for Minimal Clinically Important Difference · DSpace at SUNY Buffalo

Left open

Train and evaluate Support Vector Machines and Gradient Boosting Machines on FEA-derived temperature history features to predict DED residual stresses. Blocker: Requires the proprietary SYSWELD/ANSYS FEA simulation models or the author's private dataset of simulated temperature histories and stresses

Residual stress in metal additive manufacturing of thin-walled components: investigation and development of prediction models with respect to path planning · Scholarship at UWindsor Institutional Repository

Left open

Develop improved Support Vector Machine algorithms for doubly censored survival data. Blocker: No specific formulation or concrete algorithmic approach is provided beyond a high-level suggestion.

Survival Tree Models and Survival Ensemble Methods in Machine Learning · IRIS - UNICAM - prod

Left open

Develop an optimal engine oil emission experimental testing protocol using one-class support vector machines (OCSVM). Blocker: Requires experimental engine oil emission datasets or test bench access

A deterministic model for wear of piston ring and liner and a machine learning-based model for engine oil emissions · MIT

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.