Chapter Four · failure evidence
What Neural Network Architectures got wrong, from 82 dissertations
The records document numerous experimental evaluations where neural network architectures underperformed simpler tree-based models, linear estimators, and domain-specific baselines. Across diverse domains including tabular classification, molecular modeling, and time series forecasting, neural networks frequently struggled with overfitting on small sample sizes and lacked physical interpretability. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Neural networks underperform tree-based algorithms and linear baselines on tabular datasets
Neural networks frequently yielded inferior accuracy compared to gradient boosted decision trees and random forests when applied to tabular features. Additionally, simple linear models and decision trees regularly outperformed tuned neural architectures on sparse, categorical, or multi-metric evaluation datasets.
Tried and failed
neural networks and automated tabular ensembling applied to tabular financial tournament dataset. Outcome: worse than baseline. Reason: failed to outperform gradient boosted decision trees
Tried and failed
shallow neural networks on small tabular data applied to spatiotemporal structure descriptor prediction. Outcome: worse than baseline. Reason: Neural networks overfit on small sample sizes compared to gradient boosting trees
Tried and failed
deep neural networks on tabular sensor data applied to health disorder prediction from multi-sensor streams. Outcome: worse than baseline. Reason: failed to outperform tree-based boosting models and suffered from poor sensitivity
ADVANCING DAIRY HEALTH MANAGEMENT THROUGH INTEGRATED SENSOR TECHNOLOGIES AND MACHINE LEARNING · Cornell
Tried and failed
neural network with focal loss applied to long-tail multiclass tabular classification. Outcome: worse than baseline. Reason: underperformed classical gradient boosted trees and achieved low accuracy despite hyperparameter grid search
Tried and failed
recurrent neural networks for time-series regression applied to tabular time-series behavior forecasting. Outcome: worse than baseline. Reason: Deep sequence models underperformed gradient boosted trees on tabular sequential regression features.
Understanding and Predicting Sit-Stand Desk Usage Patterns and Willingness among Knowledge Workers: A Data-Driven Approach · Virginia Tech
Tried and failed
neural network classification applied to sparse categorical tabular data. Outcome: worse than baseline. Reason: consistently performed poorly on sparse, categorical data
Using Urban Building Energy Modeling to Develop Carbon Reduction Pathways for Cities · MIT
Tried and failed
neural networks with default hyperparameters applied to tabular customer behavior classification. Outcome: worse than baseline. Reason: suboptimal performance compared to gradient boosting without extensive tuning
Tried and failed
deep neural networks on tabular features applied to signal-background classification in high-energy physics. Outcome: worse than baseline. Reason: DNNs yielded inferior separation compared to boosted decision trees on tabular kinematic features
Tried and failed
deep neural network ensembles with input augmentation applied to tabular environmental sensor regression. Outcome: worse than baseline. Reason: localised predictive failure in specific input feature ranges, underperforming simple mean baseline
Trustworthy Soft Sensing in Water Supply Systems using Deep Learning · Virginia Tech
Lost to a baseline
End-to-end Neural Network with GNN embeddings underperformed Random Forest and XGBoost baselines across all metrics (Accuracy 0.5836, F1 0.4873, AUC-ROC 0.5405, AUC-PR 0.4672)
Predicting psychological treatment dropout using graph neural networks · OpenBU
Lost to a baseline
Extensively tuned neural network (ROC-AUC 0.910), Random Forest (0.896), AdaBoost (0.907), and Decision Tree (0.817) lost to Gradient Boosting Classifier (0.918 on NTDB cross-validation / 0.924 final).
Lost to a baseline
Neural network classifiers (MLP and PNN) were beaten by standard J48 decision trees (83.25% vs ~20-63% accuracy) and Naive Bayes for faculty classification.
Der Hannover Concordancer und das Hannover Advanced Academic Writing Corpus: Eine korpuslinguistische Software mit dem dazugehörigen Dissertationskorpus für den Einsatz in Schreibberatungen · Leibniz Universität Hannover Repository
Lost to a baseline
Neural Network (58.49% accuracy) was beaten by Random Forest (79.82%), Decision Tree (74.85%), and KNN (70.18%) on the CAT MineStar Edge dataset.
Optimization of Quarry Operations and Maintenance Schedules · Virginia Tech
Lost to a baseline
Neural Networks were consistently beaten by Random Forest, AutoML, and Logistic Regression across grades 10-12 in Tukey HSD post hoc comparisons.
Lost to a baseline
Neural Network models performed poorly at predicting top 20% sales locations out-of-sample (7%-18% accuracy) compared to Random Forests (69%) and Linear Regression (47%-51%).
Using Data To Optimize The Fulfillment And Location Decisions Of An Online Grocery Retailer · Penn
Lost to a baseline
Neural networks (MLP, TabR) in CKB and FinnGen external validation were outperformed by Elastic Net and LASSO regression
Multi-omics studies of the molecular architecture of age, ageing and age-related disorders · Oxford
Lost to a baseline
Random forests (RF), gradient boosted trees (GB), and SVM outperformed the neural network in cross-validation precision/recall on data points labeled non-somatic / ambiguous.
DISCOVERY OF NOVEL GENOMIC MARKERS AND MOLECULAR SUBTYPES IN NEUROBLASTOMA · Penn
Considered and rejected
Considered and rejected: Rejected neural networks for occupancy profile classification due to poor performance on sparse categorical data in favor of tree-based models (Random Forest).
Using Urban Building Energy Modeling to Develop Carbon Reduction Pathways for Cities · MIT
Considered and rejected
Considered and rejected: Rejected Neural Networks (MLP) for final socio-physical prediction due to poor geographic generalization and lack of interpretability compared to GBRF
Human sensor networks for natural disasters · Imperial
Graph neural networks and scientific surrogates fail to exceed simpler physical baselines
Graph neural networks and geometric surrogates trained on molecular structures or physical systems provided minimal improvements over random forests and direct semi-empirical calculations. Furthermore, these architectures encountered issues such as over-smoothing from deep message passing, numerical instability, and systematic errors on out-of-distribution targets.
Tried and failed
graph neural network on atomic coordinates applied to predicting global scalar material properties. Outcome: worse than baseline. Reason: provided only marginal accuracy improvement over random forest baselines despite much higher complexity
Using Deep Learning to Understand and Design Heterogeneous Materials · MIT
Tried and failed
graph neural networks with language model embeddings applied to chemical reaction yield prediction. Outcome: worse than baseline. Reason: failed to outperform random forests trained on molecular fingerprints and physical descriptors
Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT
Tried and failed
augmenting graph neural networks with handcrafted molecular descriptors applied to molecular property prediction from semi-empirical calculations. Outcome: worse than baseline. Reason: hyperparameter tuning and manual descriptor concatenation yielded minimal performance gains or degraded predictive accuracy
High-throughput virtual screening of molecules for photon conversion · Imperial
Tried and failed
increasing graph neural network architecture complexity applied to predicting electronic structure properties of crystals. Outcome: worse than baseline. Reason: advanced architectures provided only marginal accuracy gains while sacrificing physical interpretability
Computational and Data-Driven Design of Perturbed Metal Sites for Catalytic Transformations · Virginia Tech
Tried and failed
higher-expressivity substructure features in graph neural networks applied to molecular property prediction. Outcome: did not generalise. Reason: substructures providing higher vertex disambiguation (paths/trees) did not capture domain-relevant inductive biases like cycles
Neural function approximation on graphs: shape modelling, graph discrimination & compression · Imperial
Tried and failed
graph neural networks applied to causal inference under interference. Outcome: worse than baseline. Reason: underperformed parametric models and baselines on predictive and causal inference simulation tasks
Propensity Scores for Causal Inference under Partial Interference · Harvard
Tried and failed
broadly trained graph neural network property prediction applied to ternary crystal magnetic property screening. Outcome: did not generalise. Reason: generalized crystal graph model exhibited systematic errors on multi-component compositions without domain-specific fine-tuning
The discovery and design of rare-earth-free magnets using machine learning · UT Austin
Tried and failed
graph neural networks as surrogate models applied to optimal power flow proxies. Outcome: worse than baseline. Reason: numerical instability and lower prediction accuracy compared to standard deep neural networks
Synergizing Machine Learning and Optimization: Scalable Real-Time Risk Assessment in Power Systems · Georgia Tech
Tried and failed
pure neural network potentials or semi-empirical methods applied to molecular relative energy prediction. Outcome: worse than baseline. Reason: insufficient accuracy compared to reference quantum mechanical relative energies
Bridging data and physics for computational drug discovery · UT Austin
Tried and failed
geometric neural network surrogate models applied to conformer ensemble property prediction. Outcome: worse than baseline. Reason: direct calculation from low-level semi-empirical ensembles outperformed geometric neural network surrogates trained on random conformers
Tried and failed
explicit hydrogen representation in equivariant graph neural networks applied to chemical reaction transition state property prediction. Outcome: worse than baseline. Reason: increased noise and degraded model performance compared to implicit representation
Encoding quantum-chemical knowledge into machine-learning models of complex molecular properties · EPFL
Tried and failed
neural network regression on simulation data applied to synthetic genetic circuit optimization. Outcome: worse than baseline. Reason: ML models did not statistically outperform traditional mechanistic mathematical modeling baselines on test designs
Tried and failed
deep message passing in graph neural networks applied to molecular property prediction. Outcome: did not generalise. Reason: Excessive message passing steps caused feature collapse and over-smoothing of node embeddings.
Advances in Artificial Intelligence for Accelerated Discovery of Energy Storage Polymers · Georgia Tech
Tried and failed
graph neural network property prediction applied to molecular photoionization cross-sections. Outcome: did not generalise. Reason: systematic underprediction for high-magnitude target values outside typical training ranges
Photoionization Detection of Volatile Organic Compounds · Harvard
Tried and failed
attentive fingerprint graph neural network applied to molecular property and activity prediction. Outcome: too slow. Reason: increased computational time without consistent performance gains over graph attention networks
Exploring Graph Neural Networks for Molecular Activity Prediction · Harvard
Considered and rejected
Considered and rejected: Rejected full graph neural network approaches due to scalability and graph size bottlenecks.
Neural Network-based Methodologies for Securing Cryptographic Code · Virginia Tech
Neural networks overfit and fail to generalize when training on small sample sizes
Deep and complex neural network models suffered from severe sample inefficiency and high variance when trained on limited datasets. Consequently, these models experienced significant drops in test accuracy and could not match the generalization capabilities of kernel or linear methods.
Considered and rejected
Considered and rejected: Using neural networks for behavioral statistical modeling when sample size is small, rejected because they fail to generalize to unseen data and often overfit compared to Gaussian process regression.
Statistical Modeling of the Effects of Process Variations on Silicon Photonics · MIT
Considered and rejected
Considered and rejected: Rejected deep neural networks and complex machine learning due to sample inefficiency, overfitting, and lack of clear performance advantages over kernel and linear methods
Key Concepts and Methods in Point and Probabilistic Forecasting of Electricity Prices · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected deep neural networks for time series activity classification due to high risk of overfitting on small tabular synthetic datasets and poor explainability/transparency.
Design of a Co-Orbital Threat Identification System · Virginia Tech
Considered and rejected
Considered and rejected: Rejected OpenNN and deep learning frameworks (Keras/TensorFlow) due to narrow focus on neural networks and excessive data requirements for small tabular data
Fatigue Monitoring System · SUPSI - ARIS
Tried and failed
2D deep convolutional neural network classification applied to 3D volumetric medical image classification. Outcome: did not generalise. Reason: Severe overfitting causing poor external validation performance compared to hand-crafted feature baselines
Tried and failed
deep convolutional neural network biomarker prediction applied to lumbar spinal dual-energy x-ray absorptiometry. Outcome: did not generalise. Reason: Severe performance drop and loss of clinical outcome associations when evaluated on an external dataset.
Deep-Learning to Assess Biological Aging From Spinal Dual-energy X-ray Absorptiometry · Harvard
Lost to a baseline
Neural network models (SNN and DNN) with L2 regularization underperformed simpler SVR and RFR models in generalizability on the unseen test dataset (DNN test MSE was nearly double that of RFR).
Towards Born Qualification of AM Components: High Temperature Fatigue Testing and Microstructural Characterization of AM IN718 · Georgia Tech
Considered and rejected
Considered and rejected: Rejected deeper neural architectures (30-40 neurons per layer) for experimental data due to failure to generalize on the small dataset.
Ligand-Protein Binding Affinity Prediction Using Machine Learning Scoring Functions · IRIS - UNICAM - prod
Considered and rejected
Considered and rejected: Rejected nonlinear ML models (e.g., deep neural networks, complex kernel methods) for primary low-sample reaction optimization due to severe overfitting risks and lack of interpretability
Bayesian and Data-Efficient Strategies for the Optimization of Organic Reactions · EPFL
Considered and rejected
Considered and rejected: Rejected Neural Networks due to higher computational intensity, lack of interpretability, and susceptibility to overfitting.
Considered and rejected
Considered and rejected: Feed-forward neural network with two hidden layers was tuned and implemented but rejected due to significant overfitting on the small, variable dataset.
Silviculture for Forest Diversity: Mixed Species Management to Meet Multiple Objectives · ResearchWorks
Tried and failed
increasing model depth in transformer architectures applied to low-resource neural machine translation. Outcome: worse than baseline. Reason: deeper networks suffered from severe degradation likely due to insufficient training data for larger parameter counts
Hebrew Transformed: Machine Translation of Hebrew Using the Transformer Architecture · Harvard
Considered and rejected
Considered and rejected: Decided against using deep neural sequence models for propensity and representation estimators on 10k synthetic datasets because they performed poorly compared to simple bag-of-words logistic regression models on small sample sizes.
BALANCING THE ASSUMPTIONS OF CAUSAL INFERENCE AND NATURAL LANGUAGE PROCESSING · JScholarship
Architectural adjustments and regularizations fail to improve model performance
Modifications such as increasing layer depth, regularizing bias weights, and applying dropout did not deliver performance gains on surrogate tasks or binary classification. Other specialized adjustments including layer freezing and neural stacking failed to outperform simpler pooling baselines or full network fine-tuning.
Tried and failed
increasing hidden layer depth in neural networks applied to tabular binary classification. Outcome: worse than baseline. Reason: increasing network depth did not improve classification accuracy
Novel deep learning approaches for flight delays prediction · Cranfield
Tried and failed
dropout regularization in shallow neural networks applied to deep neural network surrogate models. Reason: did not improve model performance or generalization on the surrogate modeling task
Urban building energy modeling : from sensitivity to big data · UT Austin
Lost to a baseline
Over-parameterized extended DSPM-DE (fitting 2n+4 parameters) achieved lower training MSE than the neural ODE, but overfit and generalized poorly
Considered and rejected
Considered and rejected: Rejected single hidden-layer shallow neural networks due to inferior generalization performance compared to deep architectures with equivalent parameter counts.
Steady-State Modelling for Natural Gas Transport Network Optimization · Queens University Institutional Repository
Lost to a baseline
Algebraic Neural Networks (a-NNs) failed to accurately capture state derivative profiles on the ethylbenzene system compared to Neural ODEs despite achieving adequate interpolation of state data
Lost to a baseline
Multilayered Neural-Network Nitrogen Removal (NNNR) model had lower prediction accuracy (testing R2 = 0.6819 for TNo, 0.7211 for NH3-No; RMSE = 8.4832 and 2.1523) compared to the Stepwise-Clustered Inference (SCI) model (testing R2 = 0.7894 for TNo, 0.8559 for NH3-No; RMSE = 6.2018 and 1.7026).
Development of Multi-Layer Soil Systems for Wastewater Treatment in Rural Areas · oURspace
Considered and rejected
Considered and rejected: Regularizing the bias parameters in neural networks was rejected because it introduces a significant amount of underfitting
DEEP-LEARNING-ENHANCED MULTIPHYSICS FLOW COMPUTATIONS FOR PROPULSION APPLICATIONS · Georgia Tech
Considered and rejected
Considered and rejected: Transfer learning by freezing/constraining early neural network layers was rejected because it degraded prediction performance relative to full-network fine-tuning.
Building Blocks of Neural Network Intermolecular Interaction Potentials · Georgia Tech
Considered and rejected
Considered and rejected: Rejected neural network stacking/merging of ensemble predictions because it required fixing the number of paths beforehand and performed worse or comparable to parameter-free inverse variance merging
Tried and failed
gradient clipping applied to low-resource neural machine translation. Reason: produced no statistically significant improvement over unclipped training baseline
Regularization Techniques for Low-Resource Machine Translation · EPFL
Lost to a baseline
Transformer language model reaction embeddings failed to outperform Random Forest, Extreme Gradient Boosting, or Feedforward Neural Networks using standard fingerprint or DFT features.
Machine Learning for Chemical Reactivity Prediction: Paradigms, Challenges, and Applications · MIT
Considered and rejected
Considered and rejected: Replacing the logistic regression classifier with a new fully connected neural network layer on top of extracted CNN features, which provided no significant improvement.
Deep Learning Methods for Breaking Ocean Waves · Research Repository UCD
Black box neural networks are rejected due to lack of interpretability and physical transparency
Practitioners rejected deep neural networks because opaque representations prevented clinicians and researchers from validating feature-level explanations. These black-box models also failed to provide physical interpretability or uncover explicit closed-form multiscale relationships.
Considered and rejected
Considered and rejected: Rejected machine learning (neural networks/random forests) due to high false-positive rates and uninterpretable feature explanations.
Development of computational tools for variant calling in single-cell RNAseq · Oxford
Considered and rejected
Considered and rejected: Rejected standard Graph Neural Networks (GNNs) and tabular ML with graph features because non-network relational schema caused incompatibility and lack of interpretability
Fraud Detection in Non-Network Knowledge Graph · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Neural networks were rejected due to requiring larger datasets and lacking interpretability.
Using Predictive Models to Identify Trends Among Successful Dual-Use Startups · MIT
Considered and rejected
Considered and rejected: Rejected Black-box Deep Neural Networks/LSTMs due to vanishing gradients, lack of physical interpretability, and inability to discover underlying modal dynamics.
Data-Driven Modeling of Tracked Order Vibration in Turbofan Engine · Virginia Tech
Considered and rejected
Considered and rejected: Rejected neural networks due to difficulty in providing reliable feature-level explainability for clinicians
Hydration Status and Early Stroke Outcome · JScholarship
Considered and rejected
Considered and rejected: Rejected standard Artificial Neural Networks (ANN) due to black-box representations, lack of physical interpretability, susceptibility to overfitting, and inability to discover explicit closed-form multiscale relationships.
Multiscale models based on statistical mechanics and physically-based machine learning for the thermo-hygro-mechanical behavior of spider-silk-like hierarchical materials · IRIS - UNITN - prod
Considered and rejected
Considered and rejected: Rejected purely data-driven black-box neural networks due to lack of physical interpretability and inability to explain governing electronic factors.
Toward Designing Active ORR Catalysts via Interpretable and Explainable Machine Learning · Virginia Tech
Considered and rejected
Considered and rejected: Rejected using complex non-linear deep learning or graph neural networks (GNNs) requiring structural input, choosing single-parameter-per-element compositional linear/ReLU models to retain human-interpretable chemical heuristics.
Learning Simple Chemical Heuristics to Model and Discover Materials · MIT
Recurrent and sequence models are outmatched by traditional time series and linear methods
Recurrent neural networks and deep sequence models underperformed classical statistical baselines such as ARIMA and ARIMAX on time series forecasting tasks. In other sequential settings, deep architectures suffered from high variance or were rejected in favor of transformers and faster logistic regression.
Tried and failed
ARIMA and recurrent neural networks applied to transient emission time series prediction. Outcome: worse than baseline. Reason: severe non-linearity and limited training sample size caused overfitting and high variance compared to gradient boosting
Lost to a baseline
ARIMAX (traditional statistical baseline) considerably outperformed all deep neural networks (LSTM, GRU, CNN) on clean data, achieving 21–24% lower MASE on the Dell dataset
Optimizing Sales Forecasting, Inventory, Pricing and Sourcing Decisions · EPFL
Considered and rejected
Considered and rejected: BiLSTM-CRF architectures were rejected for named entity recognition in favor of BERT-based models because transformer-based models consistently outperform sequential recurrent neural networks.
A pipeline for data and knowledge extraction from material science literature to accelerate scientific discovery · Georgia Tech
Considered and rejected
Considered and rejected: Rejected single-firm timeseries models (ARIMA) and deep learning models (LSTM, Neural Networks) due to short time horizons (average 6 firm-years) and low interpretability.
Corporate Disclosures Decoded: Forecasting Real Decarbonization Rates · Harvard
Considered and rejected
Considered and rejected: Rejected artificial neural networks (ANNs) and support vector machines (SVMs) for econometric forecasting because literature concluded they either underperformed ARIMA or did not justify added complexity.
Considered and rejected
Considered and rejected: Rejected Transformers and Naive-Bayes based solutions as the secondary IDS in favor of Logistic Regression (LRIDS) due to LRIDS having higher accuracy after neural networks and faster execution speed.
LIDS: An Extended LSTM Based Web Intrusion Detection System With Active and Distributed Learning · Virginia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Evaluate statistical and systematic uncertainty behaviors under Random Forest, SVM, and deep neural network models on student response coding datasets. Blocker: Access to private student physics response datasets used in the thesis.
Evaluating language models applied to student thinking about experiments · Cornell
Left open
Investigate why neural network impact position error does not scale proportionally with contact duration compared to XGBoost and 2D CNN models. Blocker: Requires the physical tactile surface experimental setup or private sensor measurements from the thesis.
Deep Learning for Localized-Haptic Feedback in Tactile Surfaces · EPFL
Left open
Compare Random Forests and Support Vector Machines against Artificial Neural Networks for mixed-frequency recession and GDP forecasting. Blocker: None
Using Mixed Frequency Data to Forecast Recessions and GDP · ResearchWorks
Left open
Evaluate distributional robustness, error patterns, and network capacity relationships of neural networks on non-power-system datasets. Blocker: Lack of specific target domains, datasets, or concrete experimental designs
NEURAL NETWORK DISTRIBUTIONAL INITIAL CONDITION ROBUSTNESS IN POWER SYSTEMS · Calhoun
Left open
Train GAP-SOAP or physics-informed neural network potentials on aluminum alloy DFT datasets to improve elastic constant and cluster energy predictions. Blocker: DFT training dataset generated in thesis may not be publicly archived or complete.
Neural Network Potentials for Age Hardening Aluminum Alloys · EPFL
Left open
Increase training sample size for cascade-forward neural networks and ANFIS models to mitigate underfitting while managing memory constraints. Blocker: Lack of specific target sample sizes or defined memory optimization architectures from the thesis
Left open
Implement Graph Neural Networks with explainability methods to classify brain functional connectivity networks instead of Logistic Regression. Blocker: Requires the private preclinical functional MRI connectivity dataset collected in the thesis.
Left open
Evaluate varying decision tree depths and random forest models for embedding subspace mapping in neural network generalization prediction. Blocker: None
DEPENDABLE NEURAL NETWORKS FOR SAFETY CRITICAL TASKS · JScholarship
Left open
Develop graph neural networks and graph variational autoencoders to optimize subgraph sampling and motif search on connectome graphs. Blocker: Lacks specific target architectures, quantitative efficiency benchmarks, or concrete graph sampling algorithms specified in the thesis
Left open
Derive generalization error bounds comparing Subgraph GNNs to standard message-passing neural networks in data-limited regimes mathematically. Blocker: Requires specialized statistical learning theory proof techniques not specified in the thesis
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.