Chapter Four · failure evidence
What Recurrent Neural Networks got wrong, from 47 dissertations
Recurrent neural networks frequently encounter practical difficulties ranging from gradient pathologies and error accumulation over long sequences to inferior performance compared to tree-based baselines on tabular sequential tasks. Practitioners also reject or abandon these architectures due to computational overhead, lack of hidden state interpretability, and poor generalization under distribution shifts or inadequate regularization. These records come from PhD theses at 18 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Recurrent neural networks underperform simpler baselines and tree-based models on tabular and sequential features
Deep sequence models repeatedly perform worse than gradient boosted trees, random forests, support vector machines, and simple feedforward networks on sequential regression and classification tasks. These performance deficits are especially pronounced on tabular sensor readings, short clinical trajectories, and time-series data with limited training sample sizes.
Tried and failed
LSTM on sentiment and technical features applied to financial price direction prediction. Outcome: worse than baseline. Reason: Recurrent neural networks underperformed traditional support vector machines on multimodal tabular and sentiment features
Tried and failed
ARIMA and recurrent neural networks applied to transient emission time series prediction. Outcome: worse than baseline. Reason: severe non-linearity and limited training sample size caused overfitting and high variance compared to gradient boosting
Tried and failed
recurrent and time-delay neural networks applied to aerodynamic time-series surrogate modeling. Outcome: worse than baseline. Reason: training sensitivity caused recurrent architectures to underperform simpler feedforward networks on time-series data
Tried and failed
recurrent neural networks for time-series regression applied to tabular time-series behavior forecasting. Outcome: worse than baseline. Reason: Deep sequence models underperformed gradient boosted trees on tabular sequential regression features.
Understanding and Predicting Sit-Stand Desk Usage Patterns and Willingness among Knowledge Workers: A Data-Driven Approach · Virginia Tech
Tried and failed
Stacked Bi-LSTM applied to time-series anomaly detection. Outcome: worse than baseline. Reason: Tree-based ensembles outperformed the recurrent neural network on the tabular sensor features.
Anomaly Detection and Prevention in Smart Grid Transmission System Using Machine Learning · Texas Tech
Tried and failed
recurrent neural networks for clinical trajectory modelling applied to short tabular longitudinal laboratory trajectories. Outcome: worse than baseline. Reason: failed to outperform simple logistic regression and exhibited substantially wider confidence intervals
Tried and failed
recurrent neural networks applied to time-series trigger classification. Outcome: worse than baseline. Reason: simpler feedforward neural network performed better or was preferred for classification
Tried and failed
recurrent neural networks for tabular wear forecasting applied to component degradation time series regression. Outcome: worse than baseline. Reason: deep sequence models did not outperform tree-based regressors despite higher computational cost
Digital Twin-Driven Condition Monitoring Approach for Aircraft Carbon Brakes · Georgia Tech
Considered and rejected
Considered and rejected: Rejected Recurrent Neural Network (RNN) approach for future streamflow modeling due to superior objective function metrics achieved by Random Forest
Climate change impact on the spatial distribution of droughts in Kirindi oya and Maduru oya dry zone river basins in Sri Lanka · Institutional Repository University of Moratuwa
Vanishing and exploding gradients impede training across long sequences and stiff dynamics
Standard recurrent networks suffer from severe gradient vanishing and exploding during backpropagation through time over extended sequences or deep computational graphs. These gradient pathologies prevent models from retaining long-term dependencies and cause dynamic tracking instabilities or training failure on stiff differential equations.
Considered and rejected
Considered and rejected: Rejected standard Recurrent Neural Networks (RNNs) due to vanishing gradient issues with long-term sequential dependencies in power load cycles.
IMPROVING CYBER RESILIENCE OF SHIPBOARD POWER SYSTEMS USING MACHINE LEARNING · Calhoun
Tried and failed
Standard recurrent neural networks applied to long-term time series prediction. Reason: Gradient vanishing prevents retaining early sequence features over long horizons
Wind Power Prediction and Uncertainty Modeling for Power System Operation · HARVEST
Tried and failed
physics-informed neural networks and recurrent neural networks applied to stiff differential equations. Outcome: did not converge. Reason: severe gradient pathologies during training on stiff multiscale dynamics
Tried and failed
Vanilla recurrent neural network applied to Dynamic actuator system identification. Outcome: unstable. Reason: Vanishing and exploding gradients caused output fluctuations during periodic dynamic tracking
LSTM sequence-to-sequence based System Identification and feedforward-feedback control of piezoelectric actuators · Iowa State
Tried and failed
Training standard recurrent neural networks with backpropagation applied to long audio sequence recognition. Outcome: unstable. Reason: Exploding gradients during backpropagation through time over long sequences.
Biologically Inspired Spiking Neural Networks for Speech Recognition · EPFL
Considered and rejected
Considered and rejected: Rejected basic Simple Recurrent Neural Networks (RNNs) due to vanishing gradients that prevented learning long-range word dependencies across long sequences.
Data driven strategies for product design · UT Austin
Considered and rejected
Considered and rejected: Rejected Recurrent Neural Networks (RNNs) due to excessive resource consumption hindering real-time performance and vanishing gradient problems from frequent zero-velocity fixations
Leveraging the multi-modal nature of communication in immersive XR environments · Imperial
Considered and rejected
Considered and rejected: Rejected feedforward neural networks (FFNN) and standard recurrent neural networks (Elman and Jordan RNNs) for hydrologic sequence modeling due to inability to retain sequential memory and vulnerability to the vanishing gradient problem.
Machine Learning Methods for Modeling Streamflow in Intermittent Rivers and Ephemeral Streams · Texas Tech
Considered and rejected
Considered and rejected: Rejected standard recurrent neural networks (RNNs) for Processing Cells in favor of LSTM cells to avoid the vanishing gradient problem over deep logical circuits.
Pushing the frontier of quantum many-body simulation using classical computers · Cornell
Sequential error accumulation and history drift degrade multi-step predictions
Autoregressive rollouts and stateful recurrence accumulate errors across successive forecasting steps, leading to severe drift and decaying correlation over time. Sequence models also fail to retain early historical context over long horizons or suffer divergence when evaluating optimization steps without resetting hidden states.
Tried and failed
recurrent neural networks for time-series forecasting applied to long-term reservoir production forecasting. Outcome: worse than baseline. Reason: Severe error accumulation and inability to capture long-term decline trends compared to standard empirical baselines
Asset Development and Sweet Spot Identification in Unconventional Reservoirs Using Machine Learning Approaches · Texas Tech
Tried and failed
autoregressive LSTM for multi-step time series forecasting applied to daily aggregated emergency department arrivals. Outcome: worse than baseline. Reason: fine-grained recurrent models accumulate errors when aggregating over longer daily horizons compared to short hourly predictions
Tried and failed
Standard recurrent neural network applied to sequential degradation data modeling. Outcome: did not generalise. Reason: Captures only recent memory and fails to address long-term dependencies in the sequence
Tried and failed
recurrent neural network with convolutional layers applied to spatiotemporal fluid flow PDE dynamics. Outcome: did not generalise. Reason: fails to capture source patterns/location after initial steps and suffers decaying correlation over time
Deep Learning for Dynamical Systems: Modeling, Prediction, and Control · Georgia Tech
Tried and failed
recurrent neural networks for sequential decision making applied to error recovery action suggestion. Outcome: worse than baseline. Reason: human error accumulation over sequential steps degraded accuracy and robustness relative to non-recurrent models
Facilitating Reliable Autonomy with Human-Robot Interaction · Georgia Tech
Tried and failed
stateful recurrent neural networks in iterative solvers applied to coupled nonlinear equilibrium problems. Outcome: did not converge. Reason: evaluating sequence models across optimization steps without resetting hidden states causes fictitious history accumulation and divergence
Physics enabled Data-driven structural analysis for mechanical components and assemblies · Georgia Tech
High computational overhead and architectural constraints make recurrent models inferior to alternatives
Recurrent architectures face high inference state buffering overheads, lengthy training times, and numerical instability during higher-order hypergradient computation. Consequently, practitioners favor convolutional networks or transformers due to better spatial feature extraction, easier handling of missing data, and superior sequence modeling performance.
Considered and rejected
Considered and rejected: Rejected using Recurrent Neural Networks (RNNs) with backpropagation through time in GF due to numerical instability during higher-order hypergradient computation
On machine learning methods for time series with financial applications · Oxford
Considered and rejected
Considered and rejected: Rejected using recurrent neural networks (RNNs) for marker prediction due to difficulty handling missing values without discontinuities and higher computational cost.
From finger animation to full-body embodiment of avatars with different morphologies and proportions · EPFL
Considered and rejected
Considered and rejected: Rejected Recurrent Neural Networks (RNNs) because they use a single weight matrix across units, limiting spatial feature detection and requiring inordinate training time compared to CNNs.
Predicting Phenotypes From Novel Genomic Markers Using Deep Learning · HARVEST
Considered and rejected
Considered and rejected: BiLSTM-CRF architectures were rejected for named entity recognition in favor of BERT-based models because transformer-based models consistently outperform sequential recurrent neural networks.
A pipeline for data and knowledge extraction from material science literature to accelerate scientific discovery · Georgia Tech
Tried and failed
recurrent neural networks applied to hardware branch prediction. Outcome: worse than baseline. Reason: yielded lower prediction accuracy and high inference-engine state buffering overheads compared to CNN architectures
Using convolutional neural networks to improve branch prediction · UT Austin
Recurrent models fail to generalize across domain shifts and specialized data distributions
Continuous recurrent networks generate invalid negative values when applied to zero-inflated intermittent time series, and anomaly classifiers fail to distinguish physically similar sensor fault signatures. Additionally, models struggle with cross-task transfer, acoustic noise degradation, and synthetic text generated by modern architectures.
Tried and failed
continuous recurrent neural network regression applied to zero-inflated intermittent time series. Outcome: did not generalise. Reason: assuming a single continuous distribution failed to predict true zeros and generated invalid negative values
Machine Learning Methods for Modeling Streamflow in Intermittent Rivers and Ephemeral Streams · Texas Tech
Tried and failed
transfer learning using pretrained recurrent layers applied to cross-task text classification. Outcome: worse than baseline. Reason: None
Towards Explainable Event Detection and Extraction · Virginia Tech
Tried and failed
recurrent neural network text classifier applied to transformer-generated synthetic text detection. Outcome: did not generalise. Reason: architectures trained on older generation models fail to detect modern transformer-based outputs
Defending Against Misuse of Synthetic Media: Characterizing Real-world Challenges and Building Robust Defenses · Virginia Tech
Tried and failed
neuro-inspired feature extraction with recurrent networks applied to speech recognition under acoustic degradation. Outcome: did not generalise. Reason: models degraded on noise-vocoded speech but tolerated periodic speech, contradicting human perceptual robustness
Features of hearing: applications of machine learning to uncover the building blocks of hearing · Imperial
Tried and failed
recurrent neural network for multi-class anomaly classification applied to cyber-physical system sensor incident classification. Outcome: did not generalise. Reason: Failed to distinguish between physically similar fault signatures of missing vs corrupted sensor readings
H2OGAN: A Deep Learning Approach for Detecting and Generating Cyber-Physical Anomalies · Virginia Tech
Regularization strategies and layered structural configurations degrade recurrent performance
Applying dropout regularization to recurrent networks fails to mitigate complexity in stacked multi-layer setups and harms performance when datasets are already large. Similarly, layering unconstrained recurrent units before constrained units or adding redundant sensor features causes performance degradation without improving generalization.
Tried and failed
Dropout regularization in deep multi-layer LSTM networks applied to time series imputation. Outcome: worse than baseline. Reason: Failed to mitigate added complexity and error inflation caused by stacking multiple recurrent layers
Data Driven Early Stage Design Support for Offshore Wind Farms · Research Repository UCD
Tried and failed
applying dropout regularization to recurrent neural networks applied to time series prediction on large datasets. Outcome: worse than baseline. Reason: regularization degraded performance without improving generalization due to the sufficiently large training dataset size
Traffic Signal Phase and Timing Prediction: A Machine Learning and Controller Logic Hybrid Approach · Virginia Tech
Tried and failed
layering unconstrained recurrent units before constrained units applied to acoustic feature modeling in speech recognition. Outcome: worse than baseline. Reason: None
Biologically Inspired Spiking Neural Networks for Speech Recognition · EPFL
Tried and failed
adding redundant sensor features to recurrent networks applied to kinematic trajectory prediction. Outcome: worse than baseline. Reason: additional sensor features introduced redundancy that caused performance degradation or offered no improvement
Gait Phase Estimation and Foot Trajectory Prediction During Dynamic Walking Using Gated Recurrent Units · Virginia Tech
Opaque hidden states reduce model interpretability compared to traditional methods
Recurrent neural networks are rejected in time series applications because their hidden state representations lack interpretability for multivariate processes. Practitioners instead select traditional models such as SARIMA or Gaussian mixture hidden Markov models to retain transparent dynamics and superior efficiency.
Considered and rejected
Considered and rejected: Rejected sequence models (Gated Recurrent Units / GRU) and SDEs for temperature, HVAC consumption, and electricity spot price processes in favor of traditional time series models (SARIMA, Holt-Winters, ARX-GJR-GARCH) for superior interpretability and efficiency.
Three Essays on Energy Markets · JScholarship
Considered and rejected
Considered and rejected: Rejected recurrent neural networks (LSTMs) for temporal modeling due to opaque hidden states lacking interpretability for multivariate data, choosing Gaussian mixture HMMs instead
Introducing Productive Engagement for Social Robots Supporting Learning · EPFL
Considered and rejected
Considered and rejected: Rejected using recurrent networks to learn change detection metrics because it reduces model interpretability.
Detecting and Leveraging Changes in Temporal Data · Georgia Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Benchmark and compare the dual temporal attention recurrent model against time-series Transformers for many-to-many sequence prediction. Blocker: None
Recurrence and Temporal Attention Synergy for Optimal Time-Series Modeling and Interpretability · TXST Digital Repository
Left open
Implement recurrent architectures like LSTMs using historical sequential sensor data for RUL estimation and anomaly reconstruction. Blocker: None
Toward The Democratization of Industrial Machine Learning · Georgia Tech
Left open
Evaluate length-based text data augmentation across LSTM and recurrent convolutional neural network architectures on text classification benchmarks. Blocker: None
Analyzing Easy Data Augmentation Techniques for Text Classification · Harvard
Left open
Develop a robust training method for recurrent neural networks processing static images with both recurrence and internal noise. Blocker: None
EXPLAINING FEATURES OF SIMPLE HUMAN DECISIONS USING BAYESIAN NEURAL NETWORKS · Georgia Tech
Left open
Add a recurrent decoder or sliding attention mechanism over sequential window embeddings to incorporate temporal context in stereo-EEG seizure detection. Blocker: Requires clinical stereo-EEG dataset with sub-second channel-level seizure annotations used in the thesis
Left open
Develop recurrent neural networks or RNN-based GANs to map time-based geomechanical oilfield data. Blocker: No specific architecture, dataset, or performance targets are specified
Enhanced Oil Field Data-Wrangling using Machine Learning · Texas Tech
Left open
Implement and evaluate a recurrent neural network architecture conditioned on state history for the dynamics error predictor in Error-Aware Policies. Blocker: None
Learning Control Policies for Fall Prevention and Safety in Bipedal Locomotion · Georgia Tech
Left open
Train an LSTM-based recurrent neural network to automate coarse visual categorization of astronomical seed light curves. Blocker: None
ExoSpotter: Few Shot Relevance Feedback For Learning High Recall Exoplanet Search · MIT
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.