Chapter Four · failure evidence
What Autoencoders & Generative Modeling got wrong, from 106 dissertations
The records document widespread failure points across autoencoder and generative modeling architectures, highlighting recurring difficulties in optimization stability, representation quality, and benchmark competitiveness. Across diverse domains, these models frequently struggle with mode and posterior collapse, severe overfitting on source distributions, and underperformance relative to simpler baselines. These records come from PhD theses at 27 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Unsupervised compression and reconstruction fail to preserve task-relevant features
Standard autoencoders frequently optimize general reconstruction error at the expense of statistical interpretability or task-specific performance, sometimes treating critical anomalies as reconstructible noise. In addition, architectural bottlenecks and unconstrained representations often suffer from local optima, trivial identity mappings, or non-invertible feature spaces that undermine downstream utility.
Tried and failed
surrogate modeling on variational autoencoder latent space applied to property prediction from spatial representations. Outcome: worse than baseline. Reason: nonlinear latent embedding degraded forward predictions and increased uncertainty compared to linear dimensionality reduction
Neural Inverse Microstructure Design with Bayesian Scale-Bridging · Georgia Tech
Tried and failed
conditional variational autoencoder generating high-dimensional intermediate features applied to deep neural network feature activations. Outcome: did not converge. Reason: High feature dimensionality and non-linear classifier structure prevented effective training of small generative models
Enabling on-device domain adaptation of convolutional neural networks · Imperial
Tried and failed
Unimodal variational autoencoder applied to multimodal content representation learning. Outcome: worse than baseline. Reason: Missing cross-modal correlations reduces representation quality and reconstruction accuracy
THEORETICAL AND EMPIRICAL EXPLORATIONS OF INFLUENCER MARKETING · Penn
Tried and failed
variational autoencoder on covariance-filtered geometric features applied to single-cell chromatin dispersion clustering. Outcome: worse than baseline. Reason: removing correlated features discarded informative signals needed to separate characteristics beyond simple cluster counts
Considered and rejected
Considered and rejected: Rejected using Autoencoders for primary dimension identification because latent representations optimize reconstruction error rather than statistical interpretability
Towards soundscape fingerprinting: development, analysis and assessment of underlying acoustic dimensions to describe acoustic environments · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected an autoencoder architecture with symmetric 4-layer ConvLSTM encoder/decoder predicting autoregressively across multiple stochastic steps due to loss convergence failure and severe overfitting
Integrating Machine Learning Techniques for Streamlined Predictive Modeling in Cosmological Applications · Georgia Tech
Considered and rejected
Considered and rejected: Rejected standard Variational Autoencoder (VAE) architecture for turbofan RUL estimation because decoders fail on temporal state projections; replaced decoder with a regressor network (RVE).
Uncertainty quantification of faults in rotating machines · Texas Tech
Considered and rejected
Considered and rejected: Rejected non-convex dimensionality reduction (deep autoencoders) due to susceptibility to local optima, lack of robustness, and high trial-and-error training cost.
A Reduced Order Modeling Methodology for the Multidisciplinary Design Analysis of Hypersonic Aerial Systems · Georgia Tech
Considered and rejected
Considered and rejected: Rejected standard under-parameterized autoencoders with bottleneck layers because they cannot interpolate training examples or store them as exact fixed points.
Foundations of Machine Learning: Over-parameterization and Feature Learning · MIT
Considered and rejected
Considered and rejected: Rejected sigmoid activation functions because autoencoders failed to train effectively when handling fluctuating fields containing both positive and negative values.
Scientific machine learning for the analysis and reconstruction of turbulent flows · Imperial
Considered and rejected
Considered and rejected: Rejected training the autoencoder jointly from scratch with the decoder or using a single linear projection layer, as both lagged significantly behind pre-trained autoencoders in ROUGE metrics.
Efficient and Enhanced Text Summarization by Compressing and Data Augmentation for Transformers-Based Models · Scholarship at UWindsor Institutional Repository
Tried and failed
batch normalization applied to autoencoder anomaly detection. Reason: did not improve anomaly detection performance during model development
Developing AI Systems for Monitoring Heterogeneous Mental Health Disorders · Cornell
Tried and failed
denoising autoencoder confounder correction applied to detecting high-frequency biological outliers. Reason: autoencoder erroneously corrected true synthetic aberrations by treating them as technical noise or outliers
Tried and failed
nested autoencoder architecture applied to time series reconstruction. Reason: matched standard autoencoder reconstruction without providing significant accuracy gains
Tried and failed
L2 regularization on autoencoder dimensionality reduction applied to unsupervised anomaly detection. Reason: higher regularization penalties caused underfitting and unpredictable performance degradation
On the Effectiveness of Dimensionality Reduction for Unsupervised Structural Health Monitoring Anomaly Detection · Virginia Tech
Tried and failed
autoencoder-based anomaly detection on textured backgrounds applied to surface anomaly detection. Outcome: no signal. Reason: The network acts as a general compressor, reconstructing textured surface variations indistinguishably from actual anomalies.
Tried and failed
autoencoder feature extraction in bidirectional GAN applied to microbiome data imputation. Outcome: worse than baseline. Reason: autoencoder failed to capture phylogenetic spatial relationships as effectively as convolutional layers
Deep Learning for Enhancing Human and Environmental Health · Virginia Tech
Considered and rejected
Considered and rejected: Rejected standard autoencoders (predicting neurons from themselves) to discard non-shared private variability.
High-dimensional neuronal activity from low-dimensional latent dynamics: a solvable model · EPFL
Considered and rejected
Considered and rejected: Rejected autoencoder latent-space partition for expert model selection because autoencoders learn generic input distributions rather than task-specific model fitness.
Continuous Learning for Lightweight Machine Learning Inference at the Edge · MIT
Considered and rejected
Considered and rejected: Rejected autoencoder-based bottlenecks for PVN internal state prediction because compression limited prediction accuracy and risked trivial identity transformations without bottlenecks.
Considered and rejected
Considered and rejected: Rejected backpropagation/autoencoders for feature compression because autoencoders can reconstruct original sensitive data and require decoders.
An Efficient Multidimensional k-Anonymisation Strategy Using Self-Organising Maps · De Montfort Open Research Archive (DORA)
Considered and rejected
Considered and rejected: Rejected relying directly on raw noisy autoencoder reconstruction error for moving thresholds, adding an alternating binary convolutional filter for denoising.
A Hitchhikers Guide to Anomaly Detection: Machine Learning-based approaches to anomaly detection and Failure diagnosis in spacecraft avionics systems using low power devices & flight ready devices · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected autoencoder neural network optimizers for LDSC design because the feature extraction process is non-invertible, preventing gradient backpropagation.
From Vulnerability to Resilience: Securing Signal Transmission Against Jamming and Spoofing Attacks · Virginia Tech
Considered and rejected
Considered and rejected: Rejected fully unconstrained deep autoencoders for unsupervised dictionary learning because untying weights broke recovery of generative physical filters.
Generative models for neural time series with structured domain priors · MIT
Complex generative models and autoencoders underperform simpler baselines
Variational and standard autoencoders regularly trail traditional machine learning methods such as decision trees, random forests, one-class support vector machines, and raw feature inputs across classification and anomaly detection tasks. In several image and signal domains, simpler supervised networks or matched filters consistently outperform deep generative formulations while avoiding high training overhead.
Tried and failed
variational autoencoder applied to tabular financial earnings metrics. Outcome: worse than baseline. Reason: standard autoencoders offered simpler training and direct reconstruction without probabilistic complexity
Machine learning in finance: from earnings data analysis to algorithmic trading · Imperial
Tried and failed
semi-supervised variational autoencoders applied to protein fitness prediction with combinatorial mutations. Outcome: worse than baseline. Reason: jointly trained models failed to outperform simpler supervised baselines like linear regression
Optimizing Protein Fitness and Function with Sparse Experimental Data · Harvard
Tried and failed
sequence-to-sequence variational autoencoder applied to sketch to latent space mapping. Outcome: worse than baseline. Reason: failed to show significant improvement over simpler vector KNN regression despite higher training complexity
Lost to a baseline
Vanilla autoencoders, adversarial autoencoders (1xAAE and 4xAAE), and PCA achieved lower clustering homogeneity scores (by 6% to 9%) compared to the variational autoencoder in UPSIDE.
Understanding Cell State Transitions in Development and Disease · ResearchWorks
Lost to a baseline
Autoencoder (F2=0.8156, precision=0.469) was beaten by OCSVM (F2=0.9914, precision=0.958).
Data Analytics and Machine Learning Applications in Fermentation Processes and Molecular Property Prediction · Virginia Tech
Lost to a baseline
single-layer linear autoencoder lost to VAE in field classification (scored 0.771 vs 0.992 on 3-field drive, and 0.3934 vs 0.829 on 5-field drive)
Non-equilibrium physics: from spin glasses to machine and neural learning · MIT
Lost to a baseline
CNN autoencoder feature clustering (0.36 accuracy, 0.22 NMI) was outperformed by ViT encoder feature clustering (0.78 accuracy, 0.61 NMI)
Deep Learning and Augmented Reality for 3D human-machine interaction · IRIS - POLITO - prod
Lost to a baseline
Variational Autoencoders (VAEs) lose to standard Autoencoders (AEs) when training SNR and testing SNR perfectly match due to gaps in the latent space
Model and data driven approaches to wireless image transmission · Imperial
Considered and rejected
Considered and rejected: Rejected Variational Autoencoders (VAEs) for high-resolution semantic image synthesis of climate impacts because they generated less realistic images than GANs.
Deep Learning Emulators for Accessible Climate Projections · MIT
Tried and failed
latent diffusion models applied to high-resolution structured grayscale medical images. Outcome: worse than baseline. Reason: struggled with fine structures, failing to match generative adversarial network fidelity
Controllable synthetic algorithms and evaluations for annotated clinical image synthesis · Imperial
Tried and failed
adversarial training without generative modeling applied to semi-supervised image segmentation. Outcome: worse than baseline. Reason: distorted the training signal compared to supervised and generative alternatives
Lost to a baseline
Autoencoder-based models (mDSC <= 0.08) lost to supervised U-Net (mDSC 0.50) and GMM (mDSC 0.17) on ATLAS-T1w lesion segmentation
Probabilistic and causal reasoning in deep learning for imaging · Imperial
Lost to a baseline
On public Cartographer dataset, Autoencoder Network (w/o skips) achieved SSIM of 0.904, marginally outperforming the thesis method's 0.903 (though thesis achieved higher PSNR 16.907 vs 16.459).
Integrating Perception, Prediction and Control for Adaptive Mobile Navigation · JScholarship
Lost to a baseline
1D CNN architectures (Kachuee et al. and Acharya et al.) achieved higher sensitivity on the minority Fusion (F) class under FGSM, BIM, and HSJ adversarial attacks compared to ECG-ATK-GAN.
Lost to a baseline
Random Forest beat Autoencoders on pointwise anomaly detection accuracy (RF: 0.963 vs AE: 0.937) and F1-score (RF: 0.967 vs AE: 0.953)
An AI and data-driven approach to unwanted network traffic inspection · IRIS - POLITO - prod
Lost to a baseline
Autoencoder fault identification accuracy (0.9873) was beaten by Decision Tree (0.9898) and Neural Network (0.9897)
Fault diagnosis in aircraft fuel system components with machine learning algorithms · Cranfield
Lost to a baseline
Autoencoder and PCA dimensionality reduction were both beaten by raw feature representation (no dimensionality reduction) across all anomaly detection models (e.g., IF Top 100: 10.8 raw vs 6.6 autoencoder vs 0.2 PCA).
Lost to a baseline
Matched filter alone slightly outperformed the U-Net autoencoder on pure QPSK + AWGN due to non-linear distortion introduced by the autoencoder
On the Use of Deep Learning Models for Interference Detection and Mitigation · Virginia Tech
Lost to a baseline
Attention autoencoder was outperformed by CNN autoencoder on non-sparse TEC reconstruction (R=0.9480 vs R=0.9604).
Employing Machine Learning Techniques to Increase the Quality of Ionospheric Modeling · Georgia Tech
Lost to a baseline
Autoencoder + MC-Dropout achieved 86.73 ± 6.02% balanced accuracy and 80.46 ± 12.50% sensitivity, losing to standalone Autoencoder (88.37 ± 3.96% balanced accuracy, 82.57 ± 9.97% sensitivity).
Self-supervised learning and uncertainty estimation for surgical margin detection with mass spectrometry · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected deep learning models (e.g., autoencoders) for unknown attack detection because they require large training datasets and performed worse than tree-based models.
Diagnosis and Mitigation of Evolving Threats for Sustainable Security · Research Repository UCD
Lost to a baseline
SDVAE and SOS-DVAE had higher reconstruction loss (0.086–0.090) compared to unconstrained generative VAE (0.078) on the TST dataset when removing task condition.
Relating Traits to Electrophysiology using Factor Models · DukeSpace
Adversarial training suffers from mode collapse and optimization instability
Generative adversarial networks repeatedly encounter severe mode collapse, vanishing gradients, and training divergence when balancing generator and discriminator objectives. These persistent instabilities often cause models to produce low-diversity outputs or collapse into noise, leading researchers to discard adversarial training in favor of collaborative or non-adversarial losses.
Considered and rejected
Considered and rejected: Decided against using GANs for generative spectral modeling because of training instability, choosing Variational Autoencoders (VAEs) instead.
Information Content and Analysis of X-ray Absorption Spectroscopy and X-ray Emission Spectroscopy Using Machine Learning · ResearchWorks
Tried and failed
generative adversarial network without variational autoencoder applied to 3D voxel shape generation. Outcome: unstable. Reason: mode collapse prevented learning the latent probability distribution, generating only a few shapes
Deep Learning Based Manufacturing Capability Modeling for Process Planning Automation · Georgia Tech
Tried and failed
multi-scale gradient input to discriminator applied to generative adversarial network image synthesis. Outcome: worse than baseline. Reason: broke the generator-discriminator balance, degrading output quality
Efficient Deep Learning Computing: From TinyML to LargeLM · MIT
Tried and failed
Classical generative adversarial network applied to decay histogram distribution modeling. Outcome: unstable. Reason: Mode collapse and vanishing gradients during training on histogram data
Tried and failed
training generator ensemble without discriminator applied to generative adversarial network training. Outcome: did not converge. Reason: generator suffered severe mode collapse and converged to a single low-variance mode without adversarial feedback
Memory of Motion for Initializing Optimization in Robotics · EPFL
Tried and failed
factored latent codes for scene decomposition applied to generative image synthesis. Outcome: unstable. Reason: foreground latent code collapsed into background representation during training
Unsupervised learning of human movement from images · Imperial
Tried and failed
training GAN generator without discriminator applied to binary image generation. Outcome: did not converge. Reason: removing the adversarial feedback caused generated patterns to collapse into random noise
Algorithmic design of photonic structures with deep learning · Georgia Tech
Tried and failed
single-stage physics-constrained generative adversarial network applied to structural topology optimization design generation. Outcome: unstable. Reason: Mode collapse and severe constraint violation from combining complex physics loss with adversarial training in one stage.
Generative Design Using Deep Learning Methods for Functionality and Manufacturability · Georgia Tech
Tried and failed
differentially private generative adversarial networks applied to tabular synthetic data generation. Outcome: unstable. Reason: mode collapse caused failure to match one-way marginals across most variables
Considered and rejected
Considered and rejected: Rejected Generative Adversarial Networks (GANs) and standard INNs/MDNs due to mode collapse, vanishing gradients, lower precision, and training instabilities on lower-dimensional sub-manifolds.
Towards Efficient and Robust Robot Planning · DukeSpace
Considered and rejected
Considered and rejected: Rejected generative networks (GANs/VAEs) for creating high-uncertainty adversarial active learning candidates due to high compute, instability, and noisy unrealistic medical artifacts, opting instead for gradient/attack-based perturbation.
Weakly-supervised Learning for Cost-Effective Medical Image Analysis Weakly-Supervised learning for Label-Effective Medical Image Analysis · Research Repository UCD
Considered and rejected
Considered and rejected: Rejected Generative Adversarial Networks (GANs) for structure generation due to training instability, mode collapse, and lack of explicit likelihood/interpretable latent space.
On-Demand Inverse Design of Phononic Metamaterials via Deep Learning · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected pure gradient-based adversarial optimization for failure generation because local optima cause diversity collapse into a single failure mode.
Breaking things so you don’t have to: risk assessment and failure prediction for cyber-physical AI · MIT
Considered and rejected
Considered and rejected: GAN adversarial discriminator training in favor of collaborative L2 loss among Noise2Noise generators to reduce complexity
Denoising Low-Dose CT Images using Multi-frame techniques · HARVEST
Considered and rejected
Considered and rejected: Rejected standard GAN adversarial training for video prediction, opting for pixel-wise MSE/BCE losses paired with transformational state architecture.
Learning, Moving, And Predicting With Global Motion Representations · Penn
Tried and failed
conditional recurrent neural network generative model applied to property-conditioned molecular sequence generation. Reason: mode collapse during generation leading to low uniqueness in generated outputs
Designing Macromolecules using Machine Learning and Simulations · MIT
Considered and rejected
Considered and rejected: Rejected generative adversarial networks (GANs) for inverse CTLE design due to training instability and inability to systematically model multi-modal parameter distributions.
High-speed Channel Analysis and Design using Polynomial Chaos Theory and Machine Learning · Georgia Tech
Considered and rejected
Considered and rejected: Rejected standard Conditional Generative Adversarial Networks (CGANs) for residential load scenario forecasting because training is unstable when conditioning on continuous vector-valued historical time series.
Learning and Optimization for Efficient and Optimal Operations in Sustainable Power Systems · ResearchWorks
Variational autoencoders experience posterior collapse and latent space distortion
Imposing prior regularization and Kullback-Leibler divergence terms frequently warps latent trajectories and triggers posterior collapse instead of learning organized representations. Standard Gaussian assumptions also struggle to represent non-Gaussian data statistics, leading to noisy loss landscapes, blurry outputs, and loose variational bounds.
Tried and failed
autoregressive latent covariance structure applied to variational autoencoders. Reason: did not prevent posterior collapse when it occurred in standard VAE baseline
Tried and failed
L1 or L2 regularization on posterior parameters applied to variational autoencoder latent representations. Outcome: worse than baseline. Reason: Failed to increase representation sparsity compared to vanilla variational autoencoder baseline.
Injecting Inductive Biases into Distributed Representations of Text · Cambridge
Tried and failed
variational autoencoder for sparse latent representation applied to text sentence embeddings. Outcome: unstable. Reason: struggled to achieve steady and consistent sparsity across hyperparameter configurations
Injecting Inductive Biases into Distributed Representations of Text · Cambridge
Tried and failed
variational autoencoder for latent trajectory modeling applied to cell shape dynamics over time. Outcome: worse than baseline. Reason: prior regularisation warped latent dynamical trajectories compared to standard autoencoders
Computing Interpretable Representations of Cell Morphodynamics · Imperial
Tried and failed
variational autoencoder without auxiliary predictor network applied to structured latent space performance prediction. Outcome: no signal. Reason: unsupervised reconstruction objective failed to organize the latent space with respect to the downstream target metric
Incorporating prior knowledge to efficiently design deep learning accelerators · UT Austin
Tried and failed
variational autoencoder with explicit mutual information maximization applied to image representation learning. Outcome: worse than baseline. Reason: leads to very high KL divergence, failing to match the marginalized posterior to the prior
Novel approaches for learning representations · UT Austin
Tried and failed
neural topic modeling using variational autoencoders applied to long academic documents. Outcome: worse than baseline. Reason: consistently produced negative normalized pointwise mutual information and low topic coherence scores
Topic Modeling for Heterogeneous Digital Libraries: Tailored Approaches Using Large Language Models · Virginia Tech
Considered and rejected
Considered and rejected: Rejected standard autoencoders (AE) in favor of variational autoencoders (VAE) because standard AE latent spaces are non-continuous and hinder smooth interpolation and balanced code generation
Gestalt Perception of Biological motion with a Generative Artificial Neural Network Model · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Standard Variational AutoEncoder (VAE) with KL-divergence regularization on a single latent variable was rejected for inverse continuous localization because stochastic sampling caused noisy loss landscapes.
Condition monitoring for dry cask storage using helical guided ultrsonic waves · UT Austin
Considered and rejected
Considered and rejected: Rejected Variational Autoencoders (VAEs) for conditional human mesh recovery because they do not permit direct and tractable likelihood evaluation for downstream tasks
Considered and rejected
Considered and rejected: Rejected standard VAE in favor of deterministic autoencoder because generative latent distribution modeling was unnecessary for embedding compression.
Neural Compression for Scalable Question-Answer Retrieval · DalSpace
Tried and failed
imposing latent space regularisation constraints applied to out-of-distribution detection via reconstruction error. Outcome: worse than baseline. Reason: latent constraints degrade reconstruction fidelity compared to unconstrained autoencoders, reducing anomaly detection discriminability
Learning Representations Toward the Understanding of Out-of-Distribution for Neural Networks · Georgia Tech
Tried and failed
conditional variational autoencoder with factorised latent space applied to 3D multi-object scene completion. Reason: failed to learn a structured prior, yielding incomplete reconstructions when sampling from Gaussian noise
Scene understanding for 3D multi-object scenes: labelling, reasoning and decomposing · Imperial
Tried and failed
variational autoencoders with Gaussian priors applied to systems with strongly non-Gaussian statistics. Outcome: did not generalise. Reason: standard Gaussian latent priors poorly represent strongly non-Gaussian statistics
Physics-Driven Machine Learning for Applications in Geophysical Fluid Dynamics · MIT
Tried and failed
i.i.d. latent sampling in variational autoencoders applied to set and graph generation. Reason: It optimizes an excessively loose lower bound on the true evidence lower bound.
Equivariant Neural Architectures for Representing and Generating Graphs · EPFL
Considered and rejected
Considered and rejected: Rejected likelihood-based and flow-based generative models (VAEs, PixelRNN) due to blurry outputs, slow generation, or excessive parameter counts
Deep Learning for Localized-Haptic Feedback in Tactile Surfaces · EPFL
Generative models fail to generalize beyond source training distributions
Autoencoders and generative models frequently overfit to nominal training distributions, leading to high false-alarm rates or outright failure when applied to unseen test structures and corrupted inputs. Under sparse data, occlusions, or complex physical constraints, these architectures struggle to extrapolate valid conformational landscapes or motion trajectories.
Tried and failed
MLP-based variational autoencoder applied to cross-dataset structural anomaly detection. Outcome: did not generalise. Reason: model overfitted to source structure distributions and failed on unseen structural data
Machine learning tools for identifying structural artifacts in data · Imperial
Tried and failed
generative neural network for data augmentation applied to imbalanced 3D medical image segmentation. Outcome: did not generalise. Reason: generated augmentations caused heavy overfitting to validation data and failed to generalize to unseen test data
Learning strategies for improving neural networks for image segmentation under class imbalance · Imperial
Tried and failed
standard generative adversarial network inversion applied to image restoration under gross corruptions. Outcome: did not generalise. Reason: Restored outputs deviate heavily toward corrupted regions without prior knowledge of anomaly masks.
Robust Learning for Fine-Grained Anomaly Detection in a Data-Rich but Label-Rare Environment · Georgia Tech
Tried and failed
adversarial bias field augmentation applied to medical image segmentation domain generalization. Outcome: did not generalise. Reason: increased vulnerability to out-of-distribution spike noise artifacts, degrading segmentation performance
Improving the domain generalization and robustness of neural networks for medical imaging · Imperial
Tried and failed
compositional generative mixture models for classification applied to instance segmentation under occlusion. Outcome: did not generalise. Reason: overlapping objects with similar textures caused ambiguous boundaries and merged instance detections
Addressing Occlusion in Panoptic Segmentation · Virginia Tech
Tried and failed
variational autoencoder anomaly detection applied to semiconductor fabrication process monitoring. Outcome: overfit. Reason: overfitting on nominal training data caused high reconstruction error on unseen nominal runs, degrading classification
Applications of Probabilistic Machine Learning Models to Semiconductor Fabrication · MIT
Tried and failed
reconstruction-based neural autoencoders applied to time series anomaly detection. Outcome: did not generalise. Reason: models over-smoothed extreme values or erroneously fit out-of-range outlier points
Machine Learning Systems for Unsupervised Time Series Anomaly Detection · MIT
Tried and failed
One-Class SVM and feedforward autoencoders applied to sequential anomaly detection in logs. Outcome: did not generalise. Reason: Unable to reliably detect anomalies without supervised attack knowledge for tuning hyperparameters
Data-driven Algorithms for Critical Detection Problems: From Healthcare to Cybersecurity Defenses · Virginia Tech
Tried and failed
validation loss for autoencoder model selection applied to sequential autoencoders on sparse count data. Outcome: overfit. Reason: loss does not penalize learning the identity mapping instead of underlying latent dynamics
General and interpretable models for inferring dynamical computation in biological neural networks · Georgia Tech
Tried and failed
continuous conditional variational autoencoder applied to recursive trajectory and motion generation. Outcome: did not generalise. Reason: Continuous latent space caused blurry generations and failed to follow conditioning velocity commands over recursive rollouts
Generative Latent Motion Planning and Reinforcement Learning for Legged Locomotion · MIT
Tried and failed
variational autoencoder with normalizing flow prior applied to biomolecular conformation generation. Outcome: did not generalise. Reason: coarse-grained coordinate priors failed to capture local dihedral torsional flexibility, enforcing overly planar backbone geometries
Physically Interpretable Biomolecular Conformation Generation with A Deep Probabilistic Framework · Harvard
Tried and failed
generalist generative models for molecular dynamics applied to protein conformational landscape sampling. Outcome: did not generalise. Reason: failed to accurately sample target conformational energy landscapes compared to ground truth
Learning on Graphs with Long-Range Dependencies: Methods and Applications · EPFL
Tried and failed
variational autoencoders for sequential modeling applied to multimodal agent state estimation. Outcome: did not generalise. Reason: failed to capture multimodal distributions under sparse and partial observations
Trajectory Modeling using Generative Approaches for Scheduling, Planning, and Multi-Agent Systems · Georgia Tech
Tried and failed
transfer learning with conditional recurrent neural networks applied to generative sequence design of polymers. Outcome: data insufficient. Reason: limited labeled training data prevented meaningful sequence reconstruction and led to invalid string generation
Designing Macromolecules using Machine Learning and Simulations · MIT
Considered and rejected
Considered and rejected: Rejected conditional variational autoencoders (cVAEs) for cross-subject neural decoding mapping because directed graphical models are inflexible when adapting to new destination subjects and require deep neural networks that are difficult to fine-tune on limited data.
Score-based Approach to Analysis of Unnormalized Models and Applications · DukeSpace
Considered and rejected
Considered and rejected: Rejected using Trajectory DBSCAN and WebPPL Bayesian Programming as the core generative simulation engines because generated paths were constrained copies of training data that failed to generalize to unseen sites
Machine Learning Simulation of Pedestrians Exploring the Built Environment · MIT
Autoregressive generative forecasting accumulates compounding temporal errors
Applying generative models to multi-step recursive forecasting results in rapid error accumulation across extended prediction horizons. These temporal modeling attempts suffer from unstable rollouts, severe overfitting on time-series inputs, and inaccurate imputation.
Tried and failed
standard diffusion sampling on corrupted data models applied to generative image reconstruction. Outcome: worse than baseline. Reason: produced repetitive, low-diversity generations lacking detail without momentum adjustment
Learning generative models from corrupted data · UT Austin
Tried and failed
latent space GAN autoregressive multi-step forecasting applied to deseasonalized time series prediction. Outcome: worse than baseline. Reason: severe error accumulation during recursive forecasting in the autoencoder latent space
Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction · Virginia Tech
Tried and failed
deep generative autoregressive forecasting without state estimation applied to spatiotemporal fluid and atmospheric dynamics. Outcome: unstable. Reason: rapid accumulation of autoregressive errors over long prediction horizons
Tried and failed
direct time series generative adversarial network forecasting applied to time series prediction. Reason: mathematical formulation issues and implementation discrepancies in original code
Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction · Virginia Tech
Considered and rejected
Considered and rejected: Rejected using Generative AI (standard GANs/GAIN) for time-based production data due to severe overfitting and inaccurate time-series imputation.
Enhanced Oil Field Data-Wrangling using Machine Learning · Texas Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Perform additional error analysis on the variational autoencoder component of the SG-UREVA relation extraction model. Blocker: Lack of specific hypotheses, error categories, or targeted evaluation metrics to investigate
N-ary Cross-sentence Relation Extraction: From Supervised to Unsupervised Learning · Virginia Tech
Left open
Apply minimax Pareto fairness optimization to variational autoencoders to reduce reconstruction error disparities across demographic groups. Blocker: None
Minimax Fairness in Machine Learning · DukeSpace
Left open
Improve CNN autoencoder surrogate model generalization to unseen geometries and features for finite element error indicator prediction. Blocker: Lacks specific target geometries, performance metrics, and a concrete architectural approach for out-of-distribution generalization
Development of Surrogate Model for FEM Error Prediction using Deep Learning · Virginia Tech
Left open
Validate AutoEKF on synthetic data and benchmark against particle filter methods, or replace linearization with variational autoencoders. Blocker: None
Left open
Evaluate alternative continuous-mapping models in place of convolutional autoencoders for GAN-based synthetic EHR generation. Blocker: EHR datasets such as MIMIC typically require credentialed data access agreements
Synthetic Electronic Medical Record Generation using Generative Adversarial Networks · Virginia Tech
Left open
Optimize hyperparameters for the input stress and output finite element error autoencoder architectures. Blocker: Lack of the original synthetic FEM stress/error training datasets and detailed baseline hyperparameter search space
Development of Surrogate Model for FEM Error Prediction using Deep Learning · Virginia Tech
Left open
Replace linear experts in the Generalized Variational Autoencoder with multilayer perceptrons or partially linear mixture-of-experts formulations. Blocker: None
Supervised Variational Autoencoders for Structural Learning and Statistical Inference with Heterogeneous Data · Virginia Tech
Left open
Develop uncertainty propagation from state to value uncertainty in partially observable environments using variational autoencoders for ADFQ. Blocker: Lacks specific analytical formulation or target POMDP environments for evaluating VAE-based uncertainty propagation.
Off-Policy Temporal Difference Learning For Robotics And Autonomous Systems · Penn
Left open
Extend the CRsAE model to handle non-concentrating posteriors by integrating variational autoencoders (VAEs) for uncertainty quantification. Blocker: None
Deep Learning for Inverse Problems in Engineering and Science · Harvard
Left open
Benchmark and systematically evaluate sparse autoencoders against clustering-based concept extraction methods across standard vision and video representation learning benchmarks. Blocker: None
Disentangling Visual Concepts Across Space and Time: From Image Hierarchies to Video Dynamics · YorkSpace
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.