Chapter Four · failure evidence
What Parameter Optimization & Hyperparameter Tuning got wrong, from 47 dissertations
Doctoral researchers investigating parameter optimization and hyperparameter tuning frequently encountered issues with severe overfitting, search instability, and heavy computational demands. Across diverse applications, sophisticated tuning routines often failed to outperform untuned defaults or simpler search strategies while regularly becoming trapped in suboptimal configurations. These records come from PhD theses at 18 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Hyperparameter tuning and extended fine tuning cause severe overfitting and degrade out of sample generalization
Optimizing hyperparameters or extending fine tuning repeatedly caused models to memorize training data, overemphasize specific lag features, or suffer catastrophic forgetting. In several cases, achieving higher training likelihoods or near-perfect fits directly harmed out-of-sample prediction and validation performance.
Tried and failed
gradient-based hyperparameter optimization for sparse Gaussian processes applied to robotic manipulation dynamic modeling. Outcome: overfit. Reason: optimizers consistently overestimated signal variance, treating true functional features as noise
Decision-Making Architectures for Control of Uncertain Systems · Georgia Tech
Tried and failed
fine-tuning models solely on synthetic generated data applied to automatic speech recognition models. Outcome: overfit. Reason: training exclusively on synthetic data led to overfitting and degraded generalization to genuine speech
Voice conversion and text-to-speech for privacy protection applications · Imperial
Considered and rejected
Considered and rejected: Rejected manual divergence regularization parameter tuning (beta-NSF) in data-constrained inverse problems due to frequent severe underfitting or overfitting.
Breaking things so you don’t have to: risk assessment and failure prediction for cyber-physical AI · MIT
Considered and rejected
Considered and rejected: Rejected full unconstrained 10-sublayer free phase/amplitude parameter optimization due to severe parameter degeneracy and overfitting of noise.
Reflectivity ferromagnetic resonance for layer-resolved dynamic study of multi-layered systems · Oxford
Tried and failed
prolonged fine-tuning of pretrained diffusion models applied to medical image synthesis. Outcome: overfit. Reason: extended training caused catastrophic forgetting and overfitting, degrading reconstruction quality
Multi-contrast magnetic resonance imaging with deep generative learning · UT Austin
Tried and failed
fine-tuning pre-trained transformers with limited data applied to multivariate time series forecasting. Outcome: worse than baseline. Reason: overfitting to seasonal and occupancy patterns degraded performance below zero-shot baseline
Toward Transformer-based Large Energy Models for Smart Energy Management · Virginia Tech
Tried and failed
random forest hyperparameter tuning via grid search applied to time series error correction. Outcome: overfit. Reason: overemphasized lag-1 features, degrading downstream simulation accuracy and inflating variance
Considered and rejected
Considered and rejected: Rejected hyperparameter tuning on RF error correction via out-of-bag error in favor of default parameters because it over-emphasized lag-1 error and worsened simulation variance.
Considered and rejected
Considered and rejected: Rejected using a running-window prior ensemble because small prior variance would require tuning ad-hoc variance inflation parameters across time steps without enough proxy data to prevent overfitting.
Quantifying changes in climate and surface elevation of polar ice sheets during the last glacial-interglacial transition · ResearchWorks
Tried and failed
random restarts for marginal likelihood hyperparameter optimization applied to Gaussian process regression. Outcome: overfit. Reason: higher marginal likelihood yielded models with poorer actual predictive accuracy
Accelerating HLS Autotuning of Large, Highly-parameterized Reconfigurable SoC Mappings · Penn
Tried and failed
physics-constrained hyperparameter optimization without cross-validation applied to time series prediction in volatile regimes. Outcome: overfit. Reason: Enforcing exact energy conservation caused condensed learning with near-perfect fit that degraded out-of-sample prediction.
LSSVR + PSO = MFG · Iowa State
Tried and failed
fine-tuning singular values with extended training applied to generative adversarial network adaptation. Outcome: overfit. Reason: Extended optimization caused training set memorization despite achieving misleadingly improved quantitative evaluation metrics.
Data-Efficient Learning in Image Synthesis and Instance Segmentation · Virginia Tech
Considered and rejected
Considered and rejected: Fine-tuning generative step models for >1 epoch, rejected due to detrimental overfitting.
Equipping language models for systematic reasoning · UT Austin
Considered and rejected
Considered and rejected: Rejected STRidge and LASSO in PDE-READ due to hyperparameter fine-tuning sensitivity and bias towards over-fitted solutions
Learning Differential Equations from Noisy, Limited Data · Cornell
Sophisticated hyperparameter optimization strategies perform worse than untuned baselines or simpler random searches
Complex optimization routines such as grid search, genetic algorithms, and Bayesian methods frequently trailed untuned default model settings or simpler random sampling. Several investigations showed that untuned baseline models achieved superior or near-identical accuracy while avoiding substantial tuning effort and inflated variance.
Lost to a baseline
Kalman with GCV hyperparameter optimization performed worse than Savitzky-Golay gridsearched optima on several benchmark systems.
Open-Source Dynamical Systems Research, with a Side of (Francis) Bacon · ResearchWorks
Tried and failed
physics parameter domain randomisation applied to sim-to-real transfer in robot manipulation. Outcome: worse than baseline. Reason: did not outperform simpler random force injection despite requiring significantly higher tuning effort
Exploring sim-to-real transfer for learning-based robot manipulation · Imperial
Tried and failed
random forest Bayesian optimization for configuration tuning applied to runtime configuration optimization. Outcome: worse than baseline. Reason: direct flag-to-runtime mapping struggled to navigate high-dimensional configuration spaces effectively compared to random search
Accelerating regression testing through test environment tuning · UT Austin
Lost to a baseline
Random search with twice as many samples outperformed standard BO methods without gradient information on hyperparameter optimization tasks.
Advances in Sparse and Bayesian Optimization for Autonomous Scientific Discovery · Cornell
Considered and rejected
Considered and rejected: Rejected hyperparameter optimization of C and gamma for SVM in favor of default parameters because defaults achieved near-identical accuracy while preserving generalizability across multiple fusion schemes.
Fusion Approaches to Individual Tree Species Classification Using Multi-Source Remotely Sensed Data · YorkSpace
Lost to a baseline
In Experiment 2, Grid Search hyperparameter tuning of XGBoost (MSE 0.0078, R2 0.7602, MAPE 0.3912) and Genetic Algorithm tuning (MSE 0.0073, R2 0.7639, MAPE 0.3912) lost to the base untuned XGBoost model (MSE 0.0067, R2 0.7690, MAPE 0.3843).
Using Data Analytics and Machine Learning in Sustainable Forest Management from Remote Sensing Data · YorkSpace
Lost to a baseline
In Experiment 1, Grid Search hyperparameter tuning of XGBoost (R2 0.8600, MAPE 0.2647) and Genetic Algorithm tuning (R2 0.8635, MAPE 0.2689) lost to the base XGBoost model (R2 0.8738, MAPE 0.2525).
Using Data Analytics and Machine Learning in Sustainable Forest Management from Remote Sensing Data · YorkSpace
Lost to a baseline
Kalman smoothing with GCV hyperparameter optimization was outperformed by gridsearched Kalman smoothing and gridsearched Savitzky-Golay across multiple benchmark ODEs.
Open-Source Dynamical Systems Research, with a Side of (Francis) Bacon · ResearchWorks
Hyperparameter tuning suffers from instability and failure to generalize across varying data conditions
Tuning procedures frequently degraded when confronted with data nonidealities such as extreme outliers, leverage points, varying physical conditions, or heterogeneous textures. Authors often chose globally fixed calibrations or direct network parameterizations to avoid unstable per-scan or online adjustments that masked imaging defects.
Tried and failed
standard Gaussian process regression applied to data with leverage points and outliers. Outcome: did not generalise. Reason: mean hyperparameter optimization via weighted least squares causes severe bias from vertical outliers and bad leverage points
Robust and Data-Driven Uncertainty Quantification Methods as Real-Time Decision Support in Data-Driven Models · Virginia Tech
Tried and failed
physics-constrained kernel regression with heuristic hyperparameter optimization applied to stochastic volatility modeling across expiries. Outcome: did not generalise. Reason: Failed to handle varying expiry dates even with additional parameter calibration
LSSVR + PSO = MFG · Iowa State
Considered and rejected
Considered and rejected: Decided against application-specific hyperparameter fine-tuning to ensure strict generalizability across imaging modalities
Generalizable low-latency accelerated dynamic MRI · Iowa State
Considered and rejected
Considered and rejected: Rejected per-scan tuning/optimization of hyperparameter α across varying collimations and object positions in the feasibility study in favor of a globally fixed calibration to prevent masking imaging nonidealities.
Algorithms for Quantitative Imaging with Advanced Cone-Beam Computed Tomography Configurations · JScholarship
Considered and rejected
Considered and rejected: Treating strong convexity constants mu_i as tunable hyperparameters was rejected in favor of parametrizing a neural network mu_zeta(x, u) to eliminate manual per-datapoint hyperparameter tuning.
Engineering AI systems and AI for engineering: compositionality and physics in learning · UT Austin
Considered and rejected
Considered and rejected: Using DSSIM as a direct loss function for autoencoder training was rejected due to hyperparameter tuning instability in high-resolution, heterogeneous breast textures
Machine Learning Approaches to Improve Diagnosis and Management of Mammographic Calcifications · DukeSpace
Tried and failed
extended hyperparameter optimization applied to time series regression with extreme outliers. Reason: tuning network hyperparameters failed to eliminate or reduce extreme prediction outlier errors
Machine Learning-Based Predictive Health Model of Turbofan Engine · Virginia Tech
Considered and rejected
Considered and rejected: Rejected online test-time adaptation in favor of episodic per-core adaptation due to unstable hyperparameter tuning.
Towards High-Fidelity Prostate Tissue Characterization and Cancer Detection with Micro-Ultrasound and Deep Learning · Queens University Institutional Repository
High computational overhead and poor scalability make extensive hyperparameter tuning impractical
Exhaustive parameter searches and formal optimization algorithms encountered severe scalability bottlenecks across high-dimensional spaces. Practitioners frequently abandoned extensive tuning because the substantial increase in computation time was not justified by the marginal accuracy gains.
Considered and rejected
Considered and rejected: Rejected SMT/Max-Sat based parameter optimization for continuous control due to poor scalability, choosing Bayesian optimization instead
Programmatic reinforcement learning · UT Austin
Tried and failed
autoencoder on proper orthogonal decomposition coefficients applied to subsurface reservoir simulation. Outcome: worse than baseline. Reason: required extensive hyperparameter tuning, suffered poor consistency across scenarios, and ran slower than baselines
Fast modelling of gas reservoirs using non-intrusive reduced order modelling and machine learning · Imperial
Considered and rejected
Considered and rejected: Rejected relying solely on manual feature engineering and hyperparameter tuning in favor of AutoML using H2O library to eliminate expert bias and intensive labor.
Machine learning (ml) approaches to model interdependencies between dynamic loads and crack propagation · Cranfield
Considered and rejected
Considered and rejected: Rejected grid search for hyperparameter tuning because it was computationally infeasible across the extensive parameter space.
A DATA-DRIVEN APPROACH TO PREDICTING AUSTRALIAN BUSHFIRES · Calhoun
Considered and rejected
Considered and rejected: Rejected full hyper-parameter tuning of ML models because marginal accuracy gains were offset by large tuning time increases.
Program analysis for machine learning models · UT Austin
Considered and rejected
Considered and rejected: Neural networks were rejected for predicting phase behavior because they were less accurate than SVR/RF and required excessive hyperparameter tuning time.
Optimization of chemical enhanced oil recovery methods for naturally fractured carbonate reservoirs · UT Austin
Parameter optimization algorithms get trapped in suboptimal parameter regimes and local optima
Autonomous search routines, Bayesian optimizers, and iterative language model tuners regularly suffered from premature convergence due to low exploration or missing domain guidance. These methods became trapped in suboptimal subspaces and poor local optima instead of discovering effective global parameters.
Tried and failed
autonomous reinforcement learning parameter optimization without expert priors applied to chemical vapor deposition growth recipes. Outcome: did not converge. Reason: search frequently became trapped in poor local optima without human domain guidance
Synthesis and Applications of Large-Area Monolayer Graphene · MIT
Tried and failed
Bayesian optimization with low exploration parameter applied to neural stimulation parameter tuning. Outcome: did not converge. Reason: insufficient exploration led to premature convergence to suboptimal parameter regimes
Considered and rejected
Considered and rejected: Rejected Bayesian Optimization for hyperparameter tuning due to its tendency to get trapped in local optima rather than global optima.
Optimization of Reverse Supply Chain For End-of-life Products · Cranfield
Tried and failed
LLM-driven iterative runtime hyperparameter optimization applied to large language model serving inference. Outcome: did not converge. Reason: LLM optimizer became trapped in a suboptimal parameter subspace during iterative search
LLM-Guided Run-Time Parameter Optimization for Energy-Efficient Model Inference · Virginia Tech
Considered and rejected
Considered and rejected: Rejected Tree-structured Parzen Estimator (TPE) in favor of grid search for hyperparameter tuning due to inferior performance in initial testing.
Improving the Efficiency of Deep Reinforcement Learning based UAV Obstacle Avoidance with Edge AI · Research Repository UCD
Left open by the authors
Problems the authors named and did not get to.
Left open
Apply formal hyper-parameter optimization and overfitting safeguards to artificial neural networks used in adsorption process design. Blocker: No specific neural network architecture, baseline dataset, or optimization protocol is specified.
Computational design of multi-sorbent adsorption processes for post-combustion carbon capture · Imperial
Left open
Incorporate robust optimization techniques into simulated annealing optimal decision tree training and hyperparameter tuning to reduce overfitting to noisy data. Blocker: None
Left open
Develop initialization and hyperparameter tuning algorithms for ensemble regression methods to prevent overfitting. Blocker: None
A Sequential Modeling Approach to Explain Complex Processes and Systems · Virginia Tech
Left open
Develop and implement methods for optimal hyperparameter tuning for the ENF-ADBEL neural network. Blocker: Lack of specific optimization target, baseline performance criteria, or designated optimization methodology.
Left open
Implement automated hyperparameter optimization for the 3D-DDnet low-dose CT denoising architecture. Blocker: None
A 3D Deep Learning Architecture for Denoising Low-Dose CT Scans · Virginia Tech
Left open
Automate hyperparameter tuning for the ComputeCOVID19+ CT framework using IterML rather than perturb-and-observe methods. Blocker: None
Real-Time Computed Tomography-based Medical Diagnosis Using Deep Learning · Virginia Tech
Left open
Develop an automatic hyperparameter optimization approach for deep learning models predicting remaining useful life. Blocker: None
Optimisation of deep learning techniques on remaining useful life prediction of complex engineering systems · Cranfield
Left open
Develop an automated hyperparameter optimization method for automatic thresholding or SVD cutoff selection in 3D ultrasound super-resolution imaging. Blocker: None
Ultrafast 3-D ultrasound super-resolution imaging with a row-column array · Imperial
Left open
Incorporate Bayesian optimization, sequential experimental designs, and context information into hyperparameter tuning for image analysis models. Blocker: None
Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models · Virginia Tech
Left open
Implement automated hyperparameter tuning using grid search or Bayesian optimization for the CNN-Bi-LSTM prognostics model. Blocker: None
Intelligent Data-Driven Maintenance Planning for Marine Renewable Energy Systems · DalSpace
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.