Chapter Four · failure evidence
What Ensemble Learning got wrong, from 47 dissertations
The records detail empirical challenges encountered when designing and applying ensemble learning methods across various domains. Many attempts failed because complex ensembling strategies fell short of simpler base models, suffered from overfitting, or introduced computational bottlenecks. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Complex ensembling and stacking underperform simpler baselines or individual base models
Stacked generalization and meta-learning models frequently lost to simpler linear baselines, standalone base models, or basic voting schemes. In multiple cases, aggregating model predictions produced higher prediction error and variance than the constituent learners on their own.
Tried and failed
random forest regression applied to predicting physical parameters from relaxometry data. Outcome: worse than baseline. Reason: non-linear ensemble failed to outperform a simpler quadratic general linear model baseline
Quantitative Microstructural Imaging for Clinical Use · EPFL
Tried and failed
test-time representation optimization with MLP feature reconstruction applied to ensemble feature aggregation for downstream regression. Outcome: worse than baseline. Reason: regularization and architectural tuning failed to outperform simple ensemble averaging on downstream linear regression
BEYOND VANILLA FINETUNING: APPROACHES TO MAXIMIZE THE BENEFITS OF PRETRAINING IN COMPUTER VISION · Cornell
Lost to a baseline
Ensemble methods (DRF, Stacked Ensemble, GBM) were beaten on out-of-sample evaluation data by the simpler Mallows's Cp linear regression model due to severe overfitting.
A Sequential Modeling Approach to Explain Complex Processes and Systems · Virginia Tech
Lost to a baseline
For Clinical+ perinatal risk prediction, classic Unweighted Logistic Regression achieved a higher test AUC (0.64, 95% CI: 0.59-0.69) than the tree-based Unweighted GBM (0.62, 95% CI: 0.57-0.67) and the Unweighted GBM-Based Ensemble (0.63, 95% CI: 0.59-0.69)
Tried and failed
stacked generalization ensembles with meta-learners applied to time-series crop yield prediction. Outcome: worse than baseline. Reason: violation of the IID assumption in temporal data despite using blocked sequential cross-validation
Optimized ensemble learning and its application in agriculture · Iowa State
Tried and failed
stacked generalization ensemble with neural networks applied to high-latitude ionospheric electron density modeling. Outcome: worse than baseline. Reason: Ensemble generalization degraded significantly on out-of-distribution test data relative to individual empirical physical models.
Topside Ionospheric Modeling using Machine Learning · Georgia Tech
Lost to a baseline
Stacking Mean Regression (SMR) produced higher error and variance than simpler Voting Mean Regression across all ensembles
Key Concepts and Methods in Point and Probabilistic Forecasting of Electricity Prices · IRIS - POLITO - prod
Lost to a baseline
Average weighted ensemble achieved lower RRMSE (8.87%) than optimized weighted ensemble (9.22%) in 2019 because LASSO performed poorly in training but well in the test year.
Lost to a baseline
The first ensemble model (E1) was outperformed by both the standalone RNN and LSTM models on R2 (E1 achieved R2 0.8765 vs RNN 0.8977 and LSTM 0.9190) and error metrics (MSE 0.8532 vs RNN 0.8036, LSTM 0.7148; MAE 0.6705 vs RNN 0.6310, LSTM 0.4949)
Integrating Transactive Energy and Machine Learning For Re-Energizing Wastewater Treatment Plants · TXST Digital Repository
Lost to a baseline
Stacked ensemble model underperformed individual base models on broadleaf-only plots (SEM RMSE = 112.55 Mg ha-1 vs. GLM RMSE = 105.72 Mg ha-1, GBM RMSE = 111.41 Mg ha-1)
Advanced Techniques For Prediction of Forest Above Ground Biomass Using Satellite Remote Sensing Data · IRIS - UNITN - prod
Lost to a baseline
Averaging uncalibrated model scores in a machine-learning-only ensemble performed significantly worse than individual constituent models (Models 1, 2a, and 4).
Strong gravitational lenses in the era of wide-field surveys · Oxford
Lost to a baseline
NN ensembles retrained in an active learning loop with ensemble uncertainty-guided eABF-GaMD underperformed single NNs with GMM-based uncertainty in test set prediction error.
Lost to a baseline
Stacking ensemble model (96% accuracy) lost to standalone SVM classifier (100% accuracy) in at-risk student prediction
Gradgroom : an integrated framework for an educational early warning system with mentor matching · Institutional Repository University of Moratuwa
Considered and rejected
Considered and rejected: Rejected neural network stacking/merging of ensemble predictions because it required fixing the number of paths beforehand and performed worse or comparable to parameter-free inverse variance merging
Considered and rejected
Considered and rejected: Ensemble modeling (SuperLearner combining RF, GLMNet, BART, XGBoost) was rejected in favor of random forest alone due to equivalent performance.
HARNESSING ANTIBODY KINETICS TO IMPROVE EPIDEMIOLOGIC INFERENCE: CASE STUDIES IN CHOLERA AND SARS-COV-2 · JScholarship
Advanced weighting and aggregation strategies fail to improve over simple averaging
Techniques designed to dynamically weigh or aggregate predictions often failed to deliver benefits over standard mean ensemble averaging. In some settings, aggregation weights collapsed to a single model, initial strong learners dominated subsequent ones, or parametric layers degraded output quality.
Tried and failed
k-means clustering for sub-ensemble selection applied to spatiotemporal trajectory forecasting. Outcome: worse than baseline. Reason: clustering selected sub-ensembles yielded little to no forecast error reduction over full ensemble
Understanding tropical cyclones using machine learning with satellite imagery · Imperial
Tried and failed
boosting with computed model weight aggregation applied to weakly supervised ensemble learning. Outcome: worse than baseline. Reason: initial strong learners dominated the ensemble predictions, preventing effective contribution from subsequent weak models
On the Resource Efficiency of Language Models · Georgia Tech
Tried and failed
ensemble averaging across multiple levels of theory applied to activation free energy prediction. Outcome: worse than baseline. Reason: averaging solvation energies across parameterisation levels did not improve accuracy over single-level calculations
Solvent design assisted by mechanistic insights: methods and application to peptide synthesis · Imperial
Tried and failed
weighted ensembling of tuned gradient boosting models applied to tabular regression target prediction. Outcome: worse than baseline. Reason: the ensemble weights collapsed to the single best performing model configuration without adding diversity
The Beautiful (Computer) Game: How Data Science Will Revolutionize the World's Most Popular Sport · UT Austin
Lost to a baseline
Stacked LASSO and Stacked LightGBM achieved lower RMSE (632 kg/ha and 756 kg/ha, respectively) than the Average Ensemble (761 kg/ha) for state-level corn yield forecasts.
Optimized ensemble learning and its application in agriculture · Iowa State
Considered and rejected
Considered and rejected: Rejected model fusion / ensemble averaging across classifiers because disparate model performance (U-Net/LightGBM vs XGBoost) reduced overall quality and destroyed interpretability.
Interpretacja danych geoprzestrzennych przy użyciu wyjaśnialnych metod uczenia maszynowego Interpretation of geospatial data using explainable machine learning methods · AMUR - Repozytorium Uniwersytetu im. Adama Mickiewicza w Poz
Tried and failed
learnable aggregation of expert prediction logits applied to multimodal mixture of experts ensemble. Outcome: worse than baseline. Reason: parametric aggregation layers underperformed simple mean averaging
Multimodal Land Cover Mapping from Remote Sensing Imagery, Species Observations, and Language · EPFL
Tried and failed
averaging kernel matrices before computing attribution applied to ensemble data attribution across models. Outcome: worse than baseline. Reason: averaging intermediate kernel representations degrades attribution quality compared to ensembling final attribution predictions
Tree ensembles overfit or fail to generalize compared to regularized or linear models
Tree-based ensemble architectures overfit feature spaces and internal dataset characteristics relative to kernel methods or regularized linear models. In cross-dataset and physical regression tasks, this overfitting degraded out-of-sample generalization.
Tried and failed
merging decision trees trained on heterogeneous feature sets applied to ensemble regression modeling. Outcome: did not generalise. Reason: combining trees grown independently on relative versus absolute feature scales caused severe prediction errors
Processing and Analysis of the Seismocardiogram to Enable Estimations of Blood Volume Decompensation Status · Georgia Tech
Tried and failed
Tree-based ensemble regression models applied to crystal dielectric constant prediction. Outcome: overfit. Reason: Tree ensembles overfit the feature space compared to kernel methods like SVR and KRR
Tried and failed
non-linear ensemble models for speech emotion recognition applied to cross-dataset acoustic emotion prediction. Outcome: did not generalise. Reason: complex models overfit internal dataset characteristics compared to simpler regularized linear models
Creating Links: Building an Educational Platform to Ask Relevant Questions in Education · MIT
Linear models and simple baselines fail to capture complex non-linear patterns
Simpler linear or naive models repeatedly underperformed non-linear ensemble methods like Random Forests and Gradient Boosted Trees. They lacked the capacity to capture non-linear feature interactions and suffered when handling collinear predictors.
Tried and failed
regularized logistic regression for risk prediction applied to tabular pregnancy outcome data. Outcome: worse than baseline. Reason: linear models could not capture complex non-linear feature interactions as effectively as tree ensembles
Lost to a baseline
Linear regression model yielded significantly higher prediction error compared to Random Forest and AdaBoost ensemble models for conductivity prediction.
Real-time monitoring of neuronal cells through electrohydrodynamic patterning of flexible graphene microelectrodes · Iowa State
Lost to a baseline
Logistic regression (CV AUC 0.6673, test AUC 0.6613) and Naïve Bayes (CV AUC 0.4711, test AUC 0.4760) lost to non-linear ensemble models Gradient Boosted Trees (test AUC 0.6782) and Random Forest (test AUC 0.6727) on predicting Zone A reenlistment
PRICE ELASTICITY OF MONETARY INCENTIVES ON DRIVING DESIRED BEHAVIOR · Calhoun
Considered and rejected
Considered and rejected: Rejected linear regression from the final scATAC-Express ensemble due to poor predictive capacity and vulnerability to collinearity
Integrating genomic and multiomic data for computational analysis of gene regulation in circulating immune cells · Georgia Tech
Considered and rejected
Considered and rejected: Rejected Logistic Regression in the final ensemble because it added no AUC performance gain over tree-based models and required separate data preparation pipelines.
A DATA-DRIVEN APPROACH TO PREDICTING AUSTRALIAN BUSHFIRES · Calhoun
Considered and rejected
Considered and rejected: Rejected linear regression, Kernel Ridge, and LASSO due to inferior test MAE compared to ensemble tree methods (GBR, ETR, RFR).
Material discovery and modelling for solid-state hydrogen storage and fuel cell applications · University of Nottingham Repository
Ensemble inference incurs excessive computational overhead and latency
Deploying ensemble models often introduced severe computational and latency bottlenecks that impeded low-latency streaming and operational settings. The scaling of forward passes and model count made full ensembling impractical under tight real-time constraints.
Tried and failed
Ensemble learning algorithms applied to Low-latency streaming classification. Outcome: too slow. Reason: Inference time scaled unfavorably with ensemble size and sample count.
Driver distraction detection using experimental methods and machine learning algorithms. · Cranfield
Considered and rejected
Considered and rejected: Rejected 4DVAR and ensemble-based data assimilation (EnKF/3DEnVar) due to prohibitive operational computational costs and latency constraints.
The assimilation of surface observations in the European Alpine region · IRIS - UNITN - prod
Considered and rejected
Considered and rejected: Rejected Bayesian neural networks (e.g. MC dropout) for mapping uncertainty due to high computational overhead of multiple forward passes compared to small ensembles.
ACTIVE LEARNING OF VISION-BASED REPRESENTATIONS FOR ROBOTICS · Penn
Ensembling struggles with data quality issues and distributional assumptions
Ensembles degraded or produced negative performance metrics when subjected to misspecified parametric distribution assumptions or widespread missing diagnostic data across sites. Neural network ensembles also suffered from localized prediction failures in specific input ranges.
Tried and failed
regularized regression and tree ensembles applied to clinical outcome prediction. Outcome: worse than baseline. Reason: high rates of missing diagnostic predictor data across sites
APPLYING MATHEMATICAL MODELS TO IMPROVE CLINICAL EVALUATION AND PREDICTION · Penn
Tried and failed
deep neural network ensembles with input augmentation applied to tabular environmental sensor regression. Outcome: worse than baseline. Reason: localised predictive failure in specific input feature ranges, underperforming simple mean baseline
Trustworthy Soft Sensing in Water Supply Systems using Deep Learning · Virginia Tech
Tried and failed
ensemble model output statistics with parametric distributions applied to ensemble weather forecast postprocessing. Outcome: worse than baseline. Reason: distributional assumptions yielded poor fit, leading to degraded error metrics and negative R-squared values
Leveraging Artificial Intelligence and Numerical Weather Prediction to Build Custom Smart Home Software · Texas Tech
Left open by the authors
Problems the authors named and did not get to.
Left open
Develop ensemble prediction models combining machine learning and deep learning algorithms for ordinal disease incidence and hospitalisation targets. Blocker: None
Epidemiology and forecasting of influenza and COVID-19 · Imperial
Left open
Ensemble ESLSTM-MTF models initialized with diverse hyperparameters and custom loss functions that penalize extreme yield errors. Blocker: None
Enhancing Winter Wheat Crop Yield Predictions: A Data-Driven, Incremental and Integrative Approach with Machine Learning · Research Repository UCD
Left open
Implement ensemble methods combining fault detection and identification models trained on different feature subsets to improve prediction performance. Blocker: None
Fault Detection and Identification of Large-scale Dynamical Systems · MIT
Left open
Implement an evaluation pipeline comparing the predictive performance and resource usage of ensemble data monitors against periodic retraining baselines. Blocker: None
A Toolkit for Synthetic Data Generation and Drift Detection in Regression Scenarios · Carleton University Institutional Repository
Left open
Develop training objectives for diverse ensembles to allow sub-models to use efficient architectures at reduced training cost. Blocker: No specific mathematical formulations or target metrics are provided beyond a broad research direction
Towards Efficient and Robust Deep Neural Network Models · DukeSpace
Left open
Develop hybrid or ensemble forecasting models combining multiple RNN architectures with statistical models for multi-step wind power forecasting. Blocker: None
Deep Learning-Based Medium to Long-Term Multi-Step Ahead Wind Power Generation Forecasting · TXST Digital Repository
Left open
Integrate random forest, neural networks, CART, LASSO, ridge, and PCA into the SuperLearner-hdPS ensemble framework for small sample sizes. Blocker: None
Real-World Bleeding with Ibrutinib in B-Cell Malignancies · Penn
Left open
Implement and integrate transformer-based architectures or ensemble techniques into the solar power prediction pipeline. Blocker: None
Left open
Analyze and identify root causes for the ensemble transfer model's performance drop on the PU6 and PU8 datasets. Blocker: None
Semantic Segmentation of Satellite Imagery using Positive and Unlabeled Learning. · Scholars' Bank
Left open
Fit and prune deeper regression trees on the ensemble HFT probability scores to extract specific high-frequency trading strategies. Blocker: Requires proprietary order book data with labeled or identifiable HFT orders.
Three Essays on Financial Economics · IRIS - UNITN - prod
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.