Chapter Four · failure evidence

What Sensor Fusion & Multimodal Integration got wrong, from 94 dissertations

Across numerous multimodal integration studies, adding modalities or expanding sensor networks frequently degraded performance compared to standalone unimodal baselines. Failures stemmed from naive feature concatenation, sensory noise, synchronization offsets, and physical deployment constraints that corrupted joint representations and tracking stability. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Multi-sensor fusion underperforms single-sensor or unimodal baselines due to sensory noise and feature redundancy

25 theses · 14 institutions

Fusing all available sensor channels often increases estimation error and degrades classification compared to relying on a single dominant sensor modality. Uncorrelated signals, redundant inputs, and cross-sensor interference frequently degrade accuracy and harm cross-subject generalization.

Tried and failed

multi-sensor fusion with additional inertial measurement units applied to locomotion speed estimation. Outcome: worse than baseline. Reason: additional sensors did not reduce error and degraded user-independent cross-subject generalization

Sensor Fusion Representation of Locomotion Biomechanics with Applications in the Control of Lower Limb Prostheses · Georgia Tech

Tried and failed

multimodal fusion using all available sensors applied to beam prediction in wireless networks. Outcome: worse than baseline. Reason: excessive sensor integration introduced redundant or noisy features that slightly degraded accuracy

Reimagining wireless networks: Generative AI for integrated sensing and communications · Iowa State

Tried and failed

expanding multimodal sensor fusion for regression applied to cross-subject kinematic speed estimation. Outcome: worse than baseline. Reason: adding sensors beyond minimal IMU setup plateaued performance and increased unfiltered estimation error

Improving Intelligence of Robotic Lower-Limb Prostheses to Enhance Mobility for Individuals with Limb Loss · Georgia Tech

Tried and failed

multimodal sensor fusion using PCA applied to heterogeneous time-series signal classification. Outcome: worse than baseline. Reason: force sensor signals negatively interfered with current sensor features during fusion

Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech

Tried and failed

multimodal feature fusion with explicit posteriors applied to video action quality assessment. Outcome: worse than baseline. Reason: multimodal fusion interference degraded performance compared to single-modality baseline

Advancing Phonology-Based Sign Language Assessment: From Learner to Machine-Generated Videos · EPFL

Tried and failed

feature-level multi-band sensor fusion applied to ship wake detection in SAR imagery. Outcome: worse than baseline. Reason: multi-modal feature fusion introduced noise and degraded classification performance compared to single best band

Machine Learning and Data Fusion of Simulated Remote Sensing Data · Virginia Tech

Tried and failed

multi-axis sensor feature concatenation in naive Bayes applied to accelerometer-based state classification. Outcome: worse than baseline. Reason: additional axes introduced noise or violated conditional independence, lowering classification accuracy compared to single axis

Improved Vehicle Dynamics Sensing during Cornering for Trajectory Tracking using Robust Control and Intelligent Tires · Virginia Tech

Tried and failed

multichannel sensor fusion for anomaly detection applied to multichannel physiological time series. Outcome: worse than baseline. Reason: distant or localized anomalous activity across channels increased false positive rate during fusion

Real-time Personalized Monitoring of Neurological Disorders on Wearable Systems · EPFL

Lost to a baseline

Single combined SVM outperformed sensor fusion on PC and camera events where only 1-2 dominant sensors (power meters) existed

Ubiquitous sensing for security in smart homes · Oxford

Lost to a baseline

Single Low-Speed HS3 sensor slightly outperformed the fused MT1 60 + Low-Speed HS3 BNN model due to MT1 60 inconsistency at low power.

Uncertainty quantification of faults in rotating machines · Texas Tech

Lost to a baseline

In vacuum chicken batch testing, single-sensor MSI (0.7933 RMSEp) outperformed all early fusion configurations (e.g., MSI-FTIR early fusion at 1.2745 RMSEp).

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Lost to a baseline

Single vibration sensor (S3, 96.89%) beat dual-sensor fusion MS5 (VS+TS, 96.44%) under KAT B1 working condition.

Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech

Lost to a baseline

Single PDR modality alone (4.14 brpm MAE) beat all multi-modal fusion methods (MA-U-Net 5.14 brpm, SE-U-Net 4.44 brpm) during running

Enabling Accurate Cardiopulmonary Monitoring Using Machine Learning and a Chest-Worn Wearable Patch · Georgia Tech

Lost to a baseline

In combined aerobic+vacuum chicken batch testing, single-sensor MSI (0.8090 RMSEp) beat early-fusion MSI-FTIR (1.2714 RMSEp) and early-fusion MSI-FTIR-MSIF (1.0544 RMSEp).

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Lost to a baseline

Under LOSO, the multimodal fusion model DualLSTM-F (0.857 mean rho) did not outperform the unimodal vision model CNN-V (0.898 mean rho) or CNN-LSTM-V (0.861 mean rho)

Human-robot cooperation for teleoperation in robotic surgery · Imperial

Lost to a baseline

Wrist Accelerometer alone baseline achieved 0.81 F1 score, outperforming Wrist Gyroscope + Hip Gyroscope (0.75 F1) and multi-sensor combinations containing gyroscopes.

Towards Improving the Real Time Performance of Smartfall System · TXST Digital Repository

Lost to a baseline

In chicken thigh batch-on-batch validation, single-sensor MSI (0.6147 RMSEp aerobic, 0.7933 RMSEp vacuum) outperformed all single-sensor FTIR (1.6040 RMSEp aerobic, 1.2924 RMSEp vacuum) and MSIF models (2.2473 RMSEp aerobic, 1.0976 RMSEp vacuum).

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Considered and rejected

Considered and rejected: Rejected early fusion (concatenating DLG directly to input) because it degraded performance compared to range-image-only baselines

Pavement Crack Segmentation with Dense Local Geometry Features and Boundary Enhancement Loss · Georgia Tech

Tried and failed

multimodal RGB-D input fusion applied to 2D geometric thickness estimation. Outcome: worse than baseline. Reason: RGB modality introduced distracting appearance features that degraded depth-based geometric predictions

Multi-object scene completion using view-based network predictions · Imperial

Tried and failed

multimodal feature fusion via partial least squares applied to crop yield prediction. Reason: additional soil and field features provided redundant information already captured by remote sensing and weather

Soybean production management and remote sensing performance for grain yield prediction · Iowa State

Lost to a baseline

Full multimodal feature concatenation without SHAP feature selection performed worse or barely better than unimodal models (e.g., unaware MW RF F1 was 0.057 vs unimodal Video RF F1 of 0.081).

Multimodal Machine Learning for Automated Assessment of Attention-Related Processes during Learning · Publikationssystem UB Tuebingen

Lost to a baseline

BERT+LSTM + ResNet50 multimodal model (70.00% F1-score) performed worse than unimodal VGG19 visual classification alone (75.30% F1-score).

AI for social good: social media mining of migration discourse · Leibniz Universität Hannover Repository

Lost to a baseline

Under LOUO, multimodal DualLSTM-F (0.671 mean rho) was beaten by unimodal vision model CNN-LSTM-V (0.843 mean rho)

Human-robot cooperation for teleoperation in robotic surgery · Imperial

Tried and failed

multimodal feature concatenation for classification applied to infant gait pattern classification. Outcome: worse than baseline. Reason: integrating additional sensor features introduced noise that degraded model performance

Walking and talking: gait and its role in early development in autism · Penn

Tried and failed

retaining uncorrelated sensor features in prognostic models applied to remaining useful life prediction. Outcome: worse than baseline. Reason: uncorrelated and non-informative sensor channels introduced noise and degraded prognostic accuracy

Uncertainty quantification of faults in rotating machines · Texas Tech

Tried and failed

multimodal classification omitting primary kinematic modality applied to repetitive behavior classification. Outcome: worse than baseline. Reason: physiological signals lacked sufficient discriminative power without dominant motion features

Explainable and Robust Data-Driven Machine Learning Methods for Digital Healthcare Monitoring · Virginia Tech

Tried and failed

expanding feature set with noisy sensor modalities applied to interruptibility classification models. Outcome: worse than baseline. Reason: inclusion of noisy body orientation and audio angle features degraded classifier performance

Facilitating Reliable Autonomy with Human-Robot Interaction · Georgia Tech

Lost to a baseline

In normal conditions without sensor failure, standalone INS/GPS achieved slightly lower position RMSE (1.77 m) and velocity RMSE (0.73 m/s) compared to INS/VO/GPS DEKF (2.01 m position RMSE and 0.73 m/s velocity RMSE).

Multi-Sensor Fusion for Navigation of Ground Vehicles · Carleton University Institutional Repository

Lost to a baseline

In Motion 2 for Subject 1, event-based-only tracking was more accurate along x and y axes than the Kalman sensor fusion output (fusion error was ~24% higher for forearm).

Effective and safe framework for human-robot interaction · IRIS - POLITO - prod

Lost to a baseline

Modified stereo vision approach achieved 2.8 cm translation error, beating the proposed camera+IMU fusion method (4.7 cm), though stereo was twice as slow and lacked robust rotation.

Sensor Fusion and Stroke Learning in Robotic Table Tennis · Publikationssystem UB Tuebingen

Naive early feature concatenation and uniform integration cause modality dominance and representation collapse

19 theses · 13 institutions

Concatenating raw multimodal representations without adaptive weighting allows high-dimensional or dominant modalities to overshadow informative inputs. This uniform integration leads to representation space collapse, corrupted cross-modal alignment, and degraded downstream predictive accuracy.

Tried and failed

multimodal feature concatenation for anomaly detection applied to mechanical fault detection. Outcome: worse than baseline. Reason: redundant multimodal features degraded anomaly detection performance compared to single-sensor baselines

Practical and generally applicable condition based maintenance (CBM) system for mud pump · UT Austin

Tried and failed

direct feature addition or concatenation across modalities applied to RGB-D salient object detection. Outcome: worse than baseline. Reason: depth map noise corrupted multi-scale feature representations during naive early/mid fusion

Effective deep leaning methodologies for salient object detection · Imperial

Tried and failed

early fusion via raw feature concatenation applied to multimodal regression with high dimensional imbalance. Outcome: worse than baseline. Reason: large modality dimensionality disparities caused higher-dimensional inputs to dominate predictions

Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield

Tried and failed

concatenating multiple heterogeneous visual feature representations applied to multimodal regression prediction. Outcome: worse than baseline. Reason: feature redundancy and multicollinearity degraded predictive performance across models

Housing Price Prediction with Computer Vision and Image Features · Harvard

Tried and failed

direct multimodal imaging feature fusion applied to multimodal medical image classification. Outcome: worse than baseline. Reason: increased data complexity without clinical context failed to outperform unimodal baselines

Advancing Personalized Medicine Through Generative Artificial Intelligence · Georgia Tech

Tried and failed

concatenation feature fusion for multimodal representations applied to visual question answering. Outcome: worse than baseline. Reason: underperformed multiplicative and bilinear fusion methods across language model architectures

Interact with Earth Observation Images using AI: Transparent Methods and Evaluation for Visual Question Answering · EPFL

Tried and failed

multimodal network fed duplicate single modality applied to multimodal fusion models. Outcome: overfit. Reason: duplicate input branches increased capacity without added information, leading to memorization rather than ensembling

Multimodal and Context-Aware Computational Pathology · Harvard

Tried and failed

early fusion of visual modalities before language cross-attention applied to vision-and-language robot navigation. Outcome: worse than baseline. Reason: joint multimodal visual representation degraded cross-modal alignment with language compared to separate per-modality cross-attention

Learning 3D Robotics Perception using Inductive Priors · Georgia Tech

Tried and failed

concatenation followed by linear projection applied to multimodal feature map fusion. Reason: increased parameter count without significant performance gains over adaptive fusion

Interpretable time-series forecasting with multi-model deep learning and natural language processing (NLP) driven explainable artificial intelligence (XAI) · Iowa State

Tried and failed

naive concatenation of multimodal features applied to few-shot relation extraction. Outcome: worse than baseline. Reason: irrelevant visual noise degraded representation quality without selective multimodal fusion

Few-Shot and Zero-Shot Learning for Information Extraction · Virginia Tech

Tried and failed

multimodal fusion with uniform layer learning rates applied to multimodal imitation learning. Reason: Dominant image modalities caused network inattention to fused scalar inputs without differential learning rates.

Towards Improving and Extending Traditional Robot Autonomy with Human Guided Machine Learning · Virginia Tech

Considered and rejected

Considered and rejected: Rejected naive channel concatenation for multi-modal feature fusion between objects and RGB because it weights channels uniformly instead of selectively attending to motion.

Model-driven and Data-driven Methods for Recognizing Compositional Interactions from Videos · JScholarship

Considered and rejected

Considered and rejected: Rejected naive early fusion (feature vector concatenation) and late fusion (decision averaging) due to information loss on heterogeneous symbolic data and inability to capture cross-view correlations.

Novel methods for multi-view learning with applications in cyber security · Imperial

Considered and rejected

Considered and rejected: Rejected completely shared multimodal feature spaces/complete modality fusion because it introduces heavy bias towards concrete words and harms abstract concepts

Language Grounding in Vision · Publikationssystem UB Tuebingen

Considered and rejected

Considered and rejected: Rejected shallow modality separation (modality-specific FFNs only with shared QKV) in LLaMaFusion because shared QKV projections corrupted text representation spaces during image training.

Breaking the language model monolith · ResearchWorks

Considered and rejected

Considered and rejected: Rejected joint multi-modal fusion strategies (Linear, Bilinear, Concat, and Vanilla SA) for audio-visual egocentric gaze anticipation due to camera motion and audio-gaze reaction latency.

Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Late Fusion (sharing predicted bounding boxes) due to lack of contextual feature representations and heavy dependence on single-agent accuracy.

Enhancing Perception for Autonomous Vehicles · Queens University Institutional Repository

Tried and failed

incorporating distinct non-shared features without adaptive weighting applied to multimodal single-cell data integration. Outcome: worse than baseline. Reason: distinct features were excessively noisy relative to shared features, degrading alignment accuracy

Learning from Multiple and Heterogeneous Datasets · Penn

Tried and failed

multimodal joint fusion with single unified loss applied to multimodal clinical decision support. Reason: modality competition caused representation space collapse

Data-driven multimodal learning towards safer clinical decision support · Imperial

Multi-sensor wearable and edge deployments fail due to hardware overhead and inter-device variability

14 theses · 10 institutions

Complex multi-sensor wearable setups encounter significant practical hurdles including packet dropouts, inter-device variability, and high calibration complexity. Deploying multiple physical sensors also incurs prohibitive computational burdens and reduces participant compliance compared to simpler single-sensor configurations.

Considered and rejected

Considered and rejected: Rejected early and intermediate feature-level sensor fusion across vehicles due to high bandwidth constraints and incompatibility across diverse vehicle sensor suites.

Enhancing Perception Systems using V2V Sensor Fusion · Virginia Tech

Considered and rejected

Considered and rejected: Rejected using multi-IMU sensor fusion for model position in favor of a single master IMU to avoid inter-sensor noise and calibration discrepancies

Improving Usability for Novices in the Design of Mechatronic Devices: A Study Using Arduino Modules · Queens University Institutional Repository

Considered and rejected

Considered and rejected: Rejected activating all vehicle sensors simultaneously due to rapid battery depletion without coverage improvement.

Convergence Results for Ergodic Control of Ensembles via Iterated Function Systems · Research Repository UCD

Tried and failed

integrating heterogeneous consumer wearable sensor streams applied to remote physical activity tracking. Reason: inter-device measurement variability undermined data comparability across study participants

Rehabilitation through exercise prescription for cardiac patients using an artificial intelligence-based programme · Imperial

Tried and failed

raw wearable sensor telemetry for sequence modeling applied to joint moment estimation. Reason: sensor data suffered from packet dropout and signal saturation during dynamic movement

A Framework for Autonomous Exoskeleton Assistance Independent of Activity · Georgia Tech

Tried and failed

wearable sensor heuristic activity classification applied to travel mode and physical activity detection. Outcome: worse than baseline. Reason: built-in proprietary algorithms showed low accuracy compared to validated self-reported travel diaries

Applications of causal inference in environmental policy and transport studies · Imperial

Lost to a baseline

Wrist-worn sensors underperformed chest-worn sensors across all 5 holding assessment scenarios (0.738 vs 0.870 accuracy in Scenario 1)

Leveraging pervasive data to study and support mother-infant dyads in the wild · UT Austin

Considered and rejected

Considered and rejected: Rejected signal resampling and zero-padding for multiresolution data fusion due to excessive computational and memory overhead in wearable edge devices

Multi-sensor data fusion for ambulatory health monitoring: signal processing and deep learning techniques · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected multi-sensor IMU systems (gyroscopes/magnetometers) in favor of standalone wrist accelerometers because multi-sensor setups introduce sensor drift, synchronization errors, and reduce patient compliance.

Development and validation of a wearable sensing platform for assessment of patients undergoing breast and axillary surgery · Imperial

Considered and rejected

Considered and rejected: Rejected placing dual accelerometers on both hip and chest/wrist to avoid reducing patient compliance with multi-sensor wear.

Improving the surgical patient care pathway through use of activity monitoring · Oxford

Considered and rejected

Considered and rejected: Rejected using wearable sensor spectral monitoring data and activity diaries for quantitative prior-light-history tracking due to sensor inaccuracies/malfunctioning and inconsistent participant reporting.

Alertness in work environments : on the role of indoor daylight exposure · EPFL

Considered and rejected

Considered and rejected: Rejected cloth-based Bioharness sensor due to poor fit and low data quality.

Interfaces and Models for Improved Understanding of Real-World Communicative and Affective Nonverbal Vocalizations by Minimally Speaking Individuals · MIT

Considered and rejected

Considered and rejected: Rejected continuous high-density EMG for closed-loop peripheral phase tracking in favor of single-sensor triaxial accelerometry to minimize instrumentation and improve wearable compliance.

Probing and restoring disrupted thalamocortical interactions during Parkinson's disease and essential tremor · Oxford

Considered and rejected

Considered and rejected: Rejected bi-directional finger with two sensors design due to higher calibration complexity, sensor performance inconsistency, and mechanical interference/strain.

Capacitive Strain Sensor System for Soft-Rigid Hybrid Robotic Grippers · Harvard

Inertial navigation and kinematic state estimation diverge from compounding sensor drift and dynamic motion noise

13 theses · 10 institutions

Combining inertial measurement units and odometry data without robust compensation leads to rapid tracking divergence caused by compounding sensor bias and drift. Aggressive dynamic maneuvers, unmodeled surface friction, and environmental magnetic interference further corrupt state estimates.

Tried and failed

onboard camera and IMU sensor fusion applied to mobile robot state estimation and tracking. Outcome: no signal. Reason: low-cost onboard sensors suffered from excessive noise and drift for accurate positioning

Risk-aware and robust decision making for autonomous vehicles with reinforcement learning · Imperial

Tried and failed

redundant sensor fusion scaling applied to inertial navigation state estimation. Outcome: no signal. Reason: sensor systematic biases dominated over random noise, rendering additional sensors unhelpful

Identification and Integration of Aerodynamics into Fixed-Wing Drone Navigation · EPFL

Tried and failed

extended Kalman filter wheel odometry fusion applied to mobile robot track navigation. Outcome: unstable. Reason: unmodeled variable frictional slip between drive wheels and contact surfaces degraded state estimation

A novel railway maintenance robot for inspection and repair · Cranfield

Tried and failed

open-loop IMU integration without bias estimation applied to visual-inertial motion tracking. Outcome: unstable. Reason: sensor drift and integration errors compound with lower frame rates and lower-quality inertial sensors

Multi-modal 3D Gaussian Splatting for SLAM · UT Austin

Tried and failed

Doppler-inertial sensor fusion applied to satellite constellation radio navigation. Outcome: unstable. Reason: insufficient line-of-sight velocity vector diversity from co-aligned polar orbit trajectories

Navigation using Radio-Frequency Observables from LEO Constellations with Possible Aiding from an Inertial Navigation System · Virginia Tech

Tried and failed

direct LiDAR-inertial odometry fusion without redundancy applied to aerial state estimation. Outcome: unstable. Reason: intermittent sensor packet loss caused catastrophic estimation divergence without secondary vision-inertial fallback

Semantics-Driven Active Perception and Navigation with Aerial Robots · Penn

Tried and failed

standard complementary and Kalman filtering applied to low-cost mobile inertial sensors. Outcome: worse than baseline. Reason: gyroscope drift and high accelerometer noise during dynamic vehicle motions corrupted orientation estimation

Crowd-sourced Road Geometry and Accurate Vehicle State Estimation Using Mobile Devices · DSpace at SUNY Buffalo

Tried and failed

rolling statistics and sensor fusion orientation features applied to inertial sensor activity classification. Outcome: worse than baseline. Reason: engineered features degraded deep temporal model classification performance compared to raw sensor data

Reliable intelligence in precision agriculture: Quantifying uncertainty with evidential deep learning across animal and plant systems · Iowa State

Tried and failed

rigidly mounted visual-inertial sensors applied to dynamic robotic trajectory tracking. Outcome: worse than baseline. Reason: dynamic maneuvers induced severe motion blur and higher tracking error compared to active mechanical stabilization

Calibration and estimation for aerial robots with application to additive manufacturing · Imperial

Tried and failed

Track-to-track fusion using covariance intersection applied to multimodal multi-sensor object tracking. Outcome: worse than baseline. Reason: extreme divergence between individual sensor track error and estimated covariance

Radar and LiDAR Fusion for Scaled Vehicle Sensing · Virginia Tech

Lost to a baseline

Global-odometry (EKF fusion) had a higher standard deviation of absolute pose error (0.5667 m) compared to RTAB-Map-odometry (0.5367 m)

Autonomous localization and navigation for a railway inspection and repair system · Cranfield

Considered and rejected

Considered and rejected: Rejected inclusion of magnetometers in IMU sensor fusion due to distortion from ferromagnetic industrial equipment and the robot itself

Human upper body motion tracking for human-machine interaction in industrial applications · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Rejected IMU yaw orientation angles and thigh acceleration signals from multi-modal GRF estimation pipelines due to sensor drift and packet loss across subjects.

Experimental and Analytical Methods for Understanding User Gait Biomechanics in Response to Soft Ankle Exosuits · Harvard

Fusion algorithms and aggregation rules degrade under outliers, process noise, or sensory corruption

12 theses · 8 institutions

Common fusion rules such as sample mean aggregation or modular filtering break down when individual sensor measurements contain biases or non-zero mean noise. Centralized fusion architectures also experience severe detection failures when sensors encounter adversarial blinding attacks or severe domain distribution shifts.

Tried and failed

global sample mean aggregation applied to redundant sensor fusion with bias. Outcome: worse than baseline. Reason: vulnerable to outlier or biased sensor measurements without robust filtering

Statistical methods to obtain accurate estimates of measured process variables for redundant sensor measurements · Iowa State

Tried and failed

averaging closest pair of redundant estimates applied to redundant sensor data fusion. Outcome: worse than baseline. Reason: yielded higher mean squared error than using the simple sample median

Statistical methods to obtain accurate estimates of measured process variables for redundant sensor measurements · Iowa State

Tried and failed

past traversal visibility fusion offline adaptation applied to unsupervised 3D LiDAR object detection. Outcome: worse than baseline. Reason: caused performance drops for visible cars and non-occluded pedestrians relative to the baseline

PERCEPTION FOR AUTONOMOUS VEHICLES IN CHALLENGING WEATHERS AND OCCLUDED ENVIRONMENTS · Cornell

Lost to a baseline

Modular Bayesian fusion achieved lower accuracy than non-modular combined KNN when all sensors were fully available and trained together (accuracy ratio ~0.8-0.9).

Intent sensing for assistive technology · Oxford

Lost to a baseline

FAIEKF lost to FAUKF in multi-sensor fusion innovation sequence drift and accuracy under non-zero mean process noise (mu = 1 m).

Vision based Real-Time Navigation with Unknown and Uncooperative Space Target · Carleton University Institutional Repository

Lost to a baseline

Concatenated early multimodal feature fusion (AUROC 0.53–0.63) performed worse than unimodal HMM heart rate dynamics alone (AUROC 0.68–0.75) and late fusion (AUROC 0.72–0.82)

Multimodal assessment of neuropsychiatric disorders using audiovisual recordings · Georgia Tech

Lost to a baseline

Existing sensor fusion algorithms (F-PointNet, MV3D, AVOD) experienced perception detection failures across 98% to 99% of bundles under camera blinding attacks, and 79% to 84% under LIDAR rotation error attacks.

Secure and reliable deep learning in signal processing · Virginia Tech

Considered and rejected

Considered and rejected: Centralized sensor fusion architectures rejected in favor of federated architectures due to vulnerability/lower resilience to single-sensor failures

Robust autonomous navigation for UAVS in urban environments using machine learning · Cranfield

Considered and rejected

Considered and rejected: Rejected raw data-level 1-D stacking/matrix fusion without dimensionality reduction due to lack of error correction and inability to handle asymmetric/heterogeneous multi-modal sensors.

Development of deep meta-learning framework for cross-domain multisensory systems · Georgia Tech

Tried and failed

multimodal video feature fusion applied to cross-dataset persuasion strategy prediction. Outcome: did not generalise. Reason: large visual domain gaps between datasets impaired out-of-domain performance

Multimodal Human Behavior Modeling: From Understanding to Generation · Georgia Tech

Tried and failed

adding perceptual loss to multimodal generative model applied to sequential view synthesis from incomplete point clouds. Outcome: did not generalise. Reason: perceptual loss caused RGB-depth inconsistency, degrading multi-step rollout quality despite better single-step visual fidelity

Towards multi-modal AI systems with open-world cognition · Georgia Tech

Tried and failed

multimodal training with mutual information loss applied to cross-modality deformable image registration. Outcome: worse than baseline. Reason: unified cross-modality models underperformed compared to training separate modality-specific models

Deep Learning Methods to Process and Analyse MRI Images · Cornell

Tried and failed

multimodal context fusion with noisy visual features applied to human motion prediction. Outcome: did not converge. Reason: noisy background visual signals and inaccurate automated pose annotations caused training divergence

On the motion and action prediction using deep graph models · UT Austin

Multi-sensor integration fails when measurements suffer from timing mismatches, spatial misalignment, and association errors

6 theses · 3 institutions

Fusing observations across asynchronous or spatially separated sensors degrades tracking accuracy when timestamps and arrival rates do not align. Underestimated timing offsets, association errors, and view registration discrepancies result in the integration of out-of-sync or misaligned data.

Tried and failed

late-stage Kalman filter multi-sensor fusion applied to collaborative multi-object tracking. Outcome: worse than baseline. Reason: noisy local sensor measurements corrupted precise communicated baseline data during fusion

Enhancing Perception Systems using V2V Sensor Fusion · Virginia Tech

Tried and failed

synchronous sensor fusion state estimation applied to multi-sensor 3D state tracking. Outcome: worse than baseline. Reason: measurement arrival rate differences and timestamp mismatches across sensors degraded accuracy

3D INDOOR STATE ESTIMATION FOR RFID-BASED MOTION-CAPTURE SYSTEMS · Georgia Tech

Tried and failed

temporal multi-frame fusion with ego-motion odometry applied to monocular 3D object pose estimation. Outcome: worse than baseline. Reason: sparse low-frequency keypoint tracking and ego-motion noise degraded performance compared to single-frame geometric priors

Reshaping Perception for Autonomous Driving with Semantic Keypoints · EPFL

Tried and failed

Multi-sensor fusion across overlapping peripheral fields of view applied to Cross-traffic object state estimation. Reason: Sensor alignment and detection association failures degraded position and velocity estimation in lateral viewing angles

Assessing Effects of Object Detection Performance on Simulated Crash Outcomes for an Automated Driving System · Virginia Tech

Tried and failed

search window expansion for synchronization error applied to distributed sensor fusion. Outcome: worse than baseline. Reason: underestimating timing offsets causes fusion of corrupted, out-of-sync measurements as reliable data

Statistical Multistatic Radar with Imperfect Time Synchronization · Virginia Tech

Considered and rejected

Considered and rejected: Multi-sensor coordination frameworks: rejected due to complex cross-sensor time synchronization, bias estimation latencies, and registration errors.

Hypersonic: Real-Time Software Architecture for EM-based Radar Signal Processing and Tracking · Georgia Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Benchmark MW-FGO sensor fusion against traditional Kalman filtering across various sliding window sizes using public pedestrian GNSS/IMU datasets and GTSAM. Blocker: None

Advanced Inertial/GNSS Sensor Fusion for Smart Devices Using Factor Graph Optimization · DeustoTeka

Left open

Develop multi-sensor fusion combining IMU and FSR sensor data to reduce false positives in real-time step detection. Blocker: Requires physical wearable hardware equipped with synchronized IMU and FSR sensors or proprietary dual-sensor gait data.

Rhythms in Motion: A Biofeedback-Driven Vibrotactile Wearable for Gait Entrainment Rooted in Polyrhythmic Analysis of Afro-Diasporic Music · Harvard

Left open

Develop a multi-sensor fusion system combining radar or optical sensors for coarse target acquisition with 3D LiDAR for precision tracking of small UAS. Blocker: Requires multi-modal physical sensor hardware (LiDAR, radar, optical cameras) and synchronized real-world UAS tracking data.

DETECTION OF SMALL UNMANNED AERIAL SYSTEMS USING A 3D LIDAR SENSOR · Calhoun

Left open

Integrate fault-tolerant mechanisms and sensor fusion into the EMMA autonomous driving framework to handle missing state information and sensor failures. Blocker: None

Collaborative and safe autonomous driving through multi-agent deep reinforcement learning · Imperial

Left open

Develop and evaluate end-to-end multi-modal sensor fusion architectures for vehicle crash prediction using CARLA simulation data. Blocker: None

Vehicle crash prediction using transformer networks · Institutional Repository University of Moratuwa

Left open

Incorporate machine learning algorithms into the multi-robot localization fault detection module to detect sensor faults and failure modes. Blocker: Lack of specific ML architecture, failure mode definitions, or baseline fault detection code/dataset

From Slip Estimation to Collaborative Localization of Multi-Robot Systems Using Multi-Sensor Networks · Carleton University Institutional Repository

Left open

Evaluate whether connected vehicle sensor data alone provides sufficient accuracy for agency pavement management decision-making. Blocker: Requires access to proprietary connected vehicle fleet data and DOT pavement management standards

Pavement Surface Characteristics Evaluation Using Vehicle-Based Data Collection · Virginia Tech

Left open

Benchmark alternative machine learning models against gradient boosting for IMU/GNSS sensor fusion during GNSS outages. Blocker: No specific machine learning architectures, target metrics, or standardized benchmark datasets are defined

Improved IMU/GNSS EKF fusion using Machine Learning · Carleton University Institutional Repository

Left open

Investigate using autonomous vehicles as mobile traffic sensors to estimate traffic state and replace loop detectors or cameras. Blocker: None

Modelling mixed traffic flow of autonomous vehicles and human-driven vehicles · Imperial

Left open

Evaluate vision-based tracking using multi-sensor driving data collected from Cornell's autonomous vehicle platform. Blocker: Requires private sensor data collected from the Cornell autonomous test vehicle

PERCEPTION FOR AUTONOMOUS VEHICLES IN CHALLENGING WEATHERS AND OCCLUDED ENVIRONMENTS · Cornell

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.