Chapter Four · failure evidence

What Model Pruning & Compression got wrong, from 60 dissertations

The records document empirical evaluations of model pruning, quantization, knowledge distillation, and low-rank compression across neural network architectures. Across these experiments, techniques frequently degraded model accuracy or failed to yield computational speedups when constrained by rigid structural patterns, capacity mismatches, or hardware execution limits. These records come from PhD theses at 20 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Knowledge distillation fails when teachers and students suffer from capacity gaps or misaligned supervision

17 theses · 12 institutions

Large capacity mismatches between teacher and student models lead to severe overfitting, while low-performing or overconfident teachers fail to provide useful representation guidance. Intermediate layer matching, rigid loss weighting, and distillation objectives in continual learning often penalize necessary updates and underperform standard cross-entropy training baselines.

Tried and failed

full model knowledge distillation applied to deep neural network compression. Reason: compounded gradient interference when distilling earlier layers led to diminishing accuracy returns

Everything Is a Matrix: Minimizing Data Movement and Parameter Count Across the Machine Learning Stack · Harvard

Tried and failed

layerwise hidden representation matching in knowledge distillation applied to cross-architecture language model compression. Outcome: worse than baseline. Reason: unfiltered layerwise intermediate state matching overly constrained student learning compared to output probability distillation

On Parameter Efficiency of Neural Language Models · Georgia Tech

Tried and failed

knowledge distillation loss with exemplar replay applied to continual learning with recurring classes. Outcome: worse than baseline. Reason: distillation penalizes updating representations when classes reappear across learning sessions with replay

Lifelong Machine Learning with Data Efficiency and Knowledge Retention · EPFL

Tried and failed

knowledge distillation using sample-dependent eigenbases applied to feature extractor compression. Outcome: did not generalise. Reason: sample-dependent eigenbases produce a sub-optimal representation space that fails to transfer across diverse inputs

Lightweight model for content-style balanced photorealistic style transfer · UT Austin

Tried and failed

knowledge distillation from large teacher models applied to quantized model compression on small datasets. Outcome: overfit. Reason: Teacher-student capacity mismatch caused severe overfitting on small training datasets

Energy-efficient Neuromorphic Computing for Resource-constrained Internet of Things Devices · Virginia Tech

Tried and failed

knowledge distillation from low-performing teacher models applied to compact neural network training. Outcome: no signal. Reason: student networks cannot acquire useful representations when the teacher model has poor performance

From Symbolic Reasoning to Object Embeddings: Advanced Approaches of Knowledge Distillation in Compacted Neural Networks · Texas Tech

Tried and failed

distillation using robustly regularized teacher models applied to adversarially robust knowledge distillation. Outcome: worse than baseline. Reason: trade-off regularized teachers provided suboptimal supervision for student accuracy and robustness

The influence of data on adversarial robustness · EPFL

Tried and failed

fixed loss weighting in knowledge distillation applied to multimodal cross-modal knowledge distillation. Outcome: overfit. Reason: static distillation balance weights caused severe overfitting during knowledge transfer

Effective deep leaning methodologies for salient object detection · Imperial

Tried and failed

knowledge distillation loss applied to continual learning for image classification. Outcome: worse than baseline. Reason: degraded performance and hindered the ability to benefit from repeated exposures to classes

Shape-biased representations for object category recognition · Georgia Tech

Tried and failed

knowledge distillation loss for guidance applied to long-term time series forecasting. Outcome: worse than baseline. Reason: distillation loss degraded guidance from the source teacher model over long prediction horizons

Artificial Intelligence for Data-centric Surveillance and Forecasting of Epidemics · Georgia Tech

Tried and failed

knowledge distillation applied to conditionally gated convolutional neural networks. Outcome: no signal. Reason: small difference between ground truth and teacher output made distilled loss ineffective

Algorithm-Accelerator Co-Design for High-performance and Secure Deep Learning · Cornell

Tried and failed

knowledge distillation of fine-tuned language models applied to clinical concept presence prediction. Outcome: worse than baseline. Reason: None

Applying Language Models To Patient Health Records: Acronym Expansion, Long Document Classification and Explainable Predictions · Penn

Tried and failed

knowledge distillation with a large teacher model applied to image classification on complex datasets. Reason: large capacity gap between teacher and student limited effective knowledge transfer

From Symbolic Reasoning to Object Embeddings: Advanced Approaches of Knowledge Distillation in Compacted Neural Networks · Texas Tech

Tried and failed

knowledge distillation using higher-accuracy teacher model applied to compact object detection student network. Outcome: worse than baseline. Reason: teacher overconfidence degraded distillation quality compared to a smaller teacher model

Improving the Training of Compact Neural Networks for Visual Recognition · EPFL

Lost to a baseline

On CIFAR-10 distillation with 5000 samples/class, baseline cross-entropy student training reached 91.95%, beating FullGrad + Attention distillation (90.68%).

Gradient-based Methods for Deep Model Interpretability · EPFL

Lost to a baseline

On CIFAR-10 with standard training under full regularization (dcw), standard training (95.0% accuracy) beat Shrink & Perturb with distillation (94.7% accuracy).

On architectures and training techniques for neural networks · Oxford

Considered and rejected

Considered and rejected: Knowledge distillation: rejected in favor of domain-specific parameter pruning due to difficulties in distilling fine regulatory reasoning into small student models.

Evaluating Efficiency Gains and Security of LLM-Driven Test Generation for Computerised System Validation: A Compliance-Focused Analysis of Life Sciences Testing Processes · DSpace at Griffith College

Considered and rejected

Considered and rejected: Rejected traditional logit-based KL divergence distillation because fine-tuned teacher logits were noisy/imprecise and diluted binary classification signals.

Distilled model for contextual soft moderation on social media · OpenBU

Structured and filter pruning enforcing rigid constraints degrades accuracy and loses to simpler baselines

11 theses · 7 institutions

Imposing rigid N:M patterns, geometric filter constraints, or channel pruning inadvertently eliminates critical weights and destroys reasoning capabilities or representation quality. In several architectures and low-density regimes, these structured pruning heuristics underperform unstructured magnitude pruning and simple random pruning baselines.

Tried and failed

fixed N:M structured sparsity pruning applied to large language models. Outcome: worse than baseline. Reason: rigid ratio constraints inadvertently remove critical outlier weights, degrading model perplexity

Algorithm–Hardware Co-Design Of Digital Compute-In-Memory Architecture Supporting Flexible And Temporal N:M Sparsity · Georgia Tech

Tried and failed

structured filter pruning applied to convolutional neural network compression. Outcome: worse than baseline. Reason: degrades model accuracy significantly compared to fine-grained N:M structured sparsity patterns

Scheduling Algorithms for Low-Precision Accumulation on Energy-Efficient Deep Neural Network Hardware · Harvard

Tried and failed

post-training pruning to structured N:M sparsity applied to large language models. Outcome: worse than baseline. Reason: structured sparsity constraints severely degraded performance on knowledge-intensive downstream tasks

Chasing efficiency in the era of large language models · UT Austin

Tried and failed

pseudorandom structured pruning with strict geometric constraints applied to convolutional neural networks on ImageNet. Outcome: worse than baseline. Reason: overly rigid geometrical constraints caused severe accuracy degradation at high sparsity levels

Efficient processing of vision kernels and deep neural networks on reconfigurable computing architectures · Iowa State

Tried and failed

structured weight pruning applied to large language models for reasoning. Outcome: worse than baseline. Reason: removing large fractions of structural weights destroys complex reasoning capabilities and output coherence

Enabling Small Language Models as Efficient and Capable Agents · Virginia Tech

Tried and failed

structured N:M magnitude-based weight sparsification applied to large language model compression. Outcome: worse than baseline. Reason: magnitude pruning alone severely degrades model perplexity without post-training fine-tuning or decomposition

Structured Sparsity-Aware Hardware-Software Co-Design for Deep Neural Network Acceleration · Georgia Tech

Tried and failed

activity-based structured state pruning applied to state space sequence models. Outcome: worse than baseline. Reason: aggressive pruning ratios cause non-linear accuracy collapse while latency gains plateau

Performance profiling and activity-based state pruning for efficient Mamba inference · Iowa State

Lost to a baseline

Structured pruning (e.g., pruning full convolutional filters/channels) loses 1 or more percentage points of accuracy compared to unstructured magnitude pruning at 60% weights remaining on ResNet-50 on ImageNet.

The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT

Considered and rejected

Considered and rejected: Rejected structured N:M sparsity patterns in favor of unstructured pruning because unstructured pruning performed better and proved more robust.

Odors as ''natural language'': sparse neural networks in mammalian olfactory systems and large language models · Harvard

Lost to a baseline

Stat channel pruning on Resnet-20 performed worse than random pruning across most preservation rates.

Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace

Lost to a baseline

At very low channel preservation rates on Resnet-20, intelligent pruning methods (SNIP, SNat, Stat) were outperformed by random pruning at initialization.

Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace

Lost to a baseline

On DBSN node pruning at low preservation rates, intelligent pruning methods dropped below random pruning due to over-pruning the final bottleneck layer.

Fine Granularity is Critical for Intelligent Neural Network Pruning · YorkSpace

Left open

Investigate why structured pruning underperforms random initialization on larger networks trained with analog noise on MNIST. Blocker: None

Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech

Pruning at initialization or early in training underperforms gradual post-training pruning and random baselines

6 theses · 4 institutions

One-shot or early magnitude pruning methods inflict severe representation damage and can trigger catastrophic layer collapse. Across multiple benchmarks, post-training pruning and gradual iterative schedules consistently outperform initialization pruning, which often fails to beat random sparse initializations.

Tried and failed

iterative magnitude pruning without retraining applied to random feature models. Outcome: worse than baseline. Reason: fixed weights without re-optimization degraded representation quality compared to the unpruned minimal L2-norm solution

Adaptive and weighted optimization for efficient and robust learning · UT Austin

Tried and failed

uninformed magnitude pruning with weight regeneration applied to large language models. Outcome: worse than baseline. Reason: causes irreparable knowledge damage on complex tasks that weight regeneration fine-tuning cannot recover

Chasing efficiency in the era of large language models · UT Austin

Lost to a baseline

Magnitude pruning after training outperformed all initialization pruning methods including PHEW across high-sparsity regimes.

Leveraging sparsity in deep neural networks for training efficiency, interpretability and generalization · Georgia Tech

Lost to a baseline

Global random pruning and random reinitialization match or beat IMP at step 0 on standard ResNet-20 (88.6% and 88.8% vs 88.5% test accuracy at 16.8% density).

The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT

Lost to a baseline

Repeated full pruning and gradual pruning were beaten by simple random sparse initialization on large MNIST networks (2-3 hidden layers, 200 neurons/layer)

Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech

Considered and rejected

Considered and rejected: Rejected one-shot magnitude pruning in favor of iterative magnitude pruning because gradual pruning yields higher accuracy and sparser matching subnetworks.

The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks · MIT

Lost to a baseline

On Tiny-ImageNet, Initial (Weight) Magnitude Pruning suffered catastrophic accuracy drops due to layer collapse at lower network densities compared to PHEW, SynFlow, and SynFlow-L2.

Leveraging sparsity in deep neural networks for training efficiency, interpretability and generalization · Georgia Tech

Lost to a baseline

One-time pruning underperformed random sparsity initialization on small neural networks across multiple wait times

Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech

Lost to a baseline

On CIFAR-10 ResNet-50 pruning, 1-epoch early magnitude pruning performed worse than SNIP-MB and gradual pruning methods.

Towards efficient and reliable neural networks · Oxford

Theoretical model sparsity and compression fail to translate into practical latency reductions or hardware speedups

9 theses · 6 institutions

Unstructured weight pruning and compressed representations frequently fail to reduce execution time because sparse execution routines and metadata overheads dominate system latency. Additionally, complex pruning procedures or iterative factorizations introduce severe memory and computational overheads that exceed any downstream transmission or evaluation benefits.

Tried and failed

post-training channel pruning without fine-tuning applied to convolutional neural networks. Outcome: worse than baseline. Reason: severe accuracy degradation exceeding ten percentage points with negligible throughput improvement under hardware execution

Efficient AI model acceleration through tensor decomposition and dataflow architecture design · Imperial

Considered and rejected

Considered and rejected: Rejected unstructured pruning for learning compact hidden representations because pruned sparse weights do not reduce the number of features or inference overhead.

Achieving More with Less: Learning Generalizable Neural Networks With Less Labeled Data and Computational Overheads · Virginia Tech

Considered and rejected

Considered and rejected: Unstructured network pruning: rejected in favor of structured pruning to avoid requiring specialized sparse hardware accelerators and transmission of sparse structural metadata.

Deep joint source-channel coding for vision-based inference at the wireless edge · Imperial

Tried and failed

run-length encoding metadata compression applied to sparse tensor accelerator bandwidth optimization. Reason: bandwidth overhead is dominated by data payload rather than metadata encoding size

Systematic Modeling and Design of Sparse Deep Neural Network Accelerators · MIT

Tried and failed

weight pruning applied to bound propagation neural network verification. Outcome: too slow. Reason: Sparsity does not yield computational speedups in symbolic interval propagation routines despite reducing constraint complexity in SMT/MILP

Robust neural networks: verification, training, and repair · Imperial

Lost to a baseline

Embedded-ViT with 50% pruning at 128x128 input resolution on CPU achieved lower throughput (23.36 FPS) than unpruned DAE-Former baseline (31.32 FPS)

Weakly supervised and embedded semantic segmentation for Computer Aided Diagnostics · DSpace-CRIS at TU Wien

Lost to a baseline

The baseline accelerator from [29] is ×1.85 faster than the proposed bit-level pruning design due to its dataflow latency optimization.

Design of Efficient DNN Accelerator Architectures · HARVEST

Considered and rejected

Considered and rejected: Rejected dynamic pruning during local training because resource-constrained edge clients cannot train dense unpruned DNNs initially.

REFT: Resource-Efficient Federated Training Framework for Heterogeneous and Resource-Constrained Environments · Virginia Tech

Considered and rejected

Considered and rejected: Rejected alternating iterative SVD updates and outlier extraction during KV cache compression due to severe latency overheads unacceptable for generative inference.

On the Efficiency and Steerability of Self-Attention Mechanism of Large Language Models · Georgia Tech

Uniform and layer-agnostic pruning allocations cause severe representation collapse

8 theses · 7 institutions

Applying uniform pruning budgets or uniform low-rank matrix approximations across all layers ignores the varying sensitivity and compressibility of different architectural components. Indiscriminately removing weights from early residual blocks, skip connections, or dense language model layers disproportionately degrades critical representations and hurts downstream perplexity.

Tried and failed

post-training magnitude pruning on learned skip-connection weights applied to inter-layer depth-wise averaging modules. Outcome: worse than baseline. Reason: severely degrades model perplexity even at low pruning sparsity thresholds

Enhanced Architectures and Optimization Methods for Efficient Language Modeling · EPFL

Tried and failed

magnitude pruning applied to state space models. Outcome: worse than baseline. Reason: architecture was highly sensitive to parameter pruning with significant performance drops at moderate sparsity

Resource-efficient AI for real-world applications: Model compression, sample-efficient learning, and inference optimization across traffic safety, quantum systems, and sequential d · Iowa State

Tried and failed

uniform layer-wise pruning budget allocation applied to deep neural network compression. Outcome: worse than baseline. Reason: fails to account for non-uniform layer compressibility and importance, yielding suboptimal accuracy-compression trade-offs

Efficient Deep Learning: From Theory to Practice · MIT

Tried and failed

simultaneous pruning of early residual blocks applied to deep neural network compression. Outcome: worse than baseline. Reason: removes critical low-level representations leading to severe performance degradation

Sparsity prior in efficient deep learning based solvers and models · UT Austin

Tried and failed

backward or parallel layer-wise pruning applied to neural network model compression. Outcome: worse than baseline. Reason: deep-to-shallow or simultaneous pruning degraded accuracy compared to forward layer-by-layer pruning

Overcoming Noise and Variations In Low-Precision Neural Networks · Georgia Tech

Tried and failed

uniform low-rank matrix approximation across all layers applied to deep neural network weight matrices. Outcome: worse than baseline. Reason: indiscriminate rank reduction harms critical representations compared to targeted layer-specific pruning

Discovering and Engineering the Computation Underlying Large Intelligent Agents · MIT

Tried and failed

tensor-train decomposition for neural network compression applied to dense language model layers. Outcome: worse than baseline. Reason: degrades broad, weak lower-level feature representation required for commonsense reasoning tasks

Generative language model compression with algebraic approaches · Imperial

Tried and failed

direct computation cost loss penalty applied to dynamic neural network channel pruning. Reason: disproportionately pruned higher-FLOP layers, causing severe layer-wise imbalance compared to threshold loss

Algorithm-Accelerator Co-Design for High-performance and Secure Deep Learning · Cornell

Quantization induces precision errors and disrupts weight ranking when combined with pruning

7 theses · 6 institutions

Aggressive sub-4-bit quantization and post-training integer rounding create substantial numerical errors that severely degrade model accuracy without extensive distillation. In addition, quantizing before magnitude pruning perturbs weight values, which disrupts the relative magnitude ranking required for effective parameter pruning.

Tried and failed

post-training quantization applied to large language models across task difficulties. Outcome: did not generalise. Reason: performance degradation was non-monotonic across varying task difficulty levels compared to pruning

Chasing efficiency in the era of large language models · UT Austin

Tried and failed

quantization before magnitude pruning applied to neural network model compression. Outcome: worse than baseline. Reason: quantization perturbs weight magnitudes, disrupting pruning order and degrading perplexity

Compressing DNNs Using Microscaling Formats with Sensitivity and Sparsity · EPFL

Tried and failed

post-training quantization without knowledge distillation applied to deep neural networks. Outcome: worse than baseline. Reason: aggressive low-bit term quantization severely degrades accuracy without multi-resolution distillation guidance

Systolic Architectures for Efficient Deep Neural Network Implementations with Assured Performance · Harvard

Tried and failed

extreme sub-4-bit iterative weight quantization hierarchy applied to deep neural network weight compression. Outcome: worse than baseline. Reason: drastic rounding errors at 1-bit and 2-bit precisions severely degrade network performance

Advancing efficiency and trustworthiness : from computer vision to multimodal large language models · UT Austin

Considered and rejected

Considered and rejected: Rejected post-training integer quantization (INT8/INT4) in favor of FP16/FP32 encoder pruning to avoid accuracy degradation and retraining/calibration overhead.

On-device Learning and Inference Optimization for Lightweight Neural Networks and Transformers on Microcontrollers · IRIS - POLITO - prod

Considered and rejected

Considered and rejected: Quantization-aware training (QAT) was considered for edge model compression but not used in the thesis, using post-training quantization instead.

Robust Machine Learning Against Faults in Micro-Controllers and Stragglers in Distributed Training on the Cloud · Virginia Tech

Considered and rejected

Considered and rejected: Rejected conventional bit packing because the wide dynamic range of ML values provides minimal compression across individual tensors.

Mitigating the Impact of Data Movement in Memory-Intensive Applications · ResearchWorks

Left open by the authors

Problems the authors named and did not get to.

Left open

Analyze why BERT pruning algorithms naturally prune more neurons in deeper layers than earlier layers. Blocker: None

Sparsity prior in efficient deep learning based solvers and models · UT Austin

Left open

Develop model compression techniques using highly sparse binary masks to reduce memory footprint and inference latency without degrading accuracy. Blocker: None

Algorithms for Efficient and Robust Distributed Deep Learning · EPFL

Left open

Develop pruning techniques to recover accuracy for deep neural network models with dual-side structured sparsity (HighLight_DSSO). Blocker: Lacks specific targets, pruning algorithms, or designated network architectures to evaluate

Systematic Modeling and Design of Sparse Deep Neural Network Accelerators · MIT

Left open

Tune training hyperparameters to quantify the workload required to recover accuracy after deep learning model pruning. Blocker: None

Tools for efficient Deep Learning · Imperial

Left open

Apply pruning, feature importance, and complexity analysis methods to CNN models representing binary nanophotonic structures. Blocker: None

Machine Learning Approaches for Knowledge Discovery in Nanophotonic Structures · Georgia Tech

Left open

Implement layer-wise pruning on 4D CNNs to evaluate temporal feature retention and parameter reduction in deep layers. Blocker: None

Development of a machine-learning platform for the autonomous analysis of 3D+T calcium imaging data · UT Austin

Left open

Extend the pruning and feature importance framework to CNN models trained on binary nanophotonic structure electromagnetic simulation data. Blocker: None

Machine Learning Approaches for Knowledge Discovery in Nanophotonic Structures · Georgia Tech

Left open

Apply advanced neural network pruning to reduce connection complexity between design and response spaces in the nanophotonic inverse design framework. Blocker: None

A New Paradigm for Knowledge Discovery and Design in Nanophotonics Based on Artificial Intelligence · Georgia Tech

Left open

Apply pruning, quantization, and edge hardware acceleration to the STFT-CNN power quality disturbance classification model. Blocker: None

Power Quality Disturbance Classification in IEEE 9-Bus System using STFT and Deep Learning · Texas Tech

Left open

Apply pruning and quantization to the autoencoder-compressed transformer summarization pipeline and evaluate size-accuracy trade-offs. Blocker: None

Efficient and Enhanced Text Summarization by Compressing and Data Augmentation for Transformers-Based Models · Scholarship at UWindsor Institutional Repository

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.