Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Abstract
Debias-SparseGPT reduces pruning-induced demographic bias in compressed large language models by incorporating representational debiasing during post-training sparsification.
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.
Community
In this paper, we introduce Debias-SparseGPT, a pruning-time debiasing compression method that explicitly incorporates group differences into the compression objective.
The method is motivated by an extensive body of work showing that aggressive pruning can amplify toxic generations and stereotypical biased outputs after compression.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression (2026)
- Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed (2026)
- Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models (2026)
- F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Language Models (2026)
- The Sparsity Whisperer (2026)
- PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning (2026)
- Unified Static-Dynamic Pruning for Efficient LLM Inference (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.02496 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper