martincousseau.com

Paper, IJCNN 2025, , 2 min read, Updated

Semantic Retention and Extreme Compression in LLMs: Can We Have Both?

IN SHORT

Introduces SrCr, a way to measure how much meaning survives compression.

One-bit dithered desert dunes, the cover image for the SrCr paper.
PLATE 1. Generated plate · 1-bit · dunes

01 · In brief

#

Summary

My co-authors and I introduced SrCr, the Semantic Retention Compression Rate: a way to measure how much meaning survives when a model is compressed.

SrCr is a metric that quantifies the trade-off between model compression and semantic preservation. In the paper we examine joint compression, that is, how strategically combining pruning and quantization can yield better performance-to-compression ratios than either method alone, and we use SrCr to optimize pruning-quantization configurations.

The abstract reports that the authors' recommended combination achieves, on average, a 20% performance increase compared to an equivalent quantization-only model at the same theoretical compression rate.

I co-authored the paper with Stanislas Laborde, Antoun Yaacoub and Lionel Prevost; I am the second of the four authors. Stanislas Laborde and I contributed equally, and I started the project. It was accepted at IJCNN 2025, the International Joint Conference on Neural Networks (Rome, Italy, 30 June to 5 July 2025), and posted to arXiv on 12 May 2025.

The measure also shaped how I think about evaluation more broadly. The companion note, What Survives Compression, covers that habit: ask what meaning survived before you ask how fluent the remainder sounds.

02 · The paper

#

Abstract

Abstract of the paper as posted on arXiv (2505.07289) by Stanislas Laborde, Martin Cousseau, Antoun Yaacoub and Lionel Prevost:

The exponential growth in Large Language Model (LLM) deployment has intensified the need for efficient model compression techniques to reduce computational and memory costs. While pruning and quantization have shown promise, their combined potential remains largely unexplored. In this paper, we examine joint compression and how strategically combining pruning and quantization could yield superior performance-to-compression ratios compared to single-method approaches. Recognizing the challenges in accurately assessing LLM performance, we address key limitations of previous evaluation frameworks and introduce the Semantic Retention Compression Rate (SrCr), a novel metric that quantifies the trade-off between model compression and semantic preservation, facilitating the optimization of pruning-quantization configurations. Experiments demonstrate that our recommended combination achieves, on average, a 20% performance increase compared to an equivalent quantization-only model at the same theoretical compression rate.

03 · Links

#

Read it

04 · Reference

#

Cite

@inproceedings{laborde2025semantic,
  author        = {Laborde, Stanislas and Cousseau, Martin and Yaacoub, Antoun and Prevost, Lionel},
  title         = {Semantic Retention and Extreme Compression in {LLMs}: Can We Have Both?},
  booktitle     = {2025 International Joint Conference on Neural Networks ({IJCNN})},
  year          = {2025},
  pages         = {1--9},
  address       = {Rome, Italy},
  publisher     = {IEEE},
  doi           = {10.1109/IJCNN64981.2025.11227279},
  eprint        = {2505.07289},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}