R–D Perspective

A Unified Rate–Distortion Perspective

on Vector, Product, and Scalar Quantization

Xianghong Fang1 · Wenlong Mou1 · Yuan Yuan2 · Dehan Kong1 · Tim G. J. Rudner1,3

1University of Toronto   2Boston College   3Vijil

Q1 What should a quantization algorithm fundamentally optimize? Q2 What conditions are required for a fair intrinsic comparison of different quantization algorithms? Q3 Under these controlled conditions, how do VQ, PQ, and SQ compare in achievable distortion?

In our paper, we show that distortion is the primary objective, fair comparison requires matched latent distributions and rates, and modern VQ methods reach the lowest distortion under these controls.

Quantization is
lossy compression.

Discrete visual tokenizers compress continuous latent features into a finite set of symbols. Rate measures the representation budget. Distortion measures the information lost.

Nominal fixed-length rate
R = T log2 K

T is the token count and K is the discrete code-space cardinality.

Quantization distortion
\[\mathcal{E} = \mathbb{E}\!\left[\left\|X-\mathcal{Q}(X)\right\|_2^2\right]\]

Expected squared error measures the information removed by quantization.

This lens resolves three linked questions: what to optimize, how to compare quantizers, and which family has the strongest intrinsic rate–distortion behavior.

Overview of the unified rate-distortion perspective
Figure 1. The unified perspective. (a) Distortion links optimization to codebook use, straight-through estimator stability, and reconstruction fidelity. (b) Intrinsic comparisons must match the latent distribution and nominal rate. (c) The nested structure SQ ⊆ PQ ⊆ VQ yields 𝓔*VQ ≤ 𝓔*PQ ≤ 𝓔*SQ.

The three conclusions form one framework: optimize the quantity that measures information loss, control the factors that change it, then compare quantizer families at the same operating point.

Q1

Optimize distortion

Under mild conditions, every global distortion minimizer uses the full codebook. The converse fails: full utilization alone does not guarantee minimum distortion.

min 𝓔 ⇒ U = 1
but U = 1 ⇏ min 𝓔
Q2

Control the comparison

Latent scale changes squared distortion by a2. Token count and code-space size set the rate. Both must be matched to isolate the quantizer.

same P(X) · same T · same K
Q3

Respect the hierarchy

VQ quantizes vectors jointly. PQ factorizes them into blocks. SQ factorizes every coordinate. Each constraint narrows the set of possible solutions.

SQ ⊆ PQ ⊆ VQ

Distortion minimization is the fundamental criterion for intrinsic quantization effectiveness under a fixed rate.

Joint structure
beats fixed factorization.

VQ searches over arbitrary vector codebooks. PQ restricts each codeword to a Cartesian product of subcodebooks, while SQ applies that restriction coordinate by coordinate. The feasible sets are therefore nested.

When the source lies near a low-dimensional structure, VQ can adapt to its intrinsic dimension. Its distortion can scale as O(K−2/deff), while fixed PQ and SQ factorizations may remain tied to the ambient dimension.

VQjoint vector codebook
PQblock-factorized
SQcoordinate-factorized

Lower distortion,
better reconstruction.

We compare matched-rate quantizers across three latent-space datasets and a controlled CelebA-HQ pixel-space setting.

FamilyBest Distortion MethodBest rFID MethodDistortion ↓rFID ↓
Latent spaceImageNet-1K · T = 512 · K = 65,536
VQMMD VQMMD VQ0.2010.86
PQEMA VP2EMA VP20.2090.93
SQBSQBSQ0.2311.07
Latent spaceFFHQ · T = 512 · K = 65,536
VQMMD VQMMD VQ0.1310.85
PQEMA VP2EMA VP20.1361.05
SQBSQBSQ0.1551.54
Latent spaceCelebA-HQ · T = 512 · K = 65,536
VQWasserstein VQ / MMD VQWasserstein VQ0.1111.73
PQEMA VP2EMA VP20.1151.96
SQBSQFSQ0.1332.24
Pixel spaceCelebA-HQ · T = 4,096 · K = 65,536
VQMMD VQMMD VQ0.00213.64
PQOnline VP2 / MMD VP2MMD VP20.00233.77
SQFSQFSQ0.00324.54

Each metric reports the best result within a quantizer family. Latent-space rFID is measured after decoder adaptation; pixel-space models use single-stage training.

The family-level comparison stays consistent as the per-token coding budget changes.

Rate-distortion curves on ImageNet-1K
Figure 2. Rate–distortion curves on ImageNet-1K. Each point reports the lowest distortion among the evaluated methods in a family at the corresponding nominal per-token rate log2 K.

More rate lowers distortion for every family, yet the ordering persists at every evaluated point: VQ achieves the lowest distortion, followed by PQ and SQ.

Representation and DatasetDistortion vs. rFIDUtilization vs. rFID
Latent · ImageNet-1Kρ = 0.996p < 10−8ρ = −0.540p = 0.057
Latent · FFHQρ = 0.979p = 5.49 × 10−9ρ = −0.492p = 0.088
Latent · CelebA-HQρ = 0.944p = 1.29 × 10−6ρ = −0.559p = 0.047
Pixel · CelebA-HQρ = 0.947p = 2.91 × 10−6ρ = −0.413p = 0.182

Distortion tracks reconstruction fidelity far more closely than codebook utilization across representation spaces and datasets.

  1. Treat quantization as lossy compression with rate R = T log2 K and squared quantization error as distortion.
  2. Minimize distortion first, because optimal distortion implies full codebook use while full use does not imply optimal distortion.
  3. Match latent distributions, token counts, and composite code-space cardinalities before comparing quantizer families.
  4. Modern VQ methods achieve the lowest controlled distortion and exploit low-dimensional source structure that fixed factorizations can miss.

Build on our work.

View code ↗
@article{Fang2026ratedistortion,
  title   = {A Unified Rate--Distortion Perspective on Vector,
             Product, and Scalar Quantization},
  author  = {Fang, Xianghong and Mou, Wenlong and Yuan, Yuan and
             Kong, Dehan and Rudner, Tim G. J.},
  journal = {arXiv},
  year    = {2026}
}