Standard

SPSA View on the Straight-Through Estimator in Neural Network Quantization. / Salishev, Sergey; Makarov, Anton; Tarasova, Elizaveta; Shishin, Kirill; Granichin, Oleg.

In: IEEE Access, Vol. 14, 13.04.2026, p. 60047 - 60057.

Research output: Contribution to journalArticlepeer-review

Harvard

APA

Vancouver

Author

Salishev, Sergey ; Makarov, Anton ; Tarasova, Elizaveta ; Shishin, Kirill ; Granichin, Oleg. / SPSA View on the Straight-Through Estimator in Neural Network Quantization. In: IEEE Access. 2026 ; Vol. 14. pp. 60047 - 60057.

BibTeX

@article{a6b10cc36c3c4c64a7440268b1033e12,
title = "SPSA View on the Straight-Through Estimator in Neural Network Quantization",
abstract = "Low-bit quantization-aware training of neural networks typically relies on the straight-through estimator (STE) to learn both quantized weights and their associated scales or effective bit-widths. Building on earlier analysis of gradual differentiable noise-scale quantization, we first recall that the classical STE noise-scale gradient can be interpreted as an asymptotically optimal estimator of the rounding-noise gradient obtained by averaging over mini-batches. This viewpoint justifies replacing the exact rounding residual by any zero-mean proxy noise with matched variance, such as symmetric Bernoulli (Rademacher), Uniform or Gaussian perturbations. In this paper we make the connection between such STE-based scale updates and simultaneous perturbation stochastic approximation (SPSA) precise. We show that injecting noise into the learnable quantization scale implements a one-measurement SPSA scheme in the low-dimensional noise-scale space, and we verify the standard stochastic-approximation assumptions for a realistic quantization-aware loss with a bit-width penalty. Experiments with one-bit weight and activation quantization on ResNet-20 for CIFAR-10/100 classification, and four-bit quantization on RFDN for standard super-resolution (SR) benchmarks, demonstrate that the stochastic scale update consistently accelerates convergence of the effective bit-width while preserving final accuracy compared to empirical rounding. In a PyTorch implementation this modification requires only one additional line on top of a standard LSQ-type quantization-aware training loop, adding negligible computational and memory overhead.",
keywords = "efficient deep learning, gradient estimation, learned step size quantization, low-bit neural networks, noise-scale learning, quantization-aware training, simultaneous perturbation stochastic approximation, straight-through estimator",
author = "Sergey Salishev and Anton Makarov and Elizaveta Tarasova and Kirill Shishin and Oleg Granichin",
year = "2026",
month = apr,
day = "13",
doi = "10.1109/ACCESS.2026.3683267",
language = "English",
volume = "14",
pages = "60047 -- 60057",
journal = "IEEE Access",
issn = "2169-3536",
publisher = "Institute of Electrical and Electronics Engineers Inc.",

}

RIS

TY - JOUR

T1 - SPSA View on the Straight-Through Estimator in Neural Network Quantization

AU - Salishev, Sergey

AU - Makarov, Anton

AU - Tarasova, Elizaveta

AU - Shishin, Kirill

AU - Granichin, Oleg

PY - 2026/4/13

Y1 - 2026/4/13

N2 - Low-bit quantization-aware training of neural networks typically relies on the straight-through estimator (STE) to learn both quantized weights and their associated scales or effective bit-widths. Building on earlier analysis of gradual differentiable noise-scale quantization, we first recall that the classical STE noise-scale gradient can be interpreted as an asymptotically optimal estimator of the rounding-noise gradient obtained by averaging over mini-batches. This viewpoint justifies replacing the exact rounding residual by any zero-mean proxy noise with matched variance, such as symmetric Bernoulli (Rademacher), Uniform or Gaussian perturbations. In this paper we make the connection between such STE-based scale updates and simultaneous perturbation stochastic approximation (SPSA) precise. We show that injecting noise into the learnable quantization scale implements a one-measurement SPSA scheme in the low-dimensional noise-scale space, and we verify the standard stochastic-approximation assumptions for a realistic quantization-aware loss with a bit-width penalty. Experiments with one-bit weight and activation quantization on ResNet-20 for CIFAR-10/100 classification, and four-bit quantization on RFDN for standard super-resolution (SR) benchmarks, demonstrate that the stochastic scale update consistently accelerates convergence of the effective bit-width while preserving final accuracy compared to empirical rounding. In a PyTorch implementation this modification requires only one additional line on top of a standard LSQ-type quantization-aware training loop, adding negligible computational and memory overhead.

AB - Low-bit quantization-aware training of neural networks typically relies on the straight-through estimator (STE) to learn both quantized weights and their associated scales or effective bit-widths. Building on earlier analysis of gradual differentiable noise-scale quantization, we first recall that the classical STE noise-scale gradient can be interpreted as an asymptotically optimal estimator of the rounding-noise gradient obtained by averaging over mini-batches. This viewpoint justifies replacing the exact rounding residual by any zero-mean proxy noise with matched variance, such as symmetric Bernoulli (Rademacher), Uniform or Gaussian perturbations. In this paper we make the connection between such STE-based scale updates and simultaneous perturbation stochastic approximation (SPSA) precise. We show that injecting noise into the learnable quantization scale implements a one-measurement SPSA scheme in the low-dimensional noise-scale space, and we verify the standard stochastic-approximation assumptions for a realistic quantization-aware loss with a bit-width penalty. Experiments with one-bit weight and activation quantization on ResNet-20 for CIFAR-10/100 classification, and four-bit quantization on RFDN for standard super-resolution (SR) benchmarks, demonstrate that the stochastic scale update consistently accelerates convergence of the effective bit-width while preserving final accuracy compared to empirical rounding. In a PyTorch implementation this modification requires only one additional line on top of a standard LSQ-type quantization-aware training loop, adding negligible computational and memory overhead.

KW - efficient deep learning

KW - gradient estimation

KW - learned step size quantization

KW - low-bit neural networks

KW - noise-scale learning

KW - quantization-aware training

KW - simultaneous perturbation stochastic approximation

KW - straight-through estimator

UR - https://www.mendeley.com/catalogue/711e211d-bf31-3630-bca4-80c8a5c2e301/

U2 - 10.1109/ACCESS.2026.3683267

DO - 10.1109/ACCESS.2026.3683267

M3 - Article

VL - 14

SP - 60047

EP - 60057

JO - IEEE Access

JF - IEEE Access

SN - 2169-3536

ER -

ID: 152227010