Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
9 hours ago
- Post-training quantization (PTQ) reduces accuracy while increasing chain-of-thought (CoT) length in reasoning models across math, coding, and science QA tasks.
- In up to 52% of quantized model failures, the correct answer appears in intermediate reasoning steps but not as the final output, indicating overthinking errors.
- High token-level KL divergence between quantized and full-precision models at high-entropy positions leads quantized models to sample overthinking markers like 'wait', 'but', and 'alternatively'.
- A training-free logit penalty on curated overthinking markers reduces CoT length by 12–23% and overthinking errors by up to 58%, while preserving or improving accuracy across multiple models and benchmarks.