Hasty Briefsbeta

Bilingual

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

6 hours ago
  • ShapeLearn-Lite was released quickly with a smaller optimization budget, but full ShapeLearn models now outperform it across all tested GPUs.
  • The full ShapeLearn models set a new quality-speed frontier, with GPU-5 (IQ4_XS) reaching 99.63% of BF16 performance and recommended as default when memory permits.
  • ShapeLearn-Lite performed better in benchmark tasks than its KLD rankings suggested; KLD does not always correlate with task performance.
  • Speculative decoding with MTP or DFlash2 increases throughput on all ShapeLearn models; DFlash2 is faster but requires more memory and lacks image input support.
  • Benchmarking covered multiple GPUs (96GB to 16GB) and showed consistent ordering: larger models yield higher accuracy, smaller models yield higher throughput.