Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
6 hours ago
- ShapeLearn-Lite was released quickly with a smaller optimization budget, but full ShapeLearn models now outperform it across all tested GPUs.
- The full ShapeLearn models set a new quality-speed frontier, with GPU-5 (IQ4_XS) reaching 99.63% of BF16 performance and recommended as default when memory permits.
- ShapeLearn-Lite performed better in benchmark tasks than its KLD rankings suggested; KLD does not always correlate with task performance.
- Speculative decoding with MTP or DFlash2 increases throughput on all ShapeLearn models; DFlash2 is faster but requires more memory and lacks image input support.
- Benchmarking covered multiple GPUs (96GB to 16GB) and showed consistent ordering: larger models yield higher accuracy, smaller models yield higher throughput.