Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
4 hours ago
- Ternary Bonsai 2 27B is a compressed 27B multimodal model using ternary weights with 1.76 effective bits per weight and a 5.9GB footprint.
- It retains 98.2% of the full-precision Qwen3.8 27B performance while being over 9x smaller.
- It supports a 262K-token context window, multimodal input, and is Apache 2.0 licensed.
- Throughput reaches 143 tokens/s on RTX 5090 and 46.8 tokens/s on M5 Max, with 40% better energy efficiency than a full-precision 8B model.
- It improves reasoning, coding, vision, and agentic capabilities, enabling local deployment for coding agents, computer-use workflows, and private analysis.