Hasty Briefsbeta

Bilingual

Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

2 hours ago
  • Swift-Qwen3.8-27B is a reasoning-efficient derivative of Qwen3.8-27B that uses 58.3% fewer thinking tokens while maintaining near-identical performance (<1% loss), achieving up to 1.95x speed-up on tasks.
  • The model can be used with multiple inference methods: Transformers pipeline, vLLM server, SGLang server, Docker, and a free OpenAI-compatible API at ukisai.com.
  • Training involved penalizing reasoning-marker tokens to reduce overthinking, and includes a transfer component from ThinkingCap-Qwen3.6-27B.
  • Benchmarks (GPQA-Diamond, MMLU-Pro, AIME 2026, etc.) show significant token reduction (26-50%) with minimal accuracy change across reasoning efforts and quantized (INT4) deployments.
  • The weights are available under Swift Open License v1.0, free for research and commercial use for organizations with annual revenue under $1M.