Kimi K3, and what we can still learn from the pelican benchmark
3 hours ago
- Moonshot AI announced Kimi K3, a 2.8 trillion parameter model, promising open weights by July 27, 2026.
- Self-reported benchmarks show K3 beating Claude Opus 4.8 max and GPT-5.5 high, but falling behind Claude Fable 5 and GPT-5.6 Sol.
- Artificial Analysis reports K3 achieves an Elo of 1547 on long-horizon knowledge work, costs $0.94 per task, and uses 21% fewer output tokens than K2.6.
- Pricing is $3 per million input tokens and $15 per million output tokens, making it the most expensive model from a Chinese AI lab.
- The pelican benchmark test (generate SVG of a pelican riding a bicycle) cost $0.25 and used 13,241 reasoning tokens, revealing K3's high reasoning effort and a hidden system prompt of about 85 tokens.
- The pelican test remains useful as a quick 'hello world' for trying new models, estimating cost and reasoning, checking spatial awareness, and maintaining a tradition on Hacker News.
- Despite its limitations (no agentic evaluation), the pelican benchmark still provides value for rough comparisons within a model family and as a forcing function to actually test the model.