Hasty Briefsbeta

Bilingual

Launch HN: Tokenless (YC S26) – Automatic model switching to save money

9 hours ago
  • Tokenless reduces inference costs by routing requests to the most suitable model, canceling unnecessary ones to halve costs.
  • The service maintains output quality comparable to top models (e.g., Opus 4.8) while significantly lowering cost per task.
  • It offers drop-in compatibility with OpenAI/Anthropic APIs and provides model options like PRO, GPT 5.5, Opus 4.8, and MAX for different quality/savings balances.
  • Benchmark data shows PRO achieves 72% task completion at $0.32/task, saving up to 57% over GPT 5.5.
  • MAX mode prioritizes maximum quality by routing tasks to the best available model.
  • The platform is built by AI researchers from leading institutions and backed by Y Combinator.