Hasty Briefsbeta

Bilingual

So you want to use OpenRouter?

21 days ago
  • Different providers serving the same model can have widely varying performance on benchmarks like GPQA and TAU-Bench, with up to 20-point swings in tool-calling scores.
  • Some providers fail to process vision model inputs correctly, returning errors or incorrect descriptions while claiming success with HTTP 200.
  • The reasoning effort parameter may be accepted but ignored by some providers, requiring tracking of actual reasoning tokens per provider.
  • Quantization settings (e.g., fp8 vs fp4) are a poor indicator of model quality; benchmarks are a more reliable filter.
  • Tool calls may appear as raw text in responses due to provider parser failures, necessitating client-side parsing.
  • Reasoning models can return HTTP 200 with empty content and no tool call, which counts as a failure requiring retries.
  • Empty completions (no content, no usage object) can occur with some providers, like StreamLake and Together.
  • Providers enforce different rules for passing reasoning content back in history, causing errors if mismatched.
  • Testing from a developer's laptop may not replicate production issues like rate-limiting by IP.
  • Pinning a set of providers is risky because providers can drop models, rate-limit, or fail, leading to full service outages.