So you want to use OpenRouter?
21 days ago
- Different providers serving the same model can have widely varying performance on benchmarks like GPQA and TAU-Bench, with up to 20-point swings in tool-calling scores.
- Some providers fail to process vision model inputs correctly, returning errors or incorrect descriptions while claiming success with HTTP 200.
- The reasoning effort parameter may be accepted but ignored by some providers, requiring tracking of actual reasoning tokens per provider.
- Quantization settings (e.g., fp8 vs fp4) are a poor indicator of model quality; benchmarks are a more reliable filter.
- Tool calls may appear as raw text in responses due to provider parser failures, necessitating client-side parsing.
- Reasoning models can return HTTP 200 with empty content and no tool call, which counts as a failure requiring retries.
- Empty completions (no content, no usage object) can occur with some providers, like StreamLake and Together.
- Providers enforce different rules for passing reasoning content back in history, causing errors if mismatched.
- Testing from a developer's laptop may not replicate production issues like rate-limiting by IP.
- Pinning a set of providers is risky because providers can drop models, rate-limit, or fail, leading to full service outages.