Benchmarking retrieval for agents on messy real-world company knowledge
7 hours ago
- Public retrieval benchmarks fail for company knowledge because they don't cover real-world use cases like docs, Slack, tickets, and varied agent queries.
- The Company Knowledge Bench is a custom benchmark built from 1,000 real production queries, with each eval case containing a query, corpus snapshot, and a retrieval criterion.
- Retrieval quality is defined by completeness, minimality, and source preference, which favors authoritative and current sources over weaker ones.
- To scale benchmark creation, agents generate labels after being validated against 170 fully human-labeled eval cases, ensuring high agreement with human judgment.
- Fixed pipeline results show that adding a reranker is the cheapest big win, improving score from 0.41 to 0.50 with little latency cost.
- Query decomposition boosts score to 0.56 but more than doubles latency and adds model call costs.
- An optimized fixed pipeline (Kapa Default) scores 0.61 in 3.3 seconds, matching a frontier-model grep agent in a fifth of the time.
- Agentic retrievers show that intelligence matters: a larger reasoning model scores 0.61 vs. 0.54 for a smaller one, but with high latency and token costs.
- Kapa Deep, an optimized agentic retriever, achieves the highest score of 0.65 in about 5 seconds, with the best precision and lowest token usage per query.
- The benchmark is private because it is built from real production data, and the team plans to keep improving accuracy, latency, and cost.