We track retrieval quality as a product input and publish it instead of hiding it inside marketing claims. Figures below are Desearch’s own internal evaluation.
How well answer summaries match the source material (higher is better).
| Benchmark | Desearch | Baseline | What it measures |
|---|---|---|---|
| FreshQA | 86 | 70 | Questions that require fresh, current knowledge. |
| BrowseComp | 78 | 66 | Complex web-browsing and multi-hop retrieval tasks. |
| DeepSearch | 82 | 63 | Deep, multi-step search and synthesis. |
| FRAMES | 88 | 72 | Retrieval and factuality across framed queries. |
| HLE | 74 | 58 | Hard reasoning-style evaluation. |
Desearch internal evaluation. “Baseline” is a representative comparison point across each benchmark family. Full methodology and run dates are documented alongside product releases.
Text relevance measures how well answer summaries match the source material. The benchmark families (FreshQA, BrowseComp, DeepSearch, FRAMES, HLE) track retrieval quality on freshness, browsing depth, multi-step search, factuality, and hard reasoning.
These are Desearch's own internal evaluations, run as product-quality inputs. They are published for transparency; the same scorecards guide ranking, crawl fallback, and answer packaging.
On Desearch's text-relevance evaluation, Desearch scores 92.57, ahead of Exa (91.82) and Perplexity (90.7). See the table above for the full set.