Desearch search-quality benchmarks

    We track retrieval quality as a product input and publish it instead of hiding it inside marketing claims. Figures below are Desearch’s own internal evaluation.

    Text relevance

    How well answer summaries match the source material (higher is better).

    Desearch
    92.57
    Exa
    91.82
    Perplexity
    90.70
    Gemini
    71.58
    Andi
    23.46

    Quality across benchmark families

    BenchmarkDesearchBaselineWhat it measures
    FreshQA8670Questions that require fresh, current knowledge.
    BrowseComp7866Complex web-browsing and multi-hop retrieval tasks.
    DeepSearch8263Deep, multi-step search and synthesis.
    FRAMES8872Retrieval and factuality across framed queries.
    HLE7458Hard reasoning-style evaluation.

    Desearch internal evaluation. “Baseline” is a representative comparison point across each benchmark family. Full methodology and run dates are documented alongside product releases.

    Benchmark FAQ

    What do the Desearch benchmarks measure?

    Text relevance measures how well answer summaries match the source material. The benchmark families (FreshQA, BrowseComp, DeepSearch, FRAMES, HLE) track retrieval quality on freshness, browsing depth, multi-step search, factuality, and hard reasoning.

    Are these independent third-party results?

    These are Desearch's own internal evaluations, run as product-quality inputs. They are published for transparency; the same scorecards guide ranking, crawl fallback, and answer packaging.

    How does Desearch compare to Exa and Perplexity on relevance?

    On Desearch's text-relevance evaluation, Desearch scores 92.57, ahead of Exa (91.82) and Perplexity (90.7). See the table above for the full set.