Extract API

    Extract API

    Fetch a public URL and return clean page content as text or HTML. Each call fetches the current version of the page, not a cached snapshot.

    text
    GET https://api.desearch.ai/web/extract

    Auth

    Preferred: Authorization: $DESEARCH_API_KEY (raw key from Console). See Auth and API keys.

    curl

    bash
    curl --request GET \ 'https://api.desearch.ai/web/extract?url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FArtificial_intelligence&format=text' \ --header "Authorization: $DESEARCH_API_KEY"

    JavaScript (desearch-js@1.5.1)

    javascript
    const { Desearch } = require("desearch-js"); async function main() { const client = new Desearch(process.env.DESEARCH_API_KEY); const content = await client.extract({ url: "https://en.wikipedia.org/wiki/Artificial_intelligence", format: "text", }); console.log(String(content).slice(0, 500)); } main();

    webCrawl() on this package is a deprecated alias. New JavaScript code should call extract().

    Python (PyPI desearch-py==1.2.1)

    Published Python today exposes web_crawl, not extract. Use REST /web/extract or web_crawl until a newer PyPI release includes extract().

    python
    import asyncio import os from desearch_py import Desearch async def main(): async with Desearch(api_key=os.environ["DESEARCH_API_KEY"]) as client: content = await client.web_crawl( url="https://en.wikipedia.org/wiki/Artificial_intelligence", format="text", ) print(content[:500]) asyncio.run(main())

    Parameters (REST)

    ParamTypeNotes
    urlstringRequired
    formattext | htmlDefault text
    js"true" | "false"Optional headless render
    waitint (ms)Only with js=true

    Legacy crawl

    GET /web/crawl remains a compatibility alias. New integrations should call /web/extract.

    Response formats and errors are in the API Reference. How the crawler identifies itself and respects robots.txt is documented at desearch.ai/crawler.

    Typed clients: JavaScript SDK · Python SDK.