API Reference

    v1.0.0 ยท updated from bundled OpenAPI

    GET/web/crawl

    Crawl API

    The Desearch API allows you to crawl web content Billing metadata: successful billable responses include the `X-Desearch-Cost-Usd`, `X-Desearch-Usage-Count`, `X-Desearch-Service`, and `X-Desearch-Currency` headers. JSON object responses also include optional `cost_usd`, `usage_count`, `service`, and `currency` fields. For example, `usage_count=10` at `$0.015 / 1,000` resolves to `cost_usd=0.00015`. JSON arrays, text, and streaming responses keep their body shape and expose billing metadata only through headers.

    Base URL

    https://api.desearch.ai

    Implementation guidance

    Use Crawl API when your application needs web page contentwithout maintaining its own crawler, social-data pipeline, or retrieval infrastructure. The endpoint returns structured data that can be consumed by agents, research tools, dashboards, and retrieval-augmented generation workflows. Keep the request focused on the minimum fields your workflow needs so responses remain fast and easy to validate.

    Send requests to GET /web/crawl with your Desearch API key in the authorization header. Treat the key as a server-side secret, validate all required parameters before sending the request, and handle non-2xx responses explicitly. Production integrations should use sensible timeouts, retry only transient failures with backoff, and record the response status and request identifier for troubleshooting.

    Test the example request below before integrating it into a larger workflow. Confirm that the returned fields, links, and timestamps meet your freshness requirements, then add schema validation at the application boundary. For agentic systems, preserve source URLs with the generated answer so users can inspect evidence and your application can avoid presenting unsupported claims.

    Query Parameters

    urlstringrequired

    Url to crawl

    formatstring

    Format of the content to be returned. Options are 'html' or 'text'. Default is 'html'.

    Default: text

    jsstring

    Render the page with a headless browser before extracting content. Use for SPAs or pages whose body is populated by JavaScript. Must be exactly `true` or `false` (case-sensitive); enables a separate billing tier when set to `true`.

    Default: false

    waitstring

    Optional extra wait in milliseconds after the page loads. Only applied when ``js=true``.

    Responses

    200Content of the web link
    301Moved Permanently
    304Not Modified
    401Unauthorized
    422Validation Error
    429Too Many Requests
    500Internal Server Error
    502Crawler upstream unavailable
    Authorization

    Your API key is sent in the Authorization header.

    Request
    curl -X GET "https://api.desearch.ai/web/crawl?url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FArtificial_intelligence" \
      -H "Authorization: YOUR_API_KEY" \
      -H "Content-Type: application/json"
    ResponseSchema
    {
      "message": "Response schema not available"
    }

    Possible status codes:

    200401422429500