Extract API
Extract API
Fetch a public URL and return clean page content as text or HTML. Each call fetches the current version of the page, not a cached snapshot.
text
Auth
Preferred: Authorization: $DESEARCH_API_KEY (raw key from Console). See Auth and API keys.
curl
bash
JavaScript (desearch-js@1.5.1)
javascript
webCrawl() on this package is a deprecated alias. New JavaScript code should call extract().
Python (PyPI desearch-py==1.2.1)
Published Python today exposes web_crawl, not extract. Use REST /web/extract or web_crawl until a newer PyPI release includes extract().
python
Parameters (REST)
| Param | Type | Notes |
|---|---|---|
url | string | Required |
format | text | html | Default text |
js | "true" | "false" | Optional headless render |
wait | int (ms) | Only with js=true |
Legacy crawl
GET /web/crawl remains a compatibility alias. New integrations should call /web/extract.
Response formats and errors are in the API Reference. How the crawler identifies itself and respects robots.txt is documented at desearch.ai/crawler.
Typed clients: JavaScript SDK · Python SDK.