Data API / Tavily / Crawl
Graph-based website traversal tool using Tavily Crawl.
Walk a site from a root url and return the content of the pages it finds. Steer it with natural-language instructions plus regex path and domain filters, and bound it with max_depth, max_breadth and limit. Returns base_url and results[] with url and raw_content. Use it for broad coverage of one site — documentation, a catalogue, a competitor's blog. It answers synchronously, which post_firecrawl_crawl does not: that one runs as a background job and suits crawls too large to wait on.
Параметры
Параметров пути, запроса и заголовков нет.
Тело запроса
urlstringОбязательныйThe root URL to begin the crawl.
instructionsstringНеобязательныйNatural language instructions for the crawler.
chunks_per_sourceintegerНеобязательныйMaximum number of relevant chunks returned per source.
max_depthintegerНеобязательныйMax depth of the crawl.
max_breadthintegerНеобязательныйMax number of links to follow per level of the tree.
limitintegerНеобязательныйTotal number of links the crawler will process before stopping.
select_pathsstring[]НеобязательныйRegex patterns to select only URLs with specific path patterns.
select_domainsstring[]НеобязательныйRegex patterns to select crawling to specific domains or subdomains.
exclude_pathsstring[]НеобязательныйRegex patterns to exclude URLs with specific path patterns.
exclude_domainsstring[]НеобязательныйRegex patterns to exclude specific domains or subdomains from crawling.
allow_externalbooleanНеобязательныйInclude external domain links in the final results list.
include_imagesbooleanНеобязательныйInclude images in the crawl results.
extract_depthstringНеобязательныйДопустимые значения: basic · advanced
Depth of the extraction process.
formatstringНеобязательныйДопустимые значения: markdown · text
Format of the extracted web page content.
include_faviconbooleanНеобязательныйInclude the favicon URL for each result.
timeoutnumberНеобязательныйMaximum time in seconds to wait for the crawl operation.
include_usagebooleanНеобязательныйInclude credit usage information in the response.
Пример запроса
curl -X POST "https://openapi.felo.ai/v1/beta/tavily/crawl" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "<string>",
"instructions": "<string>",
"chunks_per_source": 0,
"max_depth": 0,
"max_breadth": 0,
"limit": 0,
"select_paths": [
"<string>"
],
"select_domains": [
"<string>"
]
}'Ответ
Поля ответа
base_urlstringНеобязательныйThe base URL that was crawled.
resultsobject[]НеобязательныйurlstringНеобязательныйURL of the crawled page.
raw_contentstringНеобязательныйExtracted raw content from the page.
faviconstringНеобязательныйFavicon URL of the crawled page.
response_timenumberНеобязательныйTime in seconds it took to complete the request.
usageobjectНеобязательныйcreditsintegerНеобязательныйCredit usage details for the request.
request_idstringНеобязательныйUnique request identifier.