数据 API / Tavily / Crawl
Graph-based website traversal tool using Tavily Crawl.
Walk a site from a root url and return the content of the pages it finds. Steer it with natural-language instructions plus regex path and domain filters, and bound it with max_depth, max_breadth and limit. Returns base_url and results[] with url and raw_content. Use it for broad coverage of one site — documentation, a catalogue, a competitor's blog. It answers synchronously, which post_firecrawl_crawl does not: that one runs as a background job and suits crawls too large to wait on.
参数
没有路径、查询或请求头参数。
请求体
urlstring必填The root URL to begin the crawl.
instructionsstring可选Natural language instructions for the crawler.
chunks_per_sourceinteger可选Maximum number of relevant chunks returned per source.
max_depthinteger可选Max depth of the crawl.
max_breadthinteger可选Max number of links to follow per level of the tree.
limitinteger可选Total number of links the crawler will process before stopping.
select_pathsstring[]可选Regex patterns to select only URLs with specific path patterns.
select_domainsstring[]可选Regex patterns to select crawling to specific domains or subdomains.
exclude_pathsstring[]可选Regex patterns to exclude URLs with specific path patterns.
exclude_domainsstring[]可选Regex patterns to exclude specific domains or subdomains from crawling.
allow_externalboolean可选Include external domain links in the final results list.
include_imagesboolean可选Include images in the crawl results.
extract_depthstring可选允许值: basic · advanced
Depth of the extraction process.
formatstring可选允许值: markdown · text
Format of the extracted web page content.
include_faviconboolean可选Include the favicon URL for each result.
timeoutnumber可选Maximum time in seconds to wait for the crawl operation.
include_usageboolean可选Include credit usage information in the response.
请求示例
curl -X POST "https://openapi.felo.ai/v1/beta/tavily/crawl" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "<string>",
"instructions": "<string>",
"chunks_per_source": 0,
"max_depth": 0,
"max_breadth": 0,
"limit": 0,
"select_paths": [
"<string>"
],
"select_domains": [
"<string>"
]
}'响应
响应字段
base_urlstring可选The base URL that was crawled.
resultsobject[]可选urlstring可选URL of the crawled page.
raw_contentstring可选Extracted raw content from the page.
faviconstring可选Favicon URL of the crawled page.
response_timenumber可选Time in seconds it took to complete the request.
usageobject可选creditsinteger可选Credit usage details for the request.
request_idstring可选Unique request identifier.