資料 API / Tavily / Crawl
Graph-based website traversal tool using Tavily Crawl.
Walk a site from a root url and return the content of the pages it finds. Steer it with natural-language instructions plus regex path and domain filters, and bound it with max_depth, max_breadth and limit. Returns base_url and results[] with url and raw_content. Use it for broad coverage of one site — documentation, a catalogue, a competitor's blog. It answers synchronously, which post_firecrawl_crawl does not: that one runs as a background job and suits crawls too large to wait on.
參數
沒有路徑、查詢或請求標頭參數。
請求主體
urlstring必填The root URL to begin the crawl.
instructionsstring選填Natural language instructions for the crawler.
chunks_per_sourceinteger選填Maximum number of relevant chunks returned per source.
max_depthinteger選填Max depth of the crawl.
max_breadthinteger選填Max number of links to follow per level of the tree.
limitinteger選填Total number of links the crawler will process before stopping.
select_pathsstring[]選填Regex patterns to select only URLs with specific path patterns.
select_domainsstring[]選填Regex patterns to select crawling to specific domains or subdomains.
exclude_pathsstring[]選填Regex patterns to exclude URLs with specific path patterns.
exclude_domainsstring[]選填Regex patterns to exclude specific domains or subdomains from crawling.
allow_externalboolean選填Include external domain links in the final results list.
include_imagesboolean選填Include images in the crawl results.
extract_depthstring選填允許值: basic · advanced
Depth of the extraction process.
formatstring選填允許值: markdown · text
Format of the extracted web page content.
include_faviconboolean選填Include the favicon URL for each result.
timeoutnumber選填Maximum time in seconds to wait for the crawl operation.
include_usageboolean選填Include credit usage information in the response.
請求範例
curl -X POST "https://openapi.felo.ai/v1/beta/tavily/crawl" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "<string>",
"instructions": "<string>",
"chunks_per_source": 0,
"max_depth": 0,
"max_breadth": 0,
"limit": 0,
"select_paths": [
"<string>"
],
"select_domains": [
"<string>"
]
}'回應
回應欄位
base_urlstring選填The base URL that was crawled.
resultsobject[]選填urlstring選填URL of the crawled page.
raw_contentstring選填Extracted raw content from the page.
faviconstring選填Favicon URL of the crawled page.
response_timenumber選填Time in seconds it took to complete the request.
usageobject選填creditsinteger選填Credit usage details for the request.
request_idstring選填Unique request identifier.