Graph-based website traversal tool using Tavily Crawl.
Web Content
Web crawling, URL discovery and page extraction.
Graph-based website traversal tool using Tavily Crawl.
Walk a site from a root url and return the content of the pages it finds. Steer it with natural-language instructions plus regex path and domain filters, and bound it with max_depth, max_breadth and limit. Returns base_url and results[] with url and raw_content. Use it for broad coverage of one site — documentation, a catalogue, a competitor's blog. It answers synchronously, which post_firecrawl_crawl does not: that one runs as a background job and suits crawls too large to wait on.
- 型: stringurl必須
The root URL to begin the crawl.
- 型: booleanallow
_external Include external domain links in the final results list.
- 型: integerchunks
_per _source 最小値:1最大値:5Maximum number of relevant chunks returned per source.
- 型: array string[]exclude
_domains Regex patterns to exclude specific domains or subdomains from crawling.
- 型: array string[]exclude
_paths Regex patterns to exclude URLs with specific path patterns.
- 型: string 列挙値extract
_depth Depth of the extraction process.
値- basic
- advanced
- 型: string 列挙値format
Format of the extracted web page content.
値- markdown
- text
- 型: booleaninclude
_favicon Include the favicon URL for each result.
- 型: booleaninclude
_images Include images in the crawl results.
- 型: booleaninclude
_usage Include credit usage information in the response.
- 型: stringinstructions
Natural language instructions for the crawler.
- 型: integerlimit最小値:1
Total number of links the crawler will process before stopping.
- 型: integermax
_breadth 最小値:1最大値:500Max number of links to follow per level of the tree.
- 型: integermax
_depth 最小値:1最大値:5Max depth of the crawl.
- 型: array string[]select
_domains Regex patterns to select crawling to specific domains or subdomains.
- 型: array string[]select
_paths Regex patterns to select only URLs with specific path patterns.
- 型: number 形式: floattimeout最小値:10最大値:150
Maximum time in seconds to wait for the crawl operation.
- 型: object200
Crawl results returned successfully.
- 型: stringbase
_url The base URL that was crawled.
- 型: stringrequest
_id Unique request identifier.
- 型: number 形式: floatresponse
_time Time in seconds it took to complete the request.
- 型: array object[]results
- favicon型: string
Favicon URL of the crawled page.
- raw型: string
_content Extracted raw content from the page.
- url型: string
URL of the crawled page.
- 型: objectusage
- credits型: integer
Credit usage details for the request.
application/json - 400
The request could not be processed. Check the request parameters.
- 401
A valid Felo API key is required.
- 402
The request cannot proceed because a billing requirement is not met.
- 403
The account is not permitted to perform this operation.
- 429
A request or spending limit has been reached.
- 502
The service could not complete the request.
- 503
The API or billing service is temporarily unavailable.
- 504
The service timed out while processing the request.
- default
The operation failed. Keep the response request ID when contacting Felo support.
curl https://openapi.felo.ai/v1/beta/tavily/crawl \
--request POST \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer YOUR_SECRET_TOKEN' \
--data '{
"url": "docs.tavily.com",
"instructions": "",
"chunks_per_source": 3,
"max_depth": 1,
"max_breadth": 20,
"limit": 50,
"select_paths": [
""
],
"select_domains": [
""
],
"exclude_paths": [
""
],
"exclude_domains": [
""
],
"allow_external": true,
"include_images": false,
"extract_depth": "basic",
"format": "markdown",
"include_favicon": false,
"timeout": 150,
"include_usage": false
}'
{
"base_url": "string",
"results": [
{
"url": "string",
"raw_content": "string",
"favicon": "string"
}
],
"response_time": 1,
"usage": {
"credits": 1
},
"request_id": "string"
}