データ API / DataForSEO / On Page
Uncrawlable Resources
The resources a crawl could not fetch, and why. Returns total_items_count, items_count and items. Reads a finished crawl, so it needs the id from post_dataforseo_on_page_submit and returns crawl_progress plus a crawl_status of max_crawl_pages, pages_in_queue and pages_crawled - check those before trusting a small result, because a crawl still running simply has less to report. These are the broken references a site owner would want first; the ones that did load are in post_dataforseo_on_page_resources. 💰 Free upstream: querying a finished crawl costs nothing, only the crawl itself does.
パラメータ
パス・クエリ・ヘッダーのパラメータはありません。
リクエストボディ
idstring必須ID of the task required field you can get this ID in the response of the Task POST endpoint example: "07131248-1535-0216-1000-17384017ad04"
limitinteger任意the maximum number of returned uncrawlable resources optional field default value: 100 maximum value: 1000
offsetinteger任意offset in the results array of returned uncrawlable resources optional field default value: 0 if you specify the 10 value, the first ten invalid resources in the results array will be omitted and the data will be provided for the successive invalid resources
order_bystring[]任意results sorting rules optional field you can use the same values as in the filters array to sort the results possible sorting types: asc - results will be sorted in the ascending order desc - results will be sorted in the descending order you should use a comma to set up a sorting type example: ["meta.content_type,desc"] note that you can set no more than three sorting rules in a single request you should use a comma to separate several sorting rules example: ["meta.content_type,asc","fetch_time,desc"]
filtersarray任意array of results filtering parameters optional field you can add several filters at once (8 filters maximum) you should set a logical operator and, or between the conditions the following operators are supported: regex, not_regex, , , >, >=, =, , in, not_in, like, not_like you can use the % operator with like and not_like to match any string of zero or more characters example: [["meta.content_type","=","image/jpeg"], "and", ["url","not_like","%/help-center/%"]]The full list of possible filters is available by this link.
リクエスト例
curl -X POST "https://openapi.felo.ai/v1/beta/dataforseo/on_page/uncrawlable_resources" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"id": "<string>",
"limit": 0,
"offset": 0,
"order_by": [
"<string>"
],
"filters": "<string>"
}'レスポンス
レスポンスのフィールド
versionstring任意the current version of the API
status_codeinteger任意general status code you can find the full list of the response codes here Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions
status_messagestring任意general informational message you can find the full list of general informational messages here
timestring任意execution time, seconds
costnumber任意total tasks cost, USD
tasks_countinteger任意the number of tasks in the tasks array
tasks_errorinteger任意the number of tasks in the tasks array returned with an error
tasksstring[]任意array of tasks
idstring任意task identifier unique task identifier in our system in the UUID format
status_codeinteger任意status code of the task generated by DataForSEO; can be within the following range: 10000-60000 you can find the full list of the response codes here
status_messagestring任意informational message of the task you can find the full list of general informational messages here
timestring任意execution time, seconds
costnumber任意cost of the task, USD
result_countinteger任意number of elements in the result array
pathstring[]任意URL path
dataobject任意contains the same parameters that you specified in the POST request
resultstring[]任意array of results
crawl_progressstring任意status of the crawling session possible values: in_progress, finished
crawl_statusobject任意details of the crawling session
max_crawl_pagesinteger任意maximum number of pages to crawl indicates the max_crawl_pages limit you specified when setting a task
pages_in_queueinteger任意number of pages that are currently in the crawling queue
pages_crawledinteger任意number of crawled pages
total_items_countinteger任意total number of uncrawlable resources found total number of uncrawlable resources found during the crawl of the target domain
items_countinteger任意number of uncrawlable resources in the items array
itemsstring[]任意array of uncrawlable resources
urlstring任意URL of the uncrawlable resource
reasonstring任意reason the resource is uncrawlable can take the following values: content_type_inconsistency
status_codeinteger任意HTTP response code returned by the uncrawlable resource possible values: 200
fetch_timestring任意date and time when the resource was fetched in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: 2026-03-09 18:20:32 +00:00
metaobject任意metadata of the uncrawlable resource
content_typestring任意actual content type of the resource
expected_content_typesstring[]任意expected content types for the resource list of content types that were expected by the crawler based on how the resource is referenced on the page