DataForSEO
SEO, search-engine results, keyword research, backlinks and AI visibility data.
OnPage API Non-indexable Pages
The crawled pages search engines will not index, each with a reason and url. Reads a finished crawl, so it needs the id from post_dataforseo_on_page_submit and returns crawl_progress plus a crawl_status of max_crawl_pages, pages_in_queue and pages_crawled - check those before trusting a small result, because a crawl still running simply has less to report. The reason field is the whole value here - noindex, canonical elsewhere, robots-blocked are very different problems with the same symptom. 💰 Free upstream: querying a finished crawl costs nothing, only the crawl itself does.
- 型: array object[]
- id型: string必須
ID of the task required field you can get this ID in the response of the Task POST endpoint example: “07131248-1535-0216-1000-17384017ad04”
- filters型: array
array of results filtering parameters optional field you can add several filters at once (8 filters maximum) you should set a logical operator and, or between the conditions the following operators are supported: regex, not_regex, , , >, >=, =, , in, not_in, like, not_like you can use the % operator with like and not_like to match any string of zero or more characters example: ["reason","=","robots_txt"][["reason","","robots_txt"], "and", ["url","not_like","%/wp-admin/%"]] [["url","not_like","%/wp-admin/%"], "and", [["reason","","meta_tag"],"or",["reason","","http_header"]]] The full list of possible filters is available by this link.
- limit型: integer
the maximum number of returned pages optional field default value: 100 maximum value: 1000
- offset型: integer
offset in the results array of returned pages optional field default value: 0 if you specify the 10 value, the first ten pages in the results array will be omitted and the data will be provided for the successive pages
- 型: object200
Successful response
- 型: numbercost
total tasks cost, USD
- 型: integerstatus
_code general status code you can find the full list of the response codes here Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions
- 型: stringstatus
_message general informational message you can find the full list of general informational messages here
- 型: array string[]tasks
array of tasks
- 型: integertasks
_count the number of tasks in the tasks array
- 型: integertasks
_error the number of tasks in the tasks array returned with an error
- 型: numbertasks
.cost cost of the task, USD
- 型: objecttasks
.data contains the same parameters that you specified in the POST request
- 型: stringtasks
.id task identifier unique task identifier in our system in the UUID format
- 型: array string[]tasks
.path URL path
- 型: array string[]tasks
.result array of results
- 型: integertasks
.result _count number of elements in the result array
- 型: stringtasks
.result .crawl _progress status of the crawling session possible values: in_progress, finished
- 型: objecttasks
.result .crawl _status details of the crawling session
- 型: integertasks
.result .crawl _status .max _crawl _pages maximum number of pages to crawl indicates the max_crawl_pages limit you specified when setting a task
- 型: integertasks
.result .crawl _status .pages _crawled number of crawled pages
- 型: integertasks
.result .crawl _status .pages _in _queue number of pages that are currently in the crawling queue
- 型: array string[]tasks
.result .items items array
- 型: integertasks
.result .items _count number of items in the results array
- 型: stringtasks
.result .reason the reason why the page is non-indexable can take the following values: robots_txt, meta_tag, http_header, attribute, too_many_redirects
- 型: integertasks
.result .total _items _count total number of relevant items in the database
- 型: stringtasks
.result .url url of the non-indexable page
- 型: integertasks
.status _code status code of the task generated by DataForSEO; can be within the following range: 10000-60000 you can find the full list of the response codes here
- 型: stringtasks
.status _message informational message of the task you can find the full list of general informational messages here
- 型: stringtasks
.time execution time, seconds
- 型: stringtime
execution time, seconds
- 型: stringversion
the current version of the API
application/json - 400
Bad request
- 401
Unauthorized
- 402
The request cannot proceed because a billing requirement is not met.
- 403
The account is not permitted to perform this operation.
- 429
Rate limit exceeded
- 500
Internal server error
- 502
The service could not complete the request.
- 503
The API or billing service is temporarily unavailable.
- 504
The service timed out while processing the request.
- default
The operation failed. Keep the response request ID when contacting Felo support.
curl https://openapi.felo.ai/v1/beta/dataforseo/on_page/non_indexable \
--request POST \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer YOUR_SECRET_TOKEN' \
--data '[
{
"id": "",
"limit": 1,
"offset": 1,
"filters": []
}
]'
{
"version": "string",
"status_code": 1,
"status_message": "string",
"time": "string",
"cost": 1,
"tasks_count": 1,
"tasks_error": 1,
"tasks": [
"string"
],
"tasks.id": "string",
"tasks.status_code": 1,
"tasks.status_message": "string",
"tasks.time": "string",
"tasks.cost": 1,
"tasks.result_count": 1,
"tasks.path": [
"string"
],
"tasks.data": {},
"tasks.result": [
"string"
],
"tasks.result.crawl_progress": "string",
"tasks.result.crawl_status": {},
"tasks.result.crawl_status.max_crawl_pages": 1,
"tasks.result.crawl_status.pages_in_queue": 1,
"tasks.result.crawl_status.pages_crawled": 1,
"tasks.result.total_items_count": 1,
"tasks.result.items_count": 1,
"tasks.result.items": [
"string"
],
"tasks.result.reason": "string",
"tasks.result.url": "string"
}