Data API / DataForSEO / On Page
OnPage API Non-indexable Pages
The crawled pages search engines will not index, each with a reason and url. Reads a finished crawl, so it needs the id from post_dataforseo_on_page_submit and returns crawl_progress plus a crawl_status of max_crawl_pages, pages_in_queue and pages_crawled - check those before trusting a small result, because a crawl still running simply has less to report. The reason field is the whole value here - noindex, canonical elsewhere, robots-blocked are very different problems with the same symptom. 💰 Free upstream: querying a finished crawl costs nothing, only the crawl itself does.
Параметры
Параметров пути, запроса и заголовков нет.
Тело запроса
idstringОбязательныйID of the task required field you can get this ID in the response of the Task POST endpoint example: “07131248-1535-0216-1000-17384017ad04”
limitintegerНеобязательныйthe maximum number of returned pages optional field default value: 100 maximum value: 1000
offsetintegerНеобязательныйoffset in the results array of returned pages optional field default value: 0 if you specify the 10 value, the first ten pages in the results array will be omitted and the data will be provided for the successive pages
filtersarrayНеобязательныйarray of results filtering parameters optional field you can add several filters at once (8 filters maximum) you should set a logical operator and, or between the conditions the following operators are supported: regex, not_regex, , , >, >=, =, , in, not_in, like, not_like you can use the % operator with like and not_like to match any string of zero or more characters example: ["reason","=","robots_txt"][["reason","","robots_txt"], "and", ["url","not_like","%/wp-admin/%"]] [["url","not_like","%/wp-admin/%"], "and", [["reason","","meta_tag"],"or",["reason","","http_header"]]] The full list of possible filters is available by this link.
Пример запроса
curl -X POST "https://openapi.felo.ai/v1/beta/dataforseo/on_page/non_indexable" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"id": "<string>",
"limit": 0,
"offset": 0,
"filters": "<string>"
}'Ответ
Поля ответа
versionstringНеобязательныйthe current version of the API
status_codeintegerНеобязательныйgeneral status code you can find the full list of the response codes here Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions
status_messagestringНеобязательныйgeneral informational message you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйtotal tasks cost, USD
tasks_countintegerНеобязательныйthe number of tasks in the tasks array
tasks_errorintegerНеобязательныйthe number of tasks in the tasks array returned with an error
tasksstring[]Необязательныйarray of tasks
idstringНеобязательныйtask identifier unique task identifier in our system in the UUID format
status_codeintegerНеобязательныйstatus code of the task generated by DataForSEO; can be within the following range: 10000-60000 you can find the full list of the response codes here
status_messagestringНеобязательныйinformational message of the task you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйcost of the task, USD
result_countintegerНеобязательныйnumber of elements in the result array
pathstring[]НеобязательныйURL path
dataobjectНеобязательныйcontains the same parameters that you specified in the POST request
resultstring[]Необязательныйarray of results
crawl_progressstringНеобязательныйstatus of the crawling session possible values: in_progress, finished
crawl_statusobjectНеобязательныйdetails of the crawling session
max_crawl_pagesintegerНеобязательныйmaximum number of pages to crawl indicates the max_crawl_pages limit you specified when setting a task
pages_in_queueintegerНеобязательныйnumber of pages that are currently in the crawling queue
pages_crawledintegerНеобязательныйnumber of crawled pages
total_items_countintegerНеобязательныйtotal number of relevant items in the database
items_countintegerНеобязательныйnumber of items in the results array
itemsstring[]Необязательныйitems array
reasonstringНеобязательныйthe reason why the page is non-indexable can take the following values: robots_txt, meta_tag, http_header, attribute, too_many_redirects
urlstringНеобязательныйurl of the non-indexable page