Data API / DataForSEO / On Page
OnPage API Resources
The non-page assets a crawl found - images, scripts, stylesheets - with their status codes and sizes. Pages with search_after_token. Reads a finished crawl, so it needs the id from post_dataforseo_on_page_submit and returns crawl_progress plus a crawl_status of max_crawl_pages, pages_in_queue and pages_crawled - check those before trusting a small result, because a crawl still running simply has less to report. For the assets the crawler could not reach at all use post_dataforseo_on_page_uncrawlable_resources. 💰 Free upstream: querying a finished crawl costs nothing, only the crawl itself does.
Параметры
Параметров пути, запроса и заголовков нет.
Тело запроса
idstringОбязательныйID of the task required field you can get this ID in the response of the Task POST endpoint example: “07131248-1535-0216-1000-17384017ad04”
urlstringНеобязательныйpage URL optional field specify this field if you want to get the resources for a specific page note that to obtain resource’s meta from a particular URL, you should specify the URL in this field; if you do not indicate a url when setting a task, resource’s meta in the results will be returned based on the data from the page where our crawler first saw the resource
limitintegerНеобязательныйthe maximum number of returned resources optional field default value: 100 maximum value: 1000
offsetintegerНеобязательныйoffset in the results array of returned resources optional field default value: 0 if you specify the 10 value, the first ten resources in the results array will be omitted and the data will be provided for the successive resources
filtersarrayНеобязательныйarray of results filtering parameters optional field you can add several filters at once (8 filters maximum) you should set a logical operator and, or between the conditions the following operators are supported: regex, not_regex, , , >, >=, =, , in, not_in, like, not_like you can use the % operator with like and not_like to match any string of zero or more characters example: ["resource_type","=","stylesheet"] [["resource_type","=","image"], "and",["checks.is_https","=",false]] [["fetch_timing.duration_time",">",1],"and",[["total_transfer_size",">",100],"or",["checks.high_loading_time","=",true]]] The full list of possible filters is available by this link.
relevant_pages_filtersarrayНеобязательныйfilter the resources by relevant pages optional field you can use this field to obtain resources from pages matching to the defined parameters you can apply the same filters here as available for the pages endpoint you can add several filters at once (8 filters maximum) you should set a logical operator and, or between the conditions the following operators are supported: regex, not_regex, , , >, >=, =, , in, not_in, like, not_like you can use the % operator with like and not_like to match any string of zero or more characters example: ["checks.no_image_title","=",true]
order_bystring[]Необязательныйresults sorting rules optional field you can use the same values as in the filters array to sort the results possible sorting types: asc – results will be sorted in the ascending order desc – results will be sorted in the descending order you should use a comma to set up a sorting type example: ["size,desc"] note that you can set no more than three sorting rules in a single request you should use a comma to separate several sorting rules example: ["size,desc","fetch_timing.fetch_end,desc"]
search_after_tokenstringНеобязательныйtoken for subsequent requests optional field provided in the identical filed of the response to each request; use this parameter to avoid timeouts while trying to obtain over 20,000 results in a single request; by specifying the unique search_after_token value from the response array, you will get the subsequent results of the initial task; search_after_token values are unique for each subsequent task ; Note: if the search_after_token is specified in the request, all other parameters should be identical to the previous request
tagstringНеобязательныйuser-defined task identifier optional field the character limit is 255 you can use this parameter to identify the task and match it with the result you will find the specified tag value in the data object of the response
Пример запроса
curl -X POST "https://openapi.felo.ai/v1/beta/dataforseo/on_page/resources" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"id": "<string>",
"url": "<string>",
"limit": 0,
"offset": 0,
"filters": "<string>",
"relevant_pages_filters": "<string>",
"order_by": [
"<string>"
],
"search_after_token": "<string>"
}'Ответ
Поля ответа
versionstringНеобязательныйthe current version of the API
status_codeintegerНеобязательныйgeneral status code you can find the full list of the response codes here Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions
status_messagestringНеобязательныйgeneral informational message you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйtotal tasks cost, USD
tasks_countintegerНеобязательныйthe number of tasks in the tasks array
tasks_errorintegerНеобязательныйthe number of tasks in the tasks array returned with an error
tasksstring[]Необязательныйarray of tasks
idstringНеобязательныйtask identifier unique task identifier in our system in the UUID format
status_codeintegerНеобязательныйstatus code of the task generated by DataForSEO; can be within the following range: 10000-60000 you can find the full list of the response codes here
status_messagestringНеобязательныйinformational message of the task you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйcost of the task, USD
result_countintegerНеобязательныйnumber of elements in the result array
pathstring[]НеобязательныйURL path
dataobjectНеобязательныйcontains the same parameters that you specified in the POST request
resultstring[]Необязательныйarray of results
crawl_progressstringНеобязательныйstatus of the crawling session possible values: in_progress, finished
crawl_statusobjectНеобязательныйdetails of the crawling session
max_crawl_pagesintegerНеобязательныйmaximum number of pages to crawl indicates the max_crawl_pages limit you specified when setting a task
pages_in_queueintegerНеобязательныйnumber of pages that are currently in the crawling queue
pages_crawledintegerНеобязательныйnumber of crawled pages
total_items_countintegerНеобязательныйtotal number of relevant items crawled
items_countintegerНеобязательныйnumber of items in the results array
itemsstring[]Необязательныйitems array
resource_typestringНеобязательныйtype of the returned resource possible types: script, image, stylesheet, broken
metaobjectНеобязательныйresource properties the value depends on the resource_type note that if you do not indicate a url when setting a task, resource’s meta is returned based on the data from the page where our crawler first saw the resource; to obtain resource’s meta from a particular url, specify that URL when setting a task
alternative_textstringНеобязательныйcontent of the image alt attribute the value depends on the resource_type
titlestringНеобязательныйtitle
original_widthintegerНеобязательныйoriginal image width in px
original_heightintegerНеобязательныйoriginal image height in px
widthintegerНеобязательныйimage width in px
heightintegerНеобязательныйimage height in px
status_codeintegerНеобязательныйstatus code of the page where a given resource is located
locationstringНеобязательныйlocation header indicates the URL to redirect a page to
urlstringНеобязательныйresource URL
sizeintegerНеобязательныйresource size indicates the size of a given resource measured in bytes
encoded_sizeintegerНеобязательныйresource size after encoding indicates the size of the encoded resource measured in bytes
total_transfer_sizeintegerНеобязательныйcompressed resource size indicates the compressed size of a given resource in bytes
fetch_timestringНеобязательныйdate and time when a resource was fetched in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: 2021-02-17 13:54:15 +00:00
fetch_timingobjectНеобязательныйresource fething time range
duration_timeintegerНеобязательныйindicates how many milliseconds it took to fetch a resource
fetch_startintegerНеобязательныйtime to start downloading the resource the amount of time a browser needs to start downloading a resource
fetch_endintegerНеобязательныйtime to complete downloading the resource the amount of time a browser needs to complete downloading a resource
cache_controlobjectНеобязательныйinstructions for caching
cachablebooleanНеобязательныйindicates whether the resource is cacheable
ttlintegerНеобязательныйtime to live the amount of time it takes for the browser to cache a resource; measured in milliseconds
checksobjectНеобязательныйresource check-ups contents of the array depend on the resource_type
no_content_encodingbooleanНеобязательныйresource with no content encoding indicates whether a page has no compression algorithm of the content; available for items with the following resource_type: script, image, stylesheet, broken
high_loading_timebooleanНеобязательныйresource with high loading time indicates whether a resource loading time exceeds 3 seconds; available for items with the following resource_type: script, image, stylesheet, broken
is_redirectbooleanНеобязательныйresource with redirects indicates whether a page with this resource has 3XX redirects to other pages; available for items with the following resource_type: script, image, stylesheet, broken
is_4xx_codebooleanНеобязательныйresource with with 4xx status code indicates whether a page with this resource has 4XX response code
is_5xx_codebooleanНеобязательныйresource with 5xx status code indicates whethera page with this resource has 5XX response code
is_brokenbooleanНеобязательныйbroken resource indicates whether a page with this resource returns 4xx, 5xx response codes or has broken elements inside the resource; available for items with the following resource_type: script, image, stylesheet, broken
is_wwwbooleanНеобязательныйpage with www indicates whether a page with this resource is on a www subdomain; available for items with the following resource_type: script, image, stylesheet, broken
is_httpsbooleanНеобязательныйpage with the https protocol available for items with the following resource_type: script, image, stylesheet, broken
is_httpbooleanНеобязательныйpage with the http protocol available for items with the following resource_type: script, image, stylesheet, broken
original_size_displayedbooleanНеобязательныйimage desplayes in its original size indicates whether the image is displayed in its original size; available for items with the following resource_type: image
is_minifiedbooleanНеобязательныйresource is minified indicates whether the content of a stylesheet or script is minified; available for items with the following resource_type: stylesheet, script
has_redirectbooleanНеобязательныйresource has a redirect available for items with the following resource_type: script, image; if the resource_type is image, this field will indicate whether other pages and/or resources have redirects pointing at the image; if the resource_type is script, this field will indicate whether the script contains a redirect
has_subrequestsbooleanНеобязательныйresource contains subrequests indicates whether the content of a stylesheet or script contain additional requests; available for items with the following resource_type: stylesheet, script
from_sitemapbooleanНеобязательныйresource was found on website’s sitemap if true, the resource was found on the sitemap of the website
resource_errorsobjectНеобязательныйresource errors and warnings
errorsstring[]Необязательныйresource errors
lineintegerНеобязательныйline where the error was found
columnintegerНеобязательныйcolumn where the error was found
messagestringНеобязательныйtext message of the error the full list of possible HTML errors can be found here
status_codeintegerНеобязательныйstatus code of the error possible values: 0 — Unidentified Error; 501 — Html Parse Error; 1501 — JS Parse Error; 2501 — CSS Parse Error; 3501 — Image Parse Error; 3502 — Image Scale Is Zero; 3503 — Image Size Is Zero; 3504 — Image Format Invalid
warningsstring[]Необязательныйresource warnings
lineintegerНеобязательныйline the warning relates to note that if "line": 0, the warning relates to the whole page
columnintegerНеобязательныйcolumn the warning relates to note that if "column": 0, the warning relates to the whole page
messagestringНеобязательныйtext message of the warning possible messages: "Has node with more than 60 childs." – HTML page has at least 1 tag nesting over 60 tags of the same level "Has more that 1500 nodes." – DOM tree contains over 1,500 elements "HTML depth more than 32 tags." – DOM depth exceeds 32 nodes
status_codeintegerНеобязательныйstatus code of the warning possible values: 0 — Unidentified Warning; 1 — Has node with more than 60 childs; 2 — Has more that 1500 nodes; 3 — HTML depth more than 32 tags
content_encodingstringНеобязательныйtype of encoding
media_typestringНеобязательныйtypes of media used to display a resource
accept_typestringНеобязательныйindicates the expected type of resource for example, if "resource_type": "broken", accept_type will indicate the type of the broken resource possible values: any, none, image, sitemap, robots, script, stylesheet, redirect, html, text, other, font
serverstringНеобязательныйserver version
last_modifiedobjectНеобязательныйcontains data on changes related to the resource if there is no data, the value will be null
headerstringНеобязательныйdate and time when the header was last modified in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: 2019-11-15 12:57:46 +00:00 if there is no data, the value will be null
sitemapstringНеобязательныйdate and time when the sitemap was last modified in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: 2019-11-15 12:57:46 +00:00 if there is no data, the value will be null
meta_tagstringНеобязательныйdate and time when the meta tag was last modified in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: 2019-11-15 12:57:46 +00:00 if there is no data, the value will be null