Data API / DataForSEO / On Page
Live OnPage API Content Parsing
Parses any url into structured content immediately, with no crawl needed - the only tool in this family that works standalone. browser_preset, browser_screen_width, browser_screen_height and custom_user_agent control the render, disable_cookie_popup removes consent overlays, store_raw_html keeps the source. Wrapped in DataForSEO's envelope: data in tasks[0].result, outcome in tasks[0].status_code - a rejected request still returns HTTP 200. For pages already inside a crawl use post_dataforseo_on_page_content_parsing, which is free.
Параметры
Параметров пути, запроса и заголовков нет.
Тело запроса
urlstringОбязательныйURL of the content to parse required field URL of the page to parse example: https://www.fujielectric.com/
custom_user_agentstringНеобязательныйcustom user agent optional field custom user agent for crawling a website example: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.116 Safari/537.36 default value: Mozilla/5.0 (compatible; RSiteAuditor)
browser_presetstringНеобязательныйpreset for browser screen parameters optional field if you use this field, you don’t need to indicate browser_screen_width, browser_screen_height, browser_screen_scale_factor possible values: desktop, mobile, tablet desktop preset will apply the following values: browser_screen_width: 1920 browser_screen_height: 1080 browser_screen_scale_factor: 1 mobile preset will apply the following values: browser_screen_width: 390 browser_screen_height: 844 browser_screen_scale_factor: 3 tablet preset will apply the following values: browser_screen_width: 1024 browser_screen_height: 1366 browser_screen_scale_factor: 2 Note: to use this parameter, set enable_javascript or enable_browser_rendering to true
browser_screen_widthintegerНеобязательныйbrowser screen width optional field you can set a custom browser screen width to perform audit for a particular device; if you use this field, you don’t need to indicate browser_preset as it will be ignored; Note: to use this parameter, set enable_javascript or enable_browser_rendering to true minimum value, in pixels: 240 maximum value, in pixels: 9999
browser_screen_heightintegerНеобязательныйbrowser screen height optional field you can set a custom browser screen height to perform audit for a particular device; if you use this field, you don’t need to indicate browser_preset as it will be ignored; Note: to use this parameter, set enable_javascript or enable_browser_rendering to true minimum value, in pixels: 240 maximum value, in pixels: 9999
browser_screen_scale_factornumberНеобязательныйbrowser screen scale factor optional field you can set a custom browser screen resolution ratio to perform audit for a particular device; if you use this field, you don’t need to indicate browser_preset as it will be ignored; Note: to use this parameter, set enable_javascript or enable_browser_rendering to true minimum value: 0.5 maximum value: 3
store_raw_htmlbooleanНеобязательныйstore HTML of a crawled page optional field set to true if you want to get the HTML of the page using the OnPage Raw HTML endpoint default value: false
disable_cookie_popupbooleanНеобязательныйdisable the cookie popup optional field set to true if you want to disable the popup requesting cookie consent from the user; default value: false
accept_languagestringНеобязательныйlanguage header for accessing the website optional field all locale formats are supported (xx, xx-XX, xxx-XX, etc.) Note: if you do not specify this parameter, some websites may deny access; in this case, pages will be returned with the "type":"broken in the response array
enable_javascriptbooleanНеобязательныйload javascript on a page optional field set to true if you want to load the scripts available on a page default value: false Note: if you use this parameter, additional charges will apply; learn more about the cost of tasks with this parameter in our help article; the cost can be calculated on the Pricing Page
enable_browser_renderingbooleanНеобязательныйemulate browser rendering to measure Core Web Vitals optional field by using this parameter you will be able to emulate a browser when loading a web page; enable_browser_rendering loads styles, images, fonts, animations, videos, and other resources on a page; default value: false set to true to obtain Core Web Vitals (FID, CLS, LCP) metrics in the response; if you use this field, enable_javascript, and load_resources parameters must be set to true Note: if you use this parameter, additional charges will apply; learn more about the cost of tasks with this parameter in our help article; the cost can be calculated on the Pricing Page
enable_xhrbooleanНеобязательныйenable XMLHttpRequest on a page optional field set to true if you want our crawler to request data from a web server using the XMLHttpRequest object default value: false if you use this field, enable_javascript must be set to true;
switch_poolbooleanНеобязательныйswitch proxy pool optional field if true, additional proxy pools will be used to obtain the requested data; the parameter can be used if a multitude of tasks is set simultaneously, resulting in occasional rate-limit and/or site_unreachable errors
ip_pool_for_scanstringНеобязательныйproxy pool optional field you can choose a location of the proxy pool that will be used to obtain the requested data; the parameter can be used if page content is inaccessible in one of the locations, resulting in occasional site_unreachable errors possible values: us, de
markdown_viewbooleanНеобязательныйreturn page content as markdown optional field if set to true, the markdown-formatted content of the page will be returned in the page_as_markdown field of the response; default value: false
Пример запроса
curl -X POST "https://openapi.felo.ai/v1/beta/dataforseo/on_page/content_parsing/live" \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "<string>",
"custom_user_agent": "<string>",
"browser_preset": "<string>",
"browser_screen_width": 0,
"browser_screen_height": 0,
"browser_screen_scale_factor": 0,
"store_raw_html": false,
"disable_cookie_popup": false
}'Ответ
Поля ответа
versionstringНеобязательныйthe current version of the API
status_codeintegerНеобязательныйgeneral status code you can find the full list of the response codes here Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions
status_messagestringНеобязательныйgeneral informational message you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйtotal tasks cost, USD
tasks_countintegerНеобязательныйthe number of tasks in the tasks array
tasks_errorintegerНеобязательныйthe number of tasks in the tasks array returned with an error
tasksstring[]Необязательныйarray of tasks
idstringНеобязательныйtask identifier unique task identifier in our system in the UUID format
status_codeintegerНеобязательныйstatus code of the task generated by DataForSEO; can be within the following range: 10000-60000 you can find the full list of the response codes here
status_messagestringНеобязательныйinformational message of the task you can find the full list of general informational messages here
timestringНеобязательныйexecution time, seconds
costnumberНеобязательныйcost of the task, USD
result_countintegerНеобязательныйnumber of elements in the result array
pathstring[]НеобязательныйURL path
dataobjectНеобязательныйcontains the same parameters that you specified in the POST request
resultstring[]Необязательныйarray of results
crawl_progressstringНеобязательныйstatus of the crawling session possible values: in_progress, finished
crawl_statusobjectНеобязательныйdetails of the crawling session
items_countintegerНеобязательныйnumber of items in the results array
itemsstring[]Необязательныйitems array
typestringНеобязательныйtype of the returned item = ‘сontent_parsing_element’
fetch_timestringНеобязательныйdate and time when the content was fetched in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: "2022-11-01 10:02:52 +00:00"
status_codeintegerНеобязательныйstatus code of the page
page_contentobjectНеобязательныйparsed content of the page
headerobjectНеобязательныйparsed content of the header
primary_contentstring[]Необязательныйprimary content on the page you can find more information about content priority calculation in this help center article
textstringНеобязательныйcontent text
urlstringНеобязательныйpage URL displayed in case the text is a link anchor
urlsstring[]Необязательныйcontains other URLs and anchors found in the content element
urlstringНеобязательныйother URL found in the content element
anchor_textstringНеобязательныйtext of the URL’s anchor
secondary_contentstring[]Необязательныйsecondary content on the page you can find more information about content priority calculation in this help center article
textstringНеобязательныйcontent text
urlstringНеобязательныйpage URL displayed in case the text is a link anchor
urlsstring[]Необязательныйcontains other URLs and anchors found in the content element
urlstringНеобязательныйother URL found in the content element
anchor_textstringНеобязательныйtext of the URL’s anchor
table_contentstring[]Необязательныйcontent of the table on the page
headerstring[]Необязательныйcontent of the header of the table
row_cellsstring[]Необязательныйcontent of the row cells of the header
textstringНеобязательныйtext in the row cell
urlsstring[]Необязательныйcontains other URLs and anchors found in the cell
urlstringНеобязательныйURL found in the cell
anchor_textstringНеобязательныйtext of the URL’s anchor
is_headerbooleanНеобязательныйindicates if the text belongs to the header
bodystring[]Необязательныйcontent of the body of the table
row_cellsstring[]Необязательныйcontent of the row cells of the header
textstringНеобязательныйtext in the row cell
urlsstring[]Необязательныйcontains other URLs and anchors found in the cell
urlstringНеобязательныйURL found in the cell
anchor_textstringНеобязательныйtext of the URL’s anchor
is_headerbooleanНеобязательныйindicates if the text belongs to the header
footerstring[]Необязательныйcontent of the footer of the table
row_cellsstring[]Необязательныйcontent of the row cells of the header
textstringНеобязательныйtext in the row cell
urlsstring[]Необязательныйcontains other URLs and anchors found in the cell
urlstringНеобязательныйURL found in the cell
anchor_textstringНеобязательныйtext of the URL’s anchor
is_headerbooleanНеобязательныйindicates if the text belongs to the header
footerobjectНеобязательныйparsed content of the footer
main_topicstring[]Необязательныйmain topic on the page you can find more information about topic priority calculation in this help center article
h_titlestringНеобязательныйmeta title
main_titlestringНеобязательныйmain title of the block
authorstringНеобязательныйcontent author name
languagestringНеобязательныйcontent language
levelstringНеобязательныйHTML level
primary_contentstring[]Необязательныйprimary content on the page you can find more information about content priority calculation in this help center article
textstringНеобязательныйcontent text
urlstringНеобязательныйpage URL displayed in case the text is a link anchor
urlsstring[]Необязательныйcontains other URLs and anchors found in the content element
urlstringНеобязательныйother URL found in the content element
anchor_textstringНеобязательныйtext of the URL’s anchor
secondary_contentstring[]Необязательныйsecondary content on the page you can find more information about content priority calculation in this help center article
secondary_topicstring[]Необязательныйsecondary topic on the page you can find more information about topic priority calculation in this help center article
h_titlestringНеобязательныйmeta title
main_titlestringНеобязательныйmain title of the block
authorstringНеобязательныйcontent author name
languagestringНеобязательныйcontent language
levelstringНеобязательныйHTML level
primary_contentstring[]Необязательныйprimary content on the page you can find more information about content priority calculation in this help center article
textstringНеобязательныйcontent text
urlstringНеобязательныйpage URL displayed in case the text is a link anchor
urlsstring[]Необязательныйcontains other URLs and anchors found in the content element
urlstringНеобязательныйother URL found in the content element
anchor_textstringНеобязательныйtext of the URL’s anchor
secondary_contentstring[]Необязательныйsecondary content on the page you can find more information about content priority calculation in this help center article
ratingsstring[]Необязательныйcontains objects with rating information for the products displayed on the page
namestringНеобязательныйrating name Note: this field is not used in this particular object, and its value is always set to null
rating_valueintegerНеобязательныйthe value of the rating
max_rating_valueintegerНеобязательныйmaximum value for the rating
rating_countintegerНеобязательныйthe amount of feedback
relative_ratingnumberНеобязательныйrelative rating can take values from 0 to 1
offersstring[]Необязательныйarray of products displayed on the page contains objects with information on products displayed on the page
namestringНеобязательныйname of the product
priceintegerНеобязательныйprice of the product
price_currencystringНеобязательныйprice currency
price_valid_untilintegerНеобязательныйdisplays the date and time until which the price is valid in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00” example: "2022-11-01 10:02:52 +00:00"
commentsstring[]Необязательныйarray of comments displayed on the page contains objects with information on comments related to displayed products
ratingobjectНеобязательныйproduct’s rating contains information about the rating a customer has given to the product
namestringНеобязательныйrating name Note: this field is not used in this particular object, and its value is always null
rating_valueintegerНеобязательныйthe value of the rating
max_rating_valueintegerНеобязательныйmaximum value for the rating
rating_countintegerНеобязательныйthe amount of feedback Note: this field is not used in this particular object, and its value is always null
relative ratingnumberНеобязательныйrelative rating can take values from 0 to 1
titlestringНеобязательныйtitle of the customer’s comment
publish_datestringНеобязательныйdate when the comment was published
authorstringНеобязательныйauthor of the comment
primary_contentstring[]Необязательныйprimary content on the page you can find more information about content priority calculation in this help center article
textstringНеобязательныйtext of the comment
urlstringНеобязательныйdisplayed in case the text is a link anchor
urlsstring[]Необязательныйcontains other URLs and anchors found in the content element
contactsobjectНеобязательныйcontact information contains contact information displayed on the page
telephonesstring[]Необязательныйarray of telephone numbers
emailsstring[]Необязательныйarray of emails
page_as_markdownstringНеобязательныйpage content in the markdown format page content in the text-to-HTML markdown format specify markdown_view as true in the request to return the value