
Convert Webpage to Markdown
Convert a webpage into clean Markdown for content pipelines
Extract readable articles from webpages or HTML.
Extract the main article with useful text formatting.
Use extracted text for indexing and content workflows.
Read the title, author and language when available.
Submit a public page URL or raw HTML.
Remove navigation and surrounding page elements.
Prepare articles for search, review and summarization.

HTTP Protocol:HTTPS
HTTP Method:POST
HTTP Endpoint:https://api.gugudata.io/v1/websitetools/readability
Response Type:application/json; charset=utf-8
DEMO Endpoint:https://api.gugudata.io/v1/websitetools/readability/demo
Live Demo:Try Interactive Demo
Full API Docs:developers.gugudata.io
| Name | Type | Is Required | Default Value | Remark |
|---|---|---|---|---|
| appkey | string | true | YOUR_APPKEY | Application key used for request authentication. Supply the value as a query parameter, form field, or multipart field according to the request content type. |
| html | string | false | YOUR_VALUE | Raw HTML content. Supply either `html` or `url`. |
| url | string | false | YOUR_VALUE | Target webpage URL. Supply either `url` or `html`. |
| Name | Type | Remark |
|---|---|---|
| DataStatus.RequestParameter | string | Normalized request parameters echoed by the service. Sensitive credentials are omitted when available. |
| DataStatus.StatusCode | integer | Application-level status code returned by the current v1 contract. |
| DataStatus.StatusDescription | string | Application-level status message returned by the current v1 contract. |
| DataStatus.ResponseDateTime | string | Response timestamp returned by the current service contract. |
| DataStatus.DataTotalCount | integer | Total number of records that match the request. |
| Data.Title | string | Article title |
| Data.Byline | string | Article author |
| Data.Dir | string | Article text direction |
| Data.Lang | string | Article language |
| Data.Content | string | Article content |
| Data.TextContent | string | Article content (without HTML tags, divided by paragraphs) |
| Data.Length | integer | Article length |
| Data.Excerpt | string | Article excerpt |
| Data.SiteName | string | Website name |
| Data.PublishedTime | array<string> | Article publication time |
| Status Code | Explanation of Status Code | Remarks |
|---|---|---|
| 200 | Request processed successfully. | Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`. |
| 400 | Invalid request parameters or request format. | Check required fields, data types, and request body format. |
| 401 | Missing or unknown application key. | Provide a valid `appkey` with the request. |
| 403 | The application key is recognized but access is not allowed. | The key may be expired, inactive, or not permitted for the requested API. |
| 429 | Request rate or trial usage limit exceeded. | Reduce concurrency or retry after the limit window resets. |
| 500 | Internal service error. | Retry later or contact support if the error persists. |
| 503 | Upstream service unavailable. | Retry later; the requested upstream dependency is temporarily unavailable. |
Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.
https://mcp.gugudata.io/mcp{
"mcpServers": {
"gugudata": {
"url": "https://mcp.gugudata.io/mcp",
"transportType": "streamable-http"
}
}
}
Convert a webpage into clean Markdown for content pipelines

Extract hyperlinks and destinations from a webpage

Extract structured, readable content from a webpage or raw HTML

Extract clean, LLM-ready article content and metadata