
LLM-ready Article Content Extraction API
Extract clean, LLM-ready article content and metadata
Extract structured, readable content from a webpage or raw HTML
Extract article content from a live public webpage or from HTML supplied directly by your application.
Receive the title, description, author, publication date, source, favicon, and article type in one record.
Isolate the primary article content from surrounding navigation, promotional blocks, and page furniture.
Capture page links, the lead image, image alternatives, and related visual context when available.
Use the returned reading-time value to support previews, prioritization, and editorial planning.
Standardize articles for monitoring, archiving, content analysis, and downstream knowledge workflows.

HTTP Protocol:HTTPS
HTTP Method:POST
HTTP Endpoint:https://api.gugudata.io/v1/article/extract
Response Type:application/json; charset=utf-8
DEMO Endpoint:https://api.gugudata.io/v1/article/extract/demo
Live Demo:Try Interactive Demo
Full API Docs:developers.gugudata.io
| Name | Type | Is Required | Default Value | Remark |
|---|---|---|---|---|
| appkey | string | true | YOUR_APPKEY | Application key used for request authentication. Supply the value as a query parameter, form field, or multipart field according to the request content type. |
| url | string | true | Target webpage URL. |
| Name | Type | Remark |
|---|---|---|
| DataStatus.StatusCode | integer | Application-level status code returned by the current v1 contract. |
| DataStatus.StatusDescription | string | Application-level status message returned by the current v1 contract. |
| DataStatus.ResponseDateTime | string | Response timestamp returned by the current service contract. |
| DataStatus.DataTotalCount | integer | Total number of records that match the request. |
| Data.url | string | Source URL of the article |
| Data.title | string | Extracted article title |
| Data.description | string | Article description/summary |
| Data.links | array<string> | Array of links contained in the article |
| Data.image | string | Main article image URL |
| Data.content | string | Extracted article content (HTML format, with ads and navigation removed) |
| Data.author | string | Article author (if available, may be empty string) |
| Data.favicon | string | Website favicon URL |
| Data.source | string | Source website domain (e.g., sohu.com) |
| Data.published | string | Article publication date/time (format: YYYY-MM-DD HH:MM) |
| Data.ttr | integer | Estimated reading time (Time to Read, in minutes) |
| Data.type | string | Article type (e.g., news, article, etc.) |
| Status Code | Explanation of Status Code | Remarks |
|---|---|---|
| 200 | Request processed successfully. | Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`. |
| 400 | Invalid request parameters or request format. | Check required fields, data types, and request body format. |
| 401 | Missing or unknown application key. | Provide a valid `appkey` with the request. |
| 403 | The application key is recognized but access is not allowed. | The key may be expired, inactive, or not permitted for the requested API. |
| 429 | Request rate or trial usage limit exceeded. | Reduce concurrency or retry after the limit window resets. |
| 500 | Internal service error. | Retry later or contact support if the error persists. |
| 503 | Upstream service unavailable. | Retry later; the requested upstream dependency is temporarily unavailable. |
Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.
https://mcp.gugudata.io/mcp{
"mcpServers": {
"gugudata": {
"url": "https://mcp.gugudata.io/mcp",
"transportType": "streamable-http"
}
}
}
Extract clean, LLM-ready article content and metadata

Extract ordered image candidates from an article

Extract hyperlinks and destinations from a webpage

Extract readable articles from webpages or HTML.