Bestseller
Multiple Format Output Highly Accurate Recognition

PDF Parsing and Formatted Output

Extract PDF text as HTML, text, Markdown, or JSON

What can this API do?

Four useful text formats

Choose HTML, plain text, Markdown, or a JSON string.

Page-aware output

Keep the original page order, with numbered sections in HTML, Markdown, and JSON.

Multilingual documents

Extract readable text already present in PDFs, including Chinese and English.

Upload or link input

Submit a PDF file or a publicly accessible PDF URL.

Browser-ready text

Display extracted text as a readable UTF-8 HTML document.

Document workflows

Prepare PDF text for search, indexing, and content processing.

PDF Parsing and Formatted Output
Annual plan
$19$49
Try it for free!
Sign in to get a trial key and test all APIs.
Sign In
Secure checkout powered by Stripe

API Document

HTTP Protocol:HTTPS

HTTP Method:POST

HTTP Endpoint:https://api.gugudata.io/v1/imagerecognition/pdf2format

Response Type:application/json; charset=utf-8

DEMO Endpoint:https://api.gugudata.io/v1/imagerecognition/pdf2format/demo

Live Demo:Try Interactive Demo

Full API Docs:developers.gugudata.io

API Request Parameters

NameTypeIs RequiredDefault ValueRemark
appkeystringtrueYOUR_APPKEYYour API key. Send it as a query parameter.
typestringtruehtmlOutput format: html, text (alias txt), markdown (alias md), or json. The result is always a string.
pdffilefilefalsePDF upload. Provide exactly one of pdffile, file, or file_url. Maximum 20 MiB and 100 pages. Encrypted, damaged, and textless PDFs are not supported.
filefilefalseCompatible alias for pdffile. Do not combine it with another input.
file_urlstringfalsehttps://storage.gugudata.io/pdf/demo.pdfPublic HTTP or HTTPS PDF URL, as an alternative to a file upload.

API Response Parameters

NameTypeRemark
dataStatus.requestParameterstringInput source, selected format, file size, and page count. The Demo shows its public sample PDF URL.
dataStatus.statusCodeintegerBusiness status code: 200 for success.
dataStatus.statusDescriptionstringBusiness status description.
dataStatus.responseDateTimestringResponse date and time.
dataStatus.dataTotalCountintegerNumber of conversion results returned.
data.resultstringExtracted HTML, plain text, Markdown, or serialized JSON containing ordered page text. Parse this string again when type=json. Does not include OCR, images, or complex layout reconstruction.

API Response Status Codes

Status CodeExplanation of Status CodeRemarks
200Request processed successfully.Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`.
400Invalid request parameters or request format.Check required fields, data types, and request body format.
401Missing or unknown application key.Provide a valid `appkey` with the request.
403The application key is recognized but access is not allowed.The key may be expired, inactive, or not permitted for the requested API.
413PDF exceeds the size or page limit.Use a PDF of at most 20 MiB and 100 pages.
422PDF cannot be converted.Check that the PDF is readable, unencrypted, and contains text. Output must not exceed 20 MiB.
429Request rate or trial usage limit exceeded.Reduce concurrency or retry after the limit window resets.
500Internal service error.Retry later or contact support if the error persists.
502PDF retrieval or conversion response failed.Check that the public PDF URL can be downloaded.
503Conversion is temporarily unavailable.Try again later.

Remote MCP

AI client access

Use this API through GuGuData Remote MCP

Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.

Endpoint https://mcp.gugudata.io/mcp
Open MCP Guide
Generic MCP configuration
{
  "mcpServers": {
    "gugudata": {
      "url": "https://mcp.gugudata.io/mcp",
      "transportType": "streamable-http"
    }
  }
}

Code Snippets