跳到正文

个人用户想在线解析、翻译 PDF,或给网页版、客户端、Zotero 插件开会员,请前往 Doc2X 网页版。本站只出售 API 调用额度,两边互不通用。

去 Doc2X 网页版

PDF 解析 API

PDF 转 Markdown / JSON,公式、表格、多栏版面一次还原

上传 PDF,得到按阅读顺序排列的逐页 Markdown;选用 v3-2026 模型还能拿到版面块、表格 HTML 和图片裁切。

状态
已上线
接口
POST /api/v2/parse/preupload,GET /api/v2/parse/status
输入
PDF 文件(含扫描版)
输出
逐页 Markdown;v3-2026 模型额外返回 JSON 版面块
计费
按解析页数
结果保留
24 小时,请及时取回
价格
¥0.02/页起,资源包低至 ¥0.01/页

效果示例

两份真实样张(学术论文与高考试卷)的解析结果。点选方框定位到对应段落,可以在渲染预览、Markdown 和 JSON 之间切换。

Attention Is All You Need · 第 6 页原页
  • 图注
  • 表格
  • 标题
  • 正文
  • 公式
解析完成attention-is-all-you-need.pdf14 个段落

Table 1: Maximum path lengths, per-layer complexity and minimum number of sequential operations for different layer types. is the sequence length, is the representation dimension, is the kernel size of convolutions and the size of the neighborhood in restricted self-attention.

Layer Type

Complexity per Layer

Sequential Operations

Maximum Path Length

Self-Attention

Recurrent

Convolutional

Self-Attention (restricted)

3.5 Positional Encoding

Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].

In this work, we use sine and cosine functions of different frequencies:

where pos is the position and is the dimension. That is,each dimension of the positional encoding corresponds to a sinusoid. The wavelengths form a geometric progression from to . We chose this function because we hypothesized it would allow the model to easily learn to attend by relative positions,since for any fixed offset can be represented as a linear function of .

We also experimented with using learned positional embeddings [9] instead, and found that the two versions produced nearly identical results (see Table 3 row (E)). We chose the sinusoidal version because it may allow the model to extrapolate to sequence lengths longer than the ones encountered during training.

4 Why Self-Attention

In this section we compare various aspects of self-attention layers to the recurrent and convolutional layers commonly used for mapping one variable-length sequence of symbol representations to another sequence of equal length ,with ,such as a hidden layer in a typical sequence transduction encoder or decoder. Motivating our use of self-attention we consider three desiderata.

One is the total computational complexity per layer. Another is the amount of computation that can be parallelized, as measured by the minimum number of sequential operations required.

The third is the path length between long-range dependencies in the network. Learning long-range dependencies is a key challenge in many sequence transduction tasks. One key factor affecting the ability to learn such dependencies is the length of the paths forward and backward signals have to traverse in the network. The shorter these paths between any combination of positions in the input and output sequences, the easier it is to learn long-range dependencies [12]. Hence we also compare the maximum path length between any two input and output positions in networks composed of the different layer types.

As noted in Table 1, a self-attention layer connects all positions with a constant number of sequentially executed operations,whereas a recurrent layer requires sequential operations. In terms of computational complexity, self-attention layers are faster than recurrent layers when the sequence

渲染 SDK 输出,与网页版、控制台看到的效果一致

样张:Attention Is All You Need · 第 6 页。左侧方框是 Doc2X 识别出的段落,按类型着色;点方框可以定位到右侧对应的结果。

能力说明

公式与表格
数学公式输出为 LaTeX,表格输出为 HTML;导出 Word 时可合并跨页表格。
复杂版面
多栏排版、图片与图注、页眉页脚都按阅读顺序还原,适合论文、教材和研报。
两种模型
v2 模型稳定成熟;v3-2026 模型额外返回版面块、父子关系、阅读顺序和图片裁切地址。
异步直传
预上传接口直接把文件传到存储,再轮询任务状态,适合批量和大文件。

调用示例

在请求头加入 Authorization: Bearer sk-xxx 即可调用,API Key 在控制台创建。

完整参数与返回说明
# 1. 申请上传地址(model 可选 v2 或 v3-2026)
curl -X POST 'https://v2.doc2x.noedgeai.com/api/v2/parse/preupload' \
  -H 'Authorization: Bearer sk-xxx' \
  -H 'Content-Type: application/json' \
  -d '{"model": "v3-2026"}'

# 2. 用返回的 url 上传 PDF
curl -X PUT "<url>" --data-binary @paper.pdf

# 3. 轮询解析状态,success 后读取逐页 Markdown
curl 'https://v2.doc2x.noedgeai.com/api/v2/parse/status?uid=<uid>' \
  -H 'Authorization: Bearer sk-xxx'

适用场景

知识库与 RAG

把制度文件、产品手册、研报解析成干净的 Markdown,按标题和版面块切分入库,召回更准。

论文与教材数字化

理工科资料里的公式、表格、参考文献完整保留,可再导出 Word 或 LaTeX 继续编辑。

题库与教辅建设

试卷和教辅批量转成结构化文本,配合图片解析处理拍照题目。

常见问题

支持扫描版 PDF 吗?

支持。扫描页会先做文字识别,再和电子版一样输出 Markdown 与结构化结果。

v2 和 v3-2026 模型怎么选?

只需要 Markdown 文本时用默认的 v2;需要版面块、阅读顺序、表格 HTML 或图片裁切(例如做 RAG 切分、版面重排)时,在预上传时传入 model: v3-2026。

解析结果保存多久?

结果临时保存 24 小时,请在任务成功后尽快取回并保存到自己的存储。

开始调用PDF 解析 API

创建 API Key 即可调用;大批量、定制或私有化需求请联系商务。