版面分析 API
v3 结构化 JSON:版面块、阅读顺序、表格 HTML、图片裁切
在 PDF 解析时选用 v3-2026 模型,额外得到每页的版面块和它们之间的父子关系,方便做 RAG 切分和版面重排。
- 状态
- 已上线
- 启用方式
- POST /api/v2/parse/preupload 时传入 model: v3-2026
- 输出
- page.md 与 page.layout.blocks
- 计费
- 按解析页数,与 PDF 解析相同
- 价格
- 随 PDF 解析按页计费,¥0.02/页起
效果示例
JSON 视图列出每个段落的 id、坐标(PDF 点,原点在左上角)与文本,按阅读顺序排列。v3-2026 模型的完整字段见 v3 JSON 文档。

- 图注
- 表格
- 标题
- 正文
- 公式
解析完成attention-is-all-you-need.pdf14 个段落
[]
{
"id": "attention:0:0",
"bbox": [107, 69, 503, 102],
"text": "Table 1: Maximum path lengths, per-layer complexity and minimum number of sequential operations for different layer types. \\( n \\) is the sequence length, \\( d \\) is the representation dimension, \\( k \\) is the kernel size of convolutions and \\( r \\) the size of the neighborhood in restricted self-attention."
},{
"id": "attention:0:1",
"bbox": [124, 114, 486, 189],
"text": "<table><tr><td>Layer Type</td><td>Complexity per Layer</td><td>Sequential Operations</td><td>Maximum Path Length</td></tr><tr><td>Self-Attention</td><td>\\( O\\left( {{n}^{2} \\cdot d}\\right) \\)</td><td>\\( O\\left( 1\\right) \\)</td><td>\\( O\\left( 1\\right) \\)</td></tr><tr><td>Recurrent</td><td>\\( O\\left( {n \\cdot {d}^{2}}\\right) \\)</td><td>\\( O\\left( n\\right) \\)</td><td>\\( O\\left( n\\right) \\)</td></tr><tr><td>Convolutional</td><td>\\( O\\left( {k \\cdot n \\cdot {d}^{2}}\\right) \\)</td><td>\\( O\\left( 1\\right) \\)</td><td>\\( O\\left( {{\\log }_{k}\\left( n\\right) }\\right) \\)</td></tr><tr><td>Self-Attention (restricted)</td><td>\\( O\\left( {r \\cdot n \\cdot d}\\right) \\)</td><td>\\( O\\left( 1\\right) \\)</td><td>\\( O\\left( {n/r}\\right) \\)</td></tr></table>"
},{
"id": "attention:0:2",
"bbox": [108, 211, 215, 223],
"text": "### 3.5 Positional Encoding"
},{
"id": "attention:0:3",
"bbox": [108, 232, 505, 298],
"text": "Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add \"positional encodings\" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension \\( {d}_{\\text{ model }} \\) as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9]."
},{
"id": "attention:0:4",
"bbox": [108, 303, 389, 315],
"text": "In this work, we use sine and cosine functions of different frequencies:"
},{
"id": "attention:0:5",
"bbox": [235, 336, 385, 349],
"text": "\\[P{E}_{\\left( pos,2i\\right) } = \\sin \\left( {{pos}/{10000}^{{2i}/{d}_{\\text{ model }}}}\\right)\\]"
},{
"id": "attention:0:6",
"bbox": [225, 353, 385, 366],
"text": "\\[P{E}_{\\left( pos,2i + 1\\right) } = \\cos \\left( {{pos}/{10000}^{{2i}/{d}_{\\text{ model }}}}\\right)\\]"
},{
"id": "attention:0:7",
"bbox": [107, 375, 503, 431],
"text": "where pos is the position and \\( i \\) is the dimension. That is,each dimension of the positional encoding corresponds to a sinusoid. The wavelengths form a geometric progression from \\( {2\\pi } \\) to \\( {10000} \\cdot {2\\pi } \\) . We chose this function because we hypothesized it would allow the model to easily learn to attend by relative positions,since for any fixed offset \\( k,P{E}_{{pos} + k} \\) can be represented as a linear function of \\( P{E}_{pos} \\) ."
},{
"id": "attention:0:8",
"bbox": [107, 435, 503, 479],
"text": "We also experimented with using learned positional embeddings [9] instead, and found that the two versions produced nearly identical results (see Table 3 row (E)). We chose the sinusoidal version because it may allow the model to extrapolate to sequence lengths longer than the ones encountered during training."
},{
"id": "attention:0:9",
"bbox": [108, 494, 225, 509],
"text": "## 4 Why Self-Attention"
},{
"id": "attention:0:10",
"bbox": [106, 519, 504, 574],
"text": "In this section we compare various aspects of self-attention layers to the recurrent and convolutional layers commonly used for mapping one variable-length sequence of symbol representations \\( \\left( {{x}_{1},\\ldots ,{x}_{n}}\\right) \\) to another sequence of equal length \\( \\left( {{z}_{1},\\ldots ,{z}_{n}}\\right) \\) ,with \\( {x}_{i},{z}_{i} \\in {\\mathbb{R}}^{d} \\) ,such as a hidden layer in a typical sequence transduction encoder or decoder. Motivating our use of self-attention we consider three desiderata."
},{
"id": "attention:0:11",
"bbox": [108, 579, 504, 601],
"text": "One is the total computational complexity per layer. Another is the amount of computation that can be parallelized, as measured by the minimum number of sequential operations required."
},{
"id": "attention:0:12",
"bbox": [107, 607, 503, 684],
"text": "The third is the path length between long-range dependencies in the network. Learning long-range dependencies is a key challenge in many sequence transduction tasks. One key factor affecting the ability to learn such dependencies is the length of the paths forward and backward signals have to traverse in the network. The shorter these paths between any combination of positions in the input and output sequences, the easier it is to learn long-range dependencies [12]. Hence we also compare the maximum path length between any two input and output positions in networks composed of the different layer types."
},{
"id": "attention:0:13",
"bbox": [107, 688, 503, 721],
"text": "As noted in Table 1, a self-attention layer connects all positions with a constant number of sequentially executed operations,whereas a recurrent layer requires \\( O\\left( n\\right) \\) sequential operations. In terms of computational complexity, self-attention layers are faster than recurrent layers when the sequence"
}段落 id、坐标(PDF 点)与文本
能力说明
- 细粒度版面块
- 标题、正文、图注、图片、表格、公式、公式编号,以及图组、表组、目录、脚注、参考文献等分组。
- 关系与顺序
- 块之间的父子关系和页内阅读顺序直接给出,不需要自己推断。
- 可用的附件
- 表格块带 HTML,图片块带裁切后的图片地址;页眉页脚等样板内容有单独标记。
调用示例
在请求头加入 Authorization: Bearer sk-xxx 即可调用,API Key 在控制台创建。
{
"page_idx": 0,
"layout": {
"blocks": [
{ "id": "b1", "type": "Title", "text": "1 Introduction" },
{ "id": "b2", "type": "Text", "text": "Large language models ..." },
{ "id": "b3", "type": "Table", "html": "<table>...</table>" },
{ "id": "b4", "type": "Figure", "url": "https://.../crop.png" }
]
}
}适用场景
RAG 精准切分
按标题和版面块切分,表格、公式独立成块,去掉页眉页脚,检索噪声更少。
版面重排与再出版
利用阅读顺序和图表分组重新排版,生成网页、电子书或新的教辅版式。
常见问题
版面分析需要单独调用接口吗?
不需要,它是 PDF 解析在 v3-2026 模型下的扩展输出,同一个任务同时返回 Markdown 和版面块。
开始调用版面分析 API
创建 API Key 即可调用;大批量、定制或私有化需求请联系商务。
