GHSA-CGC7-9QP3-86M3
Vulnerability from github – Published: 2026-10-07 20:41 – Updated: 2026-10-07 20:41Summary
The HTML, JATS, ODS (OpenDocument spreadsheet) and BoxNote backends accept table rowspan / colspan values without an upper bound. A few bytes of input, such as <td rowspan="100000000">, make docling run loops proportional to the declared span and allocate a table grid of the declared size. The result is CPU and memory exhaustion.
Details
docling/backend/html_backend.py(_get_cell_spans) parses span attributes with no upper limit. The cell-filling loop then iteratesrow_span × col_spantimes.docling/backend/jats_backend.pyanddocling/backend/boxnote_backend.pyfill their tables the same way.- The OpenDocument spreadsheet path scans the declared span range.
- Export (for example
export_to_markdown()) materialises the full grid throughTableData.gridin docling-core.
document_timeout does not bound this. It is checked between pipeline stages, and these backends convert the whole document in a single call. max_file_size and max_num_pages do not help because the payload is tiny.
Measured on 2.130.0: a 54-byte HTML file with rowspan="1e8" takes about 4.4 s of CPU, and the time grows linearly with the value. A 52-byte file with colspan="3000000" takes about 23 s and reaches 4.5 GB peak memory during Markdown export.
Impact
Denial of service of the converting process from a very small input document. Confidentiality and integrity are not affected.
Proof of concept
<table><tr><td colspan="3000000">x</td></tr></table>
from docling.document_converter import DocumentConverter
DocumentConverter().convert("span.html").document.export_to_markdown()
Patches
Fixed in docling 2.131.0 by #4414. Table spans are clamped to the HTML limits (colspan 1000, rowspan 65534) and to the actual size of the table in the HTML, JATS, BoxNote and OpenDocument spreadsheet backends, so conversion time and memory grow with the real table only.
Workarounds
Upgrade to 2.131.0. For older versions:
Run conversions of untrusted documents in a separate process with memory and CPU-time limits, or restrict allowed_formats to formats that are not affected.
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "docling"
},
"ranges": [
{
"events": [
{
"introduced": "2.0.0"
},
{
"fixed": "2.131.0"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "docling-slim"
},
"ranges": [
{
"events": [
{
"introduced": "2.92.0"
},
{
"fixed": "2.131.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-105749"
],
"database_specific": {
"cwe_ids": [
"CWE-400",
"CWE-789"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-07T20:41:02Z",
"nvd_published_at": "2026-10-05T22:16:57Z",
"severity": "MODERATE"
},
"details": "### Summary\n\nThe HTML, JATS, ODS (OpenDocument spreadsheet) and BoxNote backends accept table `rowspan` / `colspan` values without an upper bound. A few bytes of input, such as `\u003ctd rowspan=\"100000000\"\u003e`, make docling run loops proportional to the declared span and allocate a table grid of the declared size. The result is CPU and memory exhaustion.\n\n### Details\n\n- `docling/backend/html_backend.py` (`_get_cell_spans`) parses span attributes with no upper limit. The cell-filling loop then iterates `row_span \u00d7 col_span` times.\n- `docling/backend/jats_backend.py` and `docling/backend/boxnote_backend.py` fill their tables the same way.\n- The OpenDocument spreadsheet path scans the declared span range.\n- Export (for example `export_to_markdown()`) materialises the full grid through `TableData.grid` in docling-core.\n\n`document_timeout` does not bound this. It is checked between pipeline stages, and these backends convert the whole document in a single call. `max_file_size` and `max_num_pages` do not help because the payload is tiny.\n\nMeasured on 2.130.0: a 54-byte HTML file with `rowspan=\"1e8\"` takes about 4.4 s of CPU, and the time grows linearly with the value. A 52-byte file with `colspan=\"3000000\"` takes about 23 s and reaches 4.5 GB peak memory during Markdown export.\n\n### Impact\n\nDenial of service of the converting process from a very small input document. Confidentiality and integrity are not affected.\n\n### Proof of concept\n\n```html\n\u003ctable\u003e\u003ctr\u003e\u003ctd colspan=\"3000000\"\u003ex\u003c/td\u003e\u003c/tr\u003e\u003c/table\u003e\n```\n\n```python\nfrom docling.document_converter import DocumentConverter\nDocumentConverter().convert(\"span.html\").document.export_to_markdown()\n```\n\n### Patches\n\nFixed in docling 2.131.0 by [#4414](https://github.com/docling-project/docling/pull/4414). Table spans are clamped to the HTML limits (colspan 1000, rowspan 65534) and to the actual size of the table in the HTML, JATS, BoxNote and OpenDocument spreadsheet backends, so conversion time and memory grow with the real table only.\n\n### Workarounds\n\nUpgrade to 2.131.0. For older versions:\n\nRun conversions of untrusted documents in a separate process with memory and CPU-time limits, or restrict `allowed_formats` to formats that are not affected.",
"id": "GHSA-cgc7-9qp3-86m3",
"modified": "2026-10-07T20:41:02Z",
"published": "2026-10-07T20:41:02Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/docling-project/docling/security/advisories/GHSA-cgc7-9qp3-86m3"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-105749"
},
{
"type": "WEB",
"url": "https://github.com/docling-project/docling/pull/4414"
},
{
"type": "WEB",
"url": "https://github.com/docling-project/docling/commit/c5b4429cc6500a344c13edeb22e67610c2159b09"
},
{
"type": "PACKAGE",
"url": "https://github.com/docling-project/docling"
},
{
"type": "WEB",
"url": "https://github.com/docling-project/docling/releases/tag/v2.131.0"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "Docling: Unbounded table rowspan/colspan in HTML, JATS, ODS and BoxNote backends causes CPU/memory exhaustion"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.