GHSA-G53G-W8RJ-FMG7
Vulnerability from github – Published: 2026-09-08 20:31 – Updated: 2026-09-08 20:31Summary
@xmldom/xmldom's processing-instruction (PI) grammar regex exhibits quadratic-time backtracking
(ReDoS) when parsing an unterminated processing instruction. A single small XML document
containing <? + a target + a long run of whitespace and no closing ?> forces the regular
expression engine into O(n²) work, stalling the Node.js event loop. The input is parsed with
DOMParser.parseFromString under default options, so it is reachable from unauthenticated,
network-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs).
Details
The PI production in lib/grammar.js compiles (flags mu) to:
^<\?(NameChars)(?:[\x20\x09\x0D\x0A]+([Char]*?))?\?>
^^^ S+ greedy ^^^ Char*? lazy
lib/grammar.jsline 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261
In the optional tail (?:S+(Char*?))?, both the greedy separator S+ and the lazy data Char*?
match XML whitespace. When the required trailing ?> is absent, the engine must ultimately fail —
but first it tries every partition of the whitespace run between S+ and Char*?, which is O(n²)
in the length of the trailing whitespace.
The regex is executed against the entire remaining source string in two places in lib/sax.js,
so the whole whitespace tail is scanned:
parsePI— https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691parseProcessingInstruction— https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879
Affected Versions
Only the 0.9.x line is affected. lib/grammar.js (and this PI regex) was introduced in
commit 726b471 ("fix!: preserve DOCTYPE internal subset (#498)"), first released in
0.9.0-beta.9, and is unchanged through 0.9.10.
The 0.8.x line (≤ 0.8.13) and the unscoped xmldom package (≤ 0.6.0) parse PIs via a different
code path bounded by indexOf('?>') — they do not contain this regex and are not affected
by this issue. (They were not separately tested for a different PI ReDoS; the scope here is the
specific grammar.js regex.)
| Line | PI code path | Affected? |
|---|---|---|
0.9.x (0.9.0-beta.9 … 0.9.10) |
grammar.js PI regex over full remaining source |
Yes |
0.8.x (≤ 0.8.13) |
parseInstruction, bounded by indexOf('?>') |
No |
unscoped xmldom (≤ 0.6.0) |
older indexOf('?>')-bounded parsing |
No |
Proof of Concept
const { DOMParser } = require('@xmldom/xmldom');
const n = 32 * 1024;
const payload = '<a><?p' + ' '.repeat(n); // unterminated PI, no `?>`
console.time('parse');
new DOMParser().parseFromString(payload, 'text/xml');
console.timeEnd('parse');
Measured (Node 18), trailing whitespace after <?p, no ?> — time quadruples per doubling of
input length (canonical O(n²)):
| Trailing whitespace | g.PI.exec |
parseFromString |
|---|---|---|
| 2 KB | 4.4 ms | 5.1 ms |
| 4 KB | 16.9 ms | 17.0 ms |
| 8 KB | 111.4 ms | 66.3 ms |
| 16 KB | 263.8 ms | 336.5 ms |
| 32 KB | 1073.1 ms | — |
Impact
Availability only: a single parse of a small crafted document blocks the Node.js event loop for the duration of the quadratic scan (≈1 s at 32 KB; multi-second with larger inputs). No memory blow-up, no data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and parsed with default options, one request can stall a server.
Fix Applied
Fixed in @xmldom/xmldom 0.9.11 (0.9.x-only; the 0.8.x LTS line and the
unscoped xmldom package use a different, bounded PI code path and are not affected).
PR #1039 inserts a fixed-width negative lookahead
(?!\s) immediately after the greedy S+, so the separator can no longer hand whitespace back to
the lazy data group:
- var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(', Char, '*?)'), '?', /\?>/);
+ var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(?!', _SChar, ')(', Char, '*?)'), '?', /\?>/);
The change is correct, minimal, and behavior-preserving: it produces identical [target, data]
captures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl
cases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead
is fixed-width and cannot itself backtrack — a strict improvement with no new parsing risk.
Severity note
The complexity is quadratic, not exponential, so a multi-second stall requires
tens-to-hundreds of KB of input. VA:H reflects that xmldom applies no input-size limit and the
path runs on default-options parsing, so a single unbounded parse can fully stall the event loop.
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 0.9.10"
},
"package": {
"ecosystem": "npm",
"name": "@xmldom/xmldom"
},
"ranges": [
{
"events": [
{
"introduced": "0.9.0-beta.9"
},
{
"fixed": "0.9.11"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-83606"
],
"database_specific": {
"cwe_ids": [
"CWE-1333",
"CWE-400"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-08T20:31:50Z",
"nvd_published_at": "2026-09-01T15:17:38Z",
"severity": "HIGH"
},
"details": "## Summary\n\n`@xmldom/xmldom`\u0027s processing-instruction (PI) grammar regex exhibits quadratic-time backtracking\n(ReDoS) when parsing an **unterminated** processing instruction. A single small XML document\ncontaining `\u003c?` + a target + a long run of whitespace and no closing `?\u003e` forces the regular\nexpression engine into O(n\u00b2) work, stalling the Node.js event loop. The input is parsed with\n`DOMParser.parseFromString` under **default options**, so it is reachable from unauthenticated,\nnetwork-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs).\n\n## Details\n\nThe PI production in `lib/grammar.js` compiles (flags `mu`) to:\n\n```\n^\u003c\\?(NameChars)(?:[\\x20\\x09\\x0D\\x0A]+([Char]*?))?\\?\u003e\n ^^^ S+ greedy ^^^ Char*? lazy\n```\n\n- `lib/grammar.js` line 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261\n\nIn the optional tail `(?:S+(Char*?))?`, both the greedy separator `S+` and the lazy data `Char*?`\nmatch XML whitespace. When the required trailing `?\u003e` is absent, the engine must ultimately fail \u2014\nbut first it tries every partition of the whitespace run between `S+` and `Char*?`, which is O(n\u00b2)\nin the length of the trailing whitespace.\n\nThe regex is executed against the **entire remaining source string** in two places in `lib/sax.js`,\nso the whole whitespace tail is scanned:\n\n- `parsePI` \u2014 https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691\n- `parseProcessingInstruction` \u2014 https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879\n\n## Affected Versions\n\nOnly the `0.9.x` line is affected. `lib/grammar.js` (and this PI regex) was introduced in\ncommit `726b471` (\"fix!: preserve DOCTYPE internal subset (#498)\"), first released in\n**0.9.0-beta.9**, and is unchanged through **0.9.10**.\n\nThe `0.8.x` line (\u2264 0.8.13) and the unscoped `xmldom` package (\u2264 0.6.0) parse PIs via a different\ncode path bounded by `indexOf(\u0027?\u003e\u0027)` \u2014 they do **not** contain this regex and are **not affected**\nby this issue. (They were not separately tested for a *different* PI ReDoS; the scope here is the\nspecific `grammar.js` regex.)\n\n| Line | PI code path | Affected? |\n|---|---|---|\n| `0.9.x` (0.9.0-beta.9 \u2026 0.9.10) | `grammar.js` `PI` regex over full remaining source | **Yes** |\n| `0.8.x` (\u2264 0.8.13) | `parseInstruction`, bounded by `indexOf(\u0027?\u003e\u0027)` | No |\n| unscoped `xmldom` (\u2264 0.6.0) | older `indexOf(\u0027?\u003e\u0027)`-bounded parsing | No |\n\n## Proof of Concept\n\n```js\nconst { DOMParser } = require(\u0027@xmldom/xmldom\u0027);\nconst n = 32 * 1024;\nconst payload = \u0027\u003ca\u003e\u003c?p\u0027 + \u0027 \u0027.repeat(n); // unterminated PI, no `?\u003e`\nconsole.time(\u0027parse\u0027);\nnew DOMParser().parseFromString(payload, \u0027text/xml\u0027);\nconsole.timeEnd(\u0027parse\u0027);\n```\n\nMeasured (Node 18), trailing whitespace after `\u003c?p`, no `?\u003e` \u2014 time quadruples per doubling of\ninput length (canonical O(n\u00b2)):\n\n| Trailing whitespace | `g.PI.exec` | `parseFromString` |\n|---|---|---|\n| 2 KB | 4.4 ms | 5.1 ms |\n| 4 KB | 16.9 ms | 17.0 ms |\n| 8 KB | 111.4 ms | 66.3 ms |\n| 16 KB | 263.8 ms | 336.5 ms |\n| 32 KB | 1073.1 ms | \u2014 |\n\n## Impact\n\nAvailability only: a single parse of a small crafted document blocks the Node.js event loop for the\nduration of the quadratic scan (\u22481 s at 32 KB; multi-second with larger inputs). No memory blow-up,\nno data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and\nparsed with default options, one request can stall a server.\n\n## Fix Applied\n\nFixed in `@xmldom/xmldom` **0.9.11** (`0.9.x`-only; the `0.8.x` LTS line and the\nunscoped `xmldom` package use a different, bounded PI code path and are not affected).\n\nPR [#1039](https://github.com/xmldom/xmldom/pull/1039) inserts a fixed-width negative lookahead\n`(?!\\s)` immediately after the greedy `S+`, so the separator can no longer hand whitespace back to\nthe lazy data group:\n\n```\n- var PI = reg(/^\u003c\\?/, \u0027(\u0027, Name, \u0027)\u0027, regg(S, \u0027(\u0027, Char, \u0027*?)\u0027), \u0027?\u0027, /\\?\u003e/);\n+ var PI = reg(/^\u003c\\?/, \u0027(\u0027, Name, \u0027)\u0027, regg(S, \u0027(?!\u0027, _SChar, \u0027)(\u0027, Char, \u0027*?)\u0027), \u0027?\u0027, /\\?\u003e/);\n```\n\nThe change is correct, minimal, and behavior-preserving: it produces identical `[target, data]`\ncaptures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl\ncases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead\nis fixed-width and cannot itself backtrack \u2014 a strict improvement with no new parsing risk.\n\n## Severity note\n\nThe complexity is **quadratic**, not exponential, so a multi-second stall requires\ntens-to-hundreds of KB of input. `VA:H` reflects that xmldom applies **no input-size limit** and the\npath runs on default-options parsing, so a single unbounded parse can fully stall the event loop.",
"id": "GHSA-g53g-w8rj-fmg7",
"modified": "2026-09-08T20:31:50Z",
"published": "2026-09-08T20:31:50Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/xmldom/xmldom/security/advisories/GHSA-g53g-w8rj-fmg7"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-83606"
},
{
"type": "WEB",
"url": "https://github.com/xmldom/xmldom/pull/1039"
},
{
"type": "WEB",
"url": "https://github.com/xmldom/xmldom/commit/73df6b8bdbd86f904b9e8c3ab9c49aa54ef2802e"
},
{
"type": "PACKAGE",
"url": "https://github.com/xmldom/xmldom"
},
{
"type": "WEB",
"url": "https://github.com/xmldom/xmldom/releases/tag/0.9.11"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N",
"type": "CVSS_V4"
}
],
"summary": "xmldom PI grammar regex ReDoS: quadratic backtracking on unterminated processing instructions"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.