GHSA-G53G-W8RJ-FMG7

Vulnerability from github – Published: 2026-09-08 20:31 – Updated: 2026-09-08 20:31
VLAI
Summary
xmldom PI grammar regex ReDoS: quadratic backtracking on unterminated processing instructions
Details

Summary

@xmldom/xmldom's processing-instruction (PI) grammar regex exhibits quadratic-time backtracking (ReDoS) when parsing an unterminated processing instruction. A single small XML document containing <? + a target + a long run of whitespace and no closing ?> forces the regular expression engine into O(n²) work, stalling the Node.js event loop. The input is parsed with DOMParser.parseFromString under default options, so it is reachable from unauthenticated, network-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs).

Details

The PI production in lib/grammar.js compiles (flags mu) to:

^<\?(NameChars)(?:[\x20\x09\x0D\x0A]+([Char]*?))?\?>
                     ^^^ S+ greedy       ^^^ Char*? lazy
  • lib/grammar.js line 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261

In the optional tail (?:S+(Char*?))?, both the greedy separator S+ and the lazy data Char*? match XML whitespace. When the required trailing ?> is absent, the engine must ultimately fail — but first it tries every partition of the whitespace run between S+ and Char*?, which is O(n²) in the length of the trailing whitespace.

The regex is executed against the entire remaining source string in two places in lib/sax.js, so the whole whitespace tail is scanned:

  • parsePI — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691
  • parseProcessingInstruction — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879

Affected Versions

Only the 0.9.x line is affected. lib/grammar.js (and this PI regex) was introduced in commit 726b471 ("fix!: preserve DOCTYPE internal subset (#498)"), first released in 0.9.0-beta.9, and is unchanged through 0.9.10.

The 0.8.x line (≤ 0.8.13) and the unscoped xmldom package (≤ 0.6.0) parse PIs via a different code path bounded by indexOf('?>') — they do not contain this regex and are not affected by this issue. (They were not separately tested for a different PI ReDoS; the scope here is the specific grammar.js regex.)

Line PI code path Affected?
0.9.x (0.9.0-beta.9 … 0.9.10) grammar.js PI regex over full remaining source Yes
0.8.x (≤ 0.8.13) parseInstruction, bounded by indexOf('?>') No
unscoped xmldom (≤ 0.6.0) older indexOf('?>')-bounded parsing No

Proof of Concept

const { DOMParser } = require('@xmldom/xmldom');
const n = 32 * 1024;
const payload = '<a><?p' + ' '.repeat(n); // unterminated PI, no `?>`
console.time('parse');
new DOMParser().parseFromString(payload, 'text/xml');
console.timeEnd('parse');

Measured (Node 18), trailing whitespace after <?p, no ?> — time quadruples per doubling of input length (canonical O(n²)):

Trailing whitespace g.PI.exec parseFromString
2 KB 4.4 ms 5.1 ms
4 KB 16.9 ms 17.0 ms
8 KB 111.4 ms 66.3 ms
16 KB 263.8 ms 336.5 ms
32 KB 1073.1 ms —

Impact

Availability only: a single parse of a small crafted document blocks the Node.js event loop for the duration of the quadratic scan (≈1 s at 32 KB; multi-second with larger inputs). No memory blow-up, no data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and parsed with default options, one request can stall a server.

Fix Applied

Fixed in @xmldom/xmldom 0.9.11 (0.9.x-only; the 0.8.x LTS line and the unscoped xmldom package use a different, bounded PI code path and are not affected).

PR #1039 inserts a fixed-width negative lookahead (?!\s) immediately after the greedy S+, so the separator can no longer hand whitespace back to the lazy data group:

- var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(', Char, '*?)'), '?', /\?>/);
+ var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(?!', _SChar, ')(', Char, '*?)'), '?', /\?>/);

The change is correct, minimal, and behavior-preserving: it produces identical [target, data] captures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl cases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead is fixed-width and cannot itself backtrack — a strict improvement with no new parsing risk.

Severity note

The complexity is quadratic, not exponential, so a multi-second stall requires tens-to-hundreds of KB of input. VA:H reflects that xmldom applies no input-size limit and the path runs on default-options parsing, so a single unbounded parse can fully stall the event loop.

Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 0.9.10"
      },
      "package": {
        "ecosystem": "npm",
        "name": "@xmldom/xmldom"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0.9.0-beta.9"
            },
            {
              "fixed": "0.9.11"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-83606"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-1333",
      "CWE-400"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-09-08T20:31:50Z",
    "nvd_published_at": "2026-09-01T15:17:38Z",
    "severity": "HIGH"
  },
  "details": "## Summary\n\n`@xmldom/xmldom`\u0027s processing-instruction (PI) grammar regex exhibits quadratic-time backtracking\n(ReDoS) when parsing an **unterminated** processing instruction. A single small XML document\ncontaining `\u003c?` + a target + a long run of whitespace and no closing `?\u003e` forces the regular\nexpression engine into O(n\u00b2) work, stalling the Node.js event loop. The input is parsed with\n`DOMParser.parseFromString` under **default options**, so it is reachable from unauthenticated,\nnetwork-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs).\n\n## Details\n\nThe PI production in `lib/grammar.js` compiles (flags `mu`) to:\n\n```\n^\u003c\\?(NameChars)(?:[\\x20\\x09\\x0D\\x0A]+([Char]*?))?\\?\u003e\n                     ^^^ S+ greedy       ^^^ Char*? lazy\n```\n\n- `lib/grammar.js` line 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261\n\nIn the optional tail `(?:S+(Char*?))?`, both the greedy separator `S+` and the lazy data `Char*?`\nmatch XML whitespace. When the required trailing `?\u003e` is absent, the engine must ultimately fail \u2014\nbut first it tries every partition of the whitespace run between `S+` and `Char*?`, which is O(n\u00b2)\nin the length of the trailing whitespace.\n\nThe regex is executed against the **entire remaining source string** in two places in `lib/sax.js`,\nso the whole whitespace tail is scanned:\n\n- `parsePI` \u2014 https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691\n- `parseProcessingInstruction` \u2014 https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879\n\n## Affected Versions\n\nOnly the `0.9.x` line is affected. `lib/grammar.js` (and this PI regex) was introduced in\ncommit `726b471` (\"fix!: preserve DOCTYPE internal subset (#498)\"), first released in\n**0.9.0-beta.9**, and is unchanged through **0.9.10**.\n\nThe `0.8.x` line (\u2264 0.8.13) and the unscoped `xmldom` package (\u2264 0.6.0) parse PIs via a different\ncode path bounded by `indexOf(\u0027?\u003e\u0027)` \u2014 they do **not** contain this regex and are **not affected**\nby this issue. (They were not separately tested for a *different* PI ReDoS; the scope here is the\nspecific `grammar.js` regex.)\n\n| Line | PI code path | Affected? |\n|---|---|---|\n| `0.9.x` (0.9.0-beta.9 \u2026 0.9.10) | `grammar.js` `PI` regex over full remaining source | **Yes** |\n| `0.8.x` (\u2264 0.8.13) | `parseInstruction`, bounded by `indexOf(\u0027?\u003e\u0027)` | No |\n| unscoped `xmldom` (\u2264 0.6.0) | older `indexOf(\u0027?\u003e\u0027)`-bounded parsing | No |\n\n## Proof of Concept\n\n```js\nconst { DOMParser } = require(\u0027@xmldom/xmldom\u0027);\nconst n = 32 * 1024;\nconst payload = \u0027\u003ca\u003e\u003c?p\u0027 + \u0027 \u0027.repeat(n); // unterminated PI, no `?\u003e`\nconsole.time(\u0027parse\u0027);\nnew DOMParser().parseFromString(payload, \u0027text/xml\u0027);\nconsole.timeEnd(\u0027parse\u0027);\n```\n\nMeasured (Node 18), trailing whitespace after `\u003c?p`, no `?\u003e` \u2014 time quadruples per doubling of\ninput length (canonical O(n\u00b2)):\n\n| Trailing whitespace | `g.PI.exec` | `parseFromString` |\n|---|---|---|\n| 2 KB  | 4.4 ms    | 5.1 ms   |\n| 4 KB  | 16.9 ms   | 17.0 ms  |\n| 8 KB  | 111.4 ms  | 66.3 ms  |\n| 16 KB | 263.8 ms  | 336.5 ms |\n| 32 KB | 1073.1 ms | \u2014        |\n\n## Impact\n\nAvailability only: a single parse of a small crafted document blocks the Node.js event loop for the\nduration of the quadratic scan (\u22481 s at 32 KB; multi-second with larger inputs). No memory blow-up,\nno data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and\nparsed with default options, one request can stall a server.\n\n## Fix Applied\n\nFixed in `@xmldom/xmldom` **0.9.11** (`0.9.x`-only; the `0.8.x` LTS line and the\nunscoped `xmldom` package use a different, bounded PI code path and are not affected).\n\nPR [#1039](https://github.com/xmldom/xmldom/pull/1039) inserts a fixed-width negative lookahead\n`(?!\\s)` immediately after the greedy `S+`, so the separator can no longer hand whitespace back to\nthe lazy data group:\n\n```\n- var PI = reg(/^\u003c\\?/, \u0027(\u0027, Name, \u0027)\u0027, regg(S, \u0027(\u0027, Char, \u0027*?)\u0027), \u0027?\u0027, /\\?\u003e/);\n+ var PI = reg(/^\u003c\\?/, \u0027(\u0027, Name, \u0027)\u0027, regg(S, \u0027(?!\u0027, _SChar, \u0027)(\u0027, Char, \u0027*?)\u0027), \u0027?\u0027, /\\?\u003e/);\n```\n\nThe change is correct, minimal, and behavior-preserving: it produces identical `[target, data]`\ncaptures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl\ncases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead\nis fixed-width and cannot itself backtrack \u2014 a strict improvement with no new parsing risk.\n\n## Severity note\n\nThe complexity is **quadratic**, not exponential, so a multi-second stall requires\ntens-to-hundreds of KB of input. `VA:H` reflects that xmldom applies **no input-size limit** and the\npath runs on default-options parsing, so a single unbounded parse can fully stall the event loop.",
  "id": "GHSA-g53g-w8rj-fmg7",
  "modified": "2026-09-08T20:31:50Z",
  "published": "2026-09-08T20:31:50Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/xmldom/xmldom/security/advisories/GHSA-g53g-w8rj-fmg7"
    },
    {
      "type": "ADVISORY",
      "url": "https://nvd.nist.gov/vuln/detail/CVE-2026-83606"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xmldom/xmldom/pull/1039"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xmldom/xmldom/commit/73df6b8bdbd86f904b9e8c3ab9c49aa54ef2802e"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/xmldom/xmldom"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xmldom/xmldom/releases/tag/0.9.11"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N",
      "type": "CVSS_V4"
    }
  ],
  "summary": "xmldom PI grammar regex ReDoS: quadratic backtracking on unterminated processing instructions"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…