GHSA-29PJ-957V-52MC
Vulnerability from github – Published: 2026-08-06 20:30 – Updated: 2026-08-06 20:30## Summary
The AttributesExtension's href/src unsafe-link filter (AttributesHelper::filterAttributes()) can be bypassed by embedding control bytes in a javascript: URL that browsers discard before parsing the scheme. Two variants:
- Tab/newline inside the scheme — a literal ASCII TAB (0x09), CR (0x0D), or LF (0x0A), e.g.
java<TAB>script:alert(1). Per the WHATWG URL Standard's "basic URL parser" step 3, browsers "remove all ASCII tab or newline from input". - Leading C0 controls — e.g.
<0x01>javascript:alert(1). Per step 1 of the same algorithm, browsers remove any leading or trailing C0 control or space. (A leading space alone does not bypass, becauseparseAttributes()alreadytrim()s the value; other C0 bytes are not trimmed.)
The filter is a literal anchored-prefix regex (RegexHelper::isLinkPotentiallyUnsafe() / REGEX_UNSAFE_PROTOCOL) that matches neither obfuscated form, so in both cases the browser still executes javascript:alert(1).
This is confirmed reproducible even with allow_unsafe_links => false set — i.e. even applications that have followed the library's own documented hardening guidance for untrusted input remain exploitable.
This is a sibling gap in the same defense that CVE-2025-46734 (GHSA-3527-qv2q-pfvx) fixed in v2.7.0 — that fix made href/src respect allow_unsafe_links, but did not normalize control bytes before checking, so these obfuscation techniques were never covered.
Vulnerability
Files:
- src/Util/RegexHelper.php:69 (REGEX_UNSAFE_PROTOCOL), :239-242 (isLinkPotentiallyUnsafe())
- src/Extension/Attributes/Util/AttributesHelper.php:149-179 (filterAttributes())
CWE: CWE-79 (Improper Neutralization of Input During Web Page Generation / XSS) — primary
- CWE-692 (Incomplete Denylist to Cross-Site Scripting) — the anchored-prefix denylist in REGEX_UNSAFE_PROTOCOL is incomplete. This is a composite of CWE-184 and CWE-79, so it captures the full "incomplete denylist → XSS" chain on its own.
- CWE-86 (Improper Neutralization of Invalid Characters in Identifiers in Web Pages) — the specific evasion technique: control bytes embedded within the URI scheme identifier, which the browser strips before resolving it.
Root Cause
// src/Util/RegexHelper.php
public const REGEX_UNSAFE_PROTOCOL = '/^(?:javascript|vbscript|file|data):/i';
public static function isLinkPotentiallyUnsafe(string $url): bool
{
return \preg_match(self::REGEX_UNSAFE_PROTOCOL, $url) !== 0 && \preg_match(self::REGEX_SAFE_DATA_PROTOCOL, $url) === 0;
}
// src/Extension/Attributes/Util/AttributesHelper.php
foreach ($attributes as $name => $value) {
$attrNameLower = \strtolower($name);
if (! $allowUnsafeLinks && ($attrNameLower === 'href' || $attrNameLower === 'src') && \is_string($value) && RegexHelper::isLinkPotentiallyUnsafe($value)) {
unset($attributes[$name]);
continue;
}
...
The Attributes extension's own quote-value grammar (PARTIAL_DOUBLEQUOTEDVALUE = '"[^"]*"') accepts any byte except " inside quotes, including raw tab/CR/LF and other C0 controls, and parseAttributes() only trim()s (leading/trailing, and only the default charlist " \t\n\r\0\x0B" — so a leading \x01 survives). Critically, the core Markdown link-destination path (LinkParserHelper → UrlEncoder::unescapeAndEncode()) percent-encodes every control byte before this same safety check ever runs — but the Attributes extension's href/src handling has no equivalent normalization step, so the raw control byte reaches both the check and the final HTML output (Xml::escape() only escapes & < > " ', not tab/CR/LF, since they're legal bytes inside an HTML attribute).
Attack Scenario
- An application enables the (commonly-used)
AttributesExtensionand setsallow_unsafe_links => false— the project's own documented hardening step for untrusted input. - An attacker submits Markdown:
[Click me](javascript:alert(0)){href="java<TAB>script:alert(document.cookie)"}(TAB is one literal 0x09 byte). - The library emits
<a href="java<TAB>script:alert(document.cookie)">Click me</a>—isLinkPotentiallyUnsafe()doesn't match the tab-split scheme, so the filter takes no action. - A victim viewing/clicking the link has the browser strip the embedded TAB and execute
javascript:alert(document.cookie)in the victim's session — stored XSS, cookie theft, account takeover potential.
Why the payload needs an unsafe core destination. Step 2 above deliberately uses [Click me](javascript:alert(0)) rather than a normal link. LinkRenderer overwrites attrs['href'] with the node's own URL unless that URL is itself judged unsafe — so [x](https://example.com){href="java<TAB>script:..."} renders the harmless href="https://example.com", and an empty destination [x](){href="..."} renders href="". The attacker therefore supplies a core destination that the filter does catch, which suppresses the overwrite and lets the attribute-supplied href reach the final tag. This is no obstacle in practice — the attacker writes the entire Markdown document.
Two related forms that are not exploitable, noted so the fix isn't over-scoped:
- Attaching the attribute to a non-link block —
hi {href="java<TAB>script:alert(1)"}— does bypass the filter and emits<p href="java<TAB>script:alert(1)">, buthrefon a<p>is inert: there is nothing to navigate. (An earlier draft of this report described this as a "simpler, unconditional variant" of the attack; it is a filter bypass, not an XSS.) <img src>is unaffected, sinceImageRendererunconditionally overwritessrcfrom the core URL regardless of the safety verdict.
Recommended Fix
Normalize inside RegexHelper::isLinkPotentiallyUnsafe() before testing, mirroring the WHATWG URL parser's own normalization. This covers both variants, fixes every call site at once (LinkRenderer, ImageRenderer, and any third-party callers), and needs no changes in the Attributes extension.
Affected Versions
>= 1.5.0, <= 2.8.3 - every release that ships the AttributesExtension. Verified by installing each version and rendering the payloads with allow_unsafe_links => false. The attribute-value grammar (PARTIAL_DOUBLEQUOTEDVALUE = '"[^"]*"') has accepted raw control bytes since the extension was introduced, and none of the intervening parser rewrites narrowed it.
Prior Related Advisories
GHSA-3527-qv2q-pfvx / CVE-2025-46734 fixed a different Attributes-extension XSS (unallowlisted on* handlers, href/src not respecting allow_unsafe_links at all) in v2.7.0. This issue bypasses the specific href/src protection that fix introduced (the control-byte normalization gap was not part of that fix) - but the obfuscated inputs also work on older versions.
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 2.8.3"
},
"package": {
"ecosystem": "Packagist",
"name": "league/commonmark"
},
"ranges": [
{
"events": [
{
"introduced": "1.5.0"
},
{
"fixed": "2.9.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-71478"
],
"database_specific": {
"cwe_ids": [
"CWE-79",
"CWE-86",
"CWE-692"
],
"github_reviewed": true,
"github_reviewed_at": "2026-08-06T20:30:39Z",
"nvd_published_at": null,
"severity": "MODERATE"
},
"details": "\ufeff## Summary\n\nThe `AttributesExtension`\u0027s `href`/`src` unsafe-link filter (`AttributesHelper::filterAttributes()`) can be bypassed by embedding control bytes in a `javascript:` URL that browsers discard before parsing the scheme. Two variants:\n\n- **Tab/newline inside the scheme** \u2014 a literal ASCII TAB (0x09), CR (0x0D), or LF (0x0A), e.g. `java\u003cTAB\u003escript:alert(1)`. Per the WHATWG URL Standard\u0027s \"basic URL parser\" step 3, browsers \"remove all ASCII tab or newline from input\".\n- **Leading C0 controls** \u2014 e.g. `\u003c0x01\u003ejavascript:alert(1)`. Per step 1 of the same algorithm, browsers remove any leading or trailing C0 control or space. (A leading *space* alone does not bypass, because `parseAttributes()` already `trim()`s the value; other C0 bytes are not trimmed.)\n\nThe filter is a literal anchored-prefix regex (`RegexHelper::isLinkPotentiallyUnsafe()` / `REGEX_UNSAFE_PROTOCOL`) that matches neither obfuscated form, so in both cases the browser still executes `javascript:alert(1)`.\n\n**This is confirmed reproducible even with `allow_unsafe_links =\u003e false` set** \u2014 i.e. even applications that have followed the library\u0027s own documented hardening guidance for untrusted input remain exploitable.\n\nThis is a *sibling gap* in the same defense that CVE-2025-46734 (GHSA-3527-qv2q-pfvx) fixed in v2.7.0 \u2014 that fix made `href`/`src` respect `allow_unsafe_links`, but did not normalize control bytes before checking, so these obfuscation techniques were never covered.\n\n## Vulnerability\n\n**Files**:\n- `src/Util/RegexHelper.php:69` (`REGEX_UNSAFE_PROTOCOL`), `:239-242` (`isLinkPotentiallyUnsafe()`)\n- `src/Extension/Attributes/Util/AttributesHelper.php:149-179` (`filterAttributes()`)\n\n**CWE**: CWE-79 (Improper Neutralization of Input During Web Page Generation / XSS) \u2014 primary\n- CWE-692 (Incomplete Denylist to Cross-Site Scripting) \u2014 the anchored-prefix denylist in `REGEX_UNSAFE_PROTOCOL` is incomplete. This is a composite of CWE-184 and CWE-79, so it captures the full \"incomplete denylist \u2192 XSS\" chain on its own.\n- CWE-86 (Improper Neutralization of Invalid Characters in Identifiers in Web Pages) \u2014 the specific evasion technique: control bytes embedded within the URI scheme identifier, which the browser strips before resolving it.\n\n### Root Cause\n```php\n// src/Util/RegexHelper.php\npublic const REGEX_UNSAFE_PROTOCOL = \u0027/^(?:javascript|vbscript|file|data):/i\u0027;\n\npublic static function isLinkPotentiallyUnsafe(string $url): bool\n{\n return \\preg_match(self::REGEX_UNSAFE_PROTOCOL, $url) !== 0 \u0026\u0026 \\preg_match(self::REGEX_SAFE_DATA_PROTOCOL, $url) === 0;\n}\n\n// src/Extension/Attributes/Util/AttributesHelper.php\nforeach ($attributes as $name =\u003e $value) {\n $attrNameLower = \\strtolower($name);\n if (! $allowUnsafeLinks \u0026\u0026 ($attrNameLower === \u0027href\u0027 || $attrNameLower === \u0027src\u0027) \u0026\u0026 \\is_string($value) \u0026\u0026 RegexHelper::isLinkPotentiallyUnsafe($value)) {\n unset($attributes[$name]);\n continue;\n }\n ...\n```\nThe Attributes extension\u0027s own quote-value grammar (`PARTIAL_DOUBLEQUOTEDVALUE = \u0027\"[^\"]*\"\u0027`) accepts any byte except `\"` inside quotes, including raw tab/CR/LF and other C0 controls, and `parseAttributes()` only `trim()`s (leading/trailing, and only the default charlist `\" \\t\\n\\r\\0\\x0B\"` \u2014 so a leading `\\x01` survives). Critically, **the core Markdown link-destination path (`LinkParserHelper` \u2192 `UrlEncoder::unescapeAndEncode()`) percent-encodes every control byte before this same safety check ever runs \u2014 but the Attributes extension\u0027s `href`/`src` handling has no equivalent normalization step**, so the raw control byte reaches both the check and the final HTML output (`Xml::escape()` only escapes `\u0026 \u003c \u003e \" \u0027`, not tab/CR/LF, since they\u0027re legal bytes inside an HTML attribute).\n\n### Attack Scenario\n1. An application enables the (commonly-used) `AttributesExtension` and sets `allow_unsafe_links =\u003e false` \u2014 the project\u0027s own documented hardening step for untrusted input.\n2. An attacker submits Markdown: `[Click me](javascript:alert(0)){href=\"java\u003cTAB\u003escript:alert(document.cookie)\"}` (TAB is one literal 0x09 byte).\n3. The library emits `\u003ca href=\"java\u003cTAB\u003escript:alert(document.cookie)\"\u003eClick me\u003c/a\u003e` \u2014 `isLinkPotentiallyUnsafe()` doesn\u0027t match the tab-split scheme, so the filter takes no action.\n4. A victim viewing/clicking the link has the browser strip the embedded TAB and execute `javascript:alert(document.cookie)` in the victim\u0027s session \u2014 stored XSS, cookie theft, account takeover potential.\n\n**Why the payload needs an unsafe core destination.** Step 2 above deliberately uses `[Click me](javascript:alert(0))` rather than a normal link. `LinkRenderer` overwrites `attrs[\u0027href\u0027]` with the node\u0027s own URL *unless* that URL is itself judged unsafe \u2014 so `[x](https://example.com){href=\"java\u003cTAB\u003escript:...\"}` renders the harmless `href=\"https://example.com\"`, and an empty destination `[x](){href=\"...\"}` renders `href=\"\"`. The attacker therefore supplies a core destination that the filter *does* catch, which suppresses the overwrite and lets the attribute-supplied `href` reach the final tag. This is no obstacle in practice \u2014 the attacker writes the entire Markdown document.\n\nTwo related forms that are **not** exploitable, noted so the fix isn\u0027t over-scoped:\n\n- Attaching the attribute to a non-link block \u2014 `hi {href=\"java\u003cTAB\u003escript:alert(1)\"}` \u2014 does bypass the filter and emits `\u003cp href=\"java\u003cTAB\u003escript:alert(1)\"\u003e`, but `href` on a `\u003cp\u003e` is inert: there is nothing to navigate. (An earlier draft of this report described this as a \"simpler, unconditional variant\" of the attack; it is a filter bypass, not an XSS.)\n- `\u003cimg src\u003e` is unaffected, since `ImageRenderer` unconditionally overwrites `src` from the core URL regardless of the safety verdict.\n\n### Recommended Fix\n\nNormalize inside `RegexHelper::isLinkPotentiallyUnsafe()` before testing, mirroring the WHATWG URL parser\u0027s own normalization. This covers both variants, fixes every call site at once (`LinkRenderer`, `ImageRenderer`, and any third-party callers), and needs no changes in the Attributes extension.\n\n## Affected Versions\n\n**`\u003e= 1.5.0, \u003c= 2.8.3`** - every release that ships the `AttributesExtension`. Verified by installing each version and rendering the payloads with `allow_unsafe_links =\u003e false`. The attribute-value grammar (`PARTIAL_DOUBLEQUOTEDVALUE = \u0027\"[^\"]*\"\u0027`) has accepted raw control bytes since the extension was introduced, and none of the intervening parser rewrites narrowed it.\n\n## Prior Related Advisories\n\nGHSA-3527-qv2q-pfvx / CVE-2025-46734 fixed a different Attributes-extension XSS (unallowlisted `on*` handlers, `href`/`src` not respecting `allow_unsafe_links` at all) in v2.7.0. This issue bypasses the specific `href`/`src` protection that fix introduced (the control-byte normalization gap was not part of that fix) - but the obfuscated inputs also work on older versions.",
"id": "GHSA-29pj-957v-52mc",
"modified": "2026-08-06T20:30:40Z",
"published": "2026-08-06T20:30:39Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/thephpleague/commonmark/security/advisories/GHSA-29pj-957v-52mc"
},
{
"type": "WEB",
"url": "https://github.com/thephpleague/commonmark/commit/493a5aa7d65754b73846006eaff9c2c4431a8e2c"
},
{
"type": "PACKAGE",
"url": "https://github.com/thephpleague/commonmark"
},
{
"type": "WEB",
"url": "https://github.com/thephpleague/commonmark/releases/tag/2.9.0"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N",
"type": "CVSS_V3"
}
],
"summary": "league/commonmark: AttributesExtension href/src unsafe-link filter bypass via embedded control bytes"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.