GHSA-CXF4-7MRP-VVPR
Vulnerability from github – Published: 2026-09-29 18:10 – Updated: 2026-09-29 18:10Related public issue (context, not a duplicate)
Closed issue #1429 ("Handle non-UTF-8 paths", 2024-03-24) raised exactly this general concern and
even suggested detection via preg_match('//u', $path) !== 1 -- note the reporter's suggested check
explicitly compares !== 1, which would correctly treat PCRE's false return as "reject." The
maintainer's reply pointed to the PathNormalizer interface as the place to implement this. The
control-character check that ended up shipping in WhitespacePathNormalizer
(if (preg_match('#\p{C}+#u', $unixPath))) addresses the general concern but does not use the
!== 1-style comparison the original issue suggested -- it uses a bare truthy check, which is exactly
the gap this report demonstrates. So this is not a duplicate of #1429; it's a concrete bypass
surviving in the fix that issue's concern led to.
Vulnerability Details
File: src/WhitespacePathNormalizer.php, lines 22-28 (normalizePath()) -- the default
PathNormalizer used by Filesystem for every adapter (Local, FTP, SFTP, S3, AsyncAwsS3, Azure, GCS,
ZipArchive, GridFS, InMemory) unless the application supplies a custom one.
Root Cause
public function normalizePath(string $path): string
{
$unixPath = str_replace('\\', '/', $path);
if (preg_match('#\p{C}+#u', $unixPath)) {
throw CorruptedPathDetected::forPath($path);
}
...
preg_match() returns false (a PHP engine error) rather than 0 when the subject string is not
valid UTF-8 and the pattern uses the /u modifier -- PCRE can't even attempt the match. false and
0 are both falsy in PHP, and if (preg_match(...)) does not distinguish them. So a path containing
any single invalid UTF-8 byte anywhere in the string makes preg_match() fail with a "Malformed
UTF-8 characters" engine error, the if evaluates false, and CorruptedPathDetected is silently not
thrown -- even when the same string also contains literal control characters this exact check exists
to catch.
the identical payload IS correctly rejected once it's valid UTF-8:
$n->normalizePath("foo\x1bbar"); // valid UTF-8, contains ESC -> throws CorruptedPathDetected (correct)
$n->normalizePath("foo\x80\x1bbar"); // 0x80 = invalid lone UTF-8 continuation byte -> NOT thrown (bypass)
The path-traversal protection is unaffected -- it's exact byte-string comparison on
/-delimited segments, independent of UTF-8 validity:
$n->normalizePath("\x80/../../etc/passwd"); // still throws PathTraversalDetected
Recommended Fix
public function normalizePath(string $path): string
{
$unixPath = str_replace('\\', '/', $path);
$matched = preg_match('#\p{C}+#u', $unixPath);
if ($matched !== 0) {
// $matched === false means malformed UTF-8 -- must also be treated as corrupted,
// not silently allowed through.
throw CorruptedPathDetected::forPath($path);
}
...
Verification
Dynamically confirmed on league/flysystem HEAD 6837e1d / tag 3.35.2, PHP 8.4.22 CLI, end-to-end
through LocalFilesystemAdapter:
``
[1] write() succeeded -- normalizer did NOT reject the path.
[2] Actual bytes on disk: ...801b5b386d6e6f726d616c2d6c6f6f6b696e672d66696c652e7478741b5b306d...
[3] Filesystem::listContents() path contains raw ESC (0x1b): YES
[4] Raw terminal output (viacat -v`): M-^@^[[8mnormal-looking-file.txt^[[0m^[[2K^[[1Aurgent-invoice.pdf
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 3.35.2"
},
"package": {
"ecosystem": "Packagist",
"name": "league/flysystem"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "3.35.3"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-102601"
],
"database_specific": {
"cwe_ids": [
"CWE-150"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-29T18:10:14Z",
"nvd_published_at": null,
"severity": "LOW"
},
"details": "## Related public issue (context, not a duplicate)\n\nClosed issue #1429 (\"Handle non-UTF-8 paths\", 2024-03-24) raised exactly this general concern and\neven suggested detection via `preg_match(\u0027//u\u0027, $path) !== 1` -- note the reporter\u0027s suggested check\nexplicitly compares `!== 1`, which *would* correctly treat PCRE\u0027s `false` return as \"reject.\" The\nmaintainer\u0027s reply pointed to the `PathNormalizer` interface as the place to implement this. The\ncontrol-character check that ended up shipping in `WhitespacePathNormalizer`\n(`if (preg_match(\u0027#\\p{C}+#u\u0027, $unixPath))`) addresses the general concern but does **not** use the\n`!== 1`-style comparison the original issue suggested -- it uses a bare truthy check, which is exactly\nthe gap this report demonstrates. So this is not a duplicate of #1429; it\u0027s a concrete bypass\nsurviving in the fix that issue\u0027s concern led to.\n\n## Vulnerability Details\n\n**File**: `src/WhitespacePathNormalizer.php`, lines 22-28 (`normalizePath()`) -- the **default**\n`PathNormalizer` used by `Filesystem` for every adapter (Local, FTP, SFTP, S3, AsyncAwsS3, Azure, GCS,\nZipArchive, GridFS, InMemory) unless the application supplies a custom one.\n\n### Root Cause\n\n```php\npublic function normalizePath(string $path): string\n{\n $unixPath = str_replace(\u0027\\\\\u0027, \u0027/\u0027, $path);\n\n if (preg_match(\u0027#\\p{C}+#u\u0027, $unixPath)) {\n throw CorruptedPathDetected::forPath($path);\n }\n ...\n```\n\n`preg_match()` returns `false` (a PHP engine error) rather than `0` when the subject string is not\nvalid UTF-8 and the pattern uses the `/u` modifier -- PCRE can\u0027t even attempt the match. `false` and\n`0` are both falsy in PHP, and `if (preg_match(...))` does not distinguish them. So a path containing\n**any single invalid UTF-8 byte anywhere in the string** makes `preg_match()` fail with a \"Malformed\nUTF-8 characters\" engine error, the `if` evaluates false, and `CorruptedPathDetected` is silently **not**\nthrown -- even when the same string also contains literal control characters this exact check exists\nto catch.\n\nthe identical payload IS correctly rejected once it\u0027s valid UTF-8:\n```php\n$n-\u003enormalizePath(\"foo\\x1bbar\"); // valid UTF-8, contains ESC -\u003e throws CorruptedPathDetected (correct)\n$n-\u003enormalizePath(\"foo\\x80\\x1bbar\"); // 0x80 = invalid lone UTF-8 continuation byte -\u003e NOT thrown (bypass)\n```\n\nThe path-traversal protection is **unaffected** -- it\u0027s exact byte-string comparison on\n`/`-delimited segments, independent of UTF-8 validity:\n```php\n$n-\u003enormalizePath(\"\\x80/../../etc/passwd\"); // still throws PathTraversalDetected\n```\n\n### Recommended Fix\n```php\npublic function normalizePath(string $path): string\n{\n $unixPath = str_replace(\u0027\\\\\u0027, \u0027/\u0027, $path);\n $matched = preg_match(\u0027#\\p{C}+#u\u0027, $unixPath);\n\n if ($matched !== 0) {\n // $matched === false means malformed UTF-8 -- must also be treated as corrupted,\n // not silently allowed through.\n throw CorruptedPathDetected::forPath($path);\n }\n ...\n```\n\n### Verification\nDynamically confirmed on `league/flysystem` HEAD `6837e1d` / tag `3.35.2`, PHP 8.4.22 CLI, end-to-end\nthrough `LocalFilesystemAdapter`:\n```\n[1] write() succeeded -- normalizer did NOT reject the path.\n[2] Actual bytes on disk: ...801b5b386d6e6f726d616c2d6c6f6f6b696e672d66696c652e7478741b5b306d...\n[3] Filesystem::listContents() path contains raw ESC (0x1b): YES\n[4] Raw terminal output (via `cat -v`): M-^@^[[8mnormal-looking-file.txt^[[0m^[[2K^[[1Aurgent-invoice.pdf",
"id": "GHSA-cxf4-7mrp-vvpr",
"modified": "2026-09-29T18:10:14Z",
"published": "2026-09-29T18:10:14Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/thephpleague/flysystem/security/advisories/GHSA-cxf4-7mrp-vvpr"
},
{
"type": "WEB",
"url": "https://github.com/thephpleague/flysystem/commit/ef4a9a557d769b5d472c403125716706a0d9cc77"
},
{
"type": "PACKAGE",
"url": "https://github.com/thephpleague/flysystem"
},
{
"type": "WEB",
"url": "https://github.com/thephpleague/flysystem/releases/tag/3.35.3"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:N/I:L/A:N",
"type": "CVSS_V3"
}
],
"summary": "Flysystem: WhitespacePathNormalizer\u0027s control-character (CorruptedPathDetected) check is bypassed by malformed UTF-8 in the path, affecting every adapter"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.