GHSA-8MGP-746C-J5XP
Vulnerability from github – Published: 2026-09-02 14:35 – Updated: 2026-09-02 14:35Summary
Several model-artifact APIs still treat caller-controlled model paths as ordinary filenames even when NLTK path security is enforced. The same outside-root paths are rejected by guarded helpers, but these public read and write flows still use raw file APIs.
Details
- Vulnerability type: File sandbox bypass
- Affected component:
TransitionParser.train,TransitionParser.parse,AveragedPerceptron.save,AveragedPerceptron.load,PerceptronTagger.save_to_json,save_maxent_params - Affected versions: Published
3.9.4and current sourcev3.10.0-rc2both reproduced. - Patched versions: Not yet patched
- Root cause: Model import and export helpers use built-in
open()on caller-controlled paths instead of pathsec-aware helpers.
TransitionParser.train() writes outside allowed roots, TransitionParser.parse() reads outside allowed roots, AveragedPerceptron bypasses the sandbox in both directions, and adjacent read-side helpers in the same family already show the intended guarded behavior. I confirmed outside-root reads and writes while pathsec.open() or the guarded sibling helpers rejected the same paths.
PoC
Preconditions
- The application enables pathsec enforcement and lets untrusted workflows choose model import or export paths.
Steps
1. Enable pathsec.ENFORCE=True and restrict allowed roots to a dedicated sandbox directory.
2. Use public model import or export APIs with paths that point outside that root.
3. Observe the same paths are rejected by negative-control guarded helpers such as pathsec.open(), PerceptronTagger.load_from_json(), or load_maxent_params().
4. Observe the vulnerable APIs still read or write outside-root files successfully.
Minimal reproducible excerpt
transition_train_exists True
transition_parse_loader_read_bytes 13
averaged_load_keys ['bias']
maxent_save wrote ['alwayson.tab', 'labels.txt']
Impact
Consumers that rely on pathsec for local containment can be tricked into reading or overwriting files outside approved roots through normal model persistence and loading APIs.
Remediation
Route all model-path file access through nltk.pathsec.open() or existing pathsec-aware helpers, and add regression tests that pair each vulnerable API with a negative control on the same path.
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "nltk"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"last_affected": "3.10.3"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-81726"
],
"database_specific": {
"cwe_ids": [
"CWE-22",
"CWE-59",
"CWE-73"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-02T14:35:04Z",
"nvd_published_at": null,
"severity": "HIGH"
},
"details": "### Summary\n\nSeveral model-artifact APIs still treat caller-controlled model paths as ordinary filenames even when NLTK path security is enforced. The same outside-root paths are rejected by guarded helpers, but these public read and write flows still use raw file APIs.\n\n### Details\n\n- **Vulnerability type:** File sandbox bypass\n- **Affected component:** `TransitionParser.train`, `TransitionParser.parse`, `AveragedPerceptron.save`, `AveragedPerceptron.load`, `PerceptronTagger.save_to_json`, `save_maxent_params`\n- **Affected versions:** Published `3.9.4` and current source `v3.10.0-rc2` both reproduced.\n- **Patched versions:** Not yet patched\n- **Root cause:** Model import and export helpers use built-in `open()` on caller-controlled paths instead of pathsec-aware helpers.\n\n`TransitionParser.train()` writes outside allowed roots, `TransitionParser.parse()` reads outside allowed roots, `AveragedPerceptron` bypasses the sandbox in both directions, and adjacent read-side helpers in the same family already show the intended guarded behavior. I confirmed outside-root reads and writes while `pathsec.open()` or the guarded sibling helpers rejected the same paths.\n\n### PoC\n\n**Preconditions**\n- The application enables `pathsec` enforcement and lets untrusted workflows choose model import or export paths.\n\n**Steps**\n1. Enable `pathsec.ENFORCE=True` and restrict allowed roots to a dedicated sandbox directory.\n2. Use public model import or export APIs with paths that point outside that root.\n3. Observe the same paths are rejected by negative-control guarded helpers such as `pathsec.open()`, `PerceptronTagger.load_from_json()`, or `load_maxent_params()`.\n4. Observe the vulnerable APIs still read or write outside-root files successfully.\n\n**Minimal reproducible excerpt**\n\n```text\ntransition_train_exists True\ntransition_parse_loader_read_bytes 13\naveraged_load_keys [\u0027bias\u0027]\nmaxent_save wrote [\u0027alwayson.tab\u0027, \u0027labels.txt\u0027]\n```\n\n### Impact\n\nConsumers that rely on `pathsec` for local containment can be tricked into reading or overwriting files outside approved roots through normal model persistence and loading APIs.\n\n### Remediation\n\nRoute all model-path file access through `nltk.pathsec.open()` or existing pathsec-aware helpers, and add regression tests that pair each vulnerable API with a negative control on the same path.",
"id": "GHSA-8mgp-746c-j5xp",
"modified": "2026-09-02T14:35:04Z",
"published": "2026-09-02T14:35:04Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/security/advisories/GHSA-8mgp-746c-j5xp"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-81726"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/pull/3757"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/pull/3759"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/pull/3813"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/commit/2a92b71827d754ae8920261e7ed0c4bb283ab2d7"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/commit/a44a7af69bca87e92d9c4a701fcbbe4512e8d450"
},
{
"type": "WEB",
"url": "https://github.com/nltk/nltk/commit/cbc98458b43de5f792f0382583c16df39e5c5117"
},
{
"type": "PACKAGE",
"url": "https://github.com/nltk/nltk"
},
{
"type": "WEB",
"url": "https://github.com/pypa/advisory-database/tree/main/vulns/nltk/PYSEC-2026-3740.yaml"
},
{
"type": "WEB",
"url": "https://www.vulncheck.com/advisories/nltk-through-3.10.3-path-traversal-via-model-artifact-apis"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:L/A:L",
"type": "CVSS_V3"
},
{
"score": "CVSS:4.0/AV:N/AC:H/AT:N/PR:N/UI:N/VC:H/VI:L/VA:L/SC:N/SI:N/SA:N",
"type": "CVSS_V4"
}
],
"summary": "NLTK: Model-artifact APIs bypass pathsec and touch files outside allowed roots"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.