<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://db.gcve.eu/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-09-28T08:13:32.401874+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>info@gcve.eu</email>
  </author>
  <link href="https://db.gcve.eu" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://db.gcve.eu/vuln/brew-acronym-cve-2026-81725</id>
    <title>BREW-acronym-CVE-2026-81725 — NLTK: Pl196xCorpusReader has quadratic ReDoS on malformed TEI blocks</title>
    <updated>2026-09-28T08:13:32.405432+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> Homebrew: acronym</p>
<p>### Summary</p>
<p>`Pl196xCorpusReader` still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.</p>
<p>### Details</p>
<p>- **Vulnerability type:** Regular-expression denial of service
- **Affected component:** `nltk.corpus.reader.pl196x.TEICorpusView.read_block` and `Pl196xCorpusReader` public methods
- **Affected versions:** Published `3.9.4` and current source `v3.10.0-rc2` both reproduced.
- **Patched versions:** Not yet patched
- **Root cause:** Lazy `.*?` whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.</p>
<p>The parser uses regexes for paragraphs, sentences, and word tags across the whole `&lt;text&gt;` block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed `&lt;p&gt;` tags doubled, through normal public calls such as `words()` and `tagged_words()`.</p>
<p>### PoC</p>
<p>**Preconditions**
- The application parses attacker-influenced PL196X or TEI-like corpus files through public reader APIs.</p>
<p>**Steps**
1. Create a corpus file with a valid header followed by a `&lt;text&gt;` block that contains many opening tags and no matching closing tags.
2. Instantiate `Pl196xCorpusReader` on that corpus.
3. Call `words()` or `tagged_words()`…</p></div>
    </content>
    <link href="https://db.gcve.eu/vuln/brew-acronym-cve-2026-81725"/>
  </entry>
  <entry>
    <id>https://db.gcve.eu/vuln/cve-2026-81725</id>
    <title>CVE-2026-81725 — NLTK before 3.10.3 Regular Expression Denial of Service via Pl196xCorpusReader</title>
    <updated>2026-09-28T08:13:32.405548+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> nltk</p>
<p>NLTK before 3.10.3 contains a regular expression denial of service vulnerability in Pl196xCorpusReader that allows attackers to cause quadratic CPU consumption by supplying malformed TEI blocks with many unmatched opening tags. Attackers can exploit lazy regex patterns in the read_block method through public APIs like words() and tagged_words() to force repeated rescans and achieve near-quadratic runtime growth.</p></div>
    </content>
    <link href="https://db.gcve.eu/vuln/cve-2026-81725"/>
  </entry>
  <entry>
    <id>https://db.gcve.eu/vuln/pysec-2026-3752</id>
    <title>PYSEC-2026-3752</title>
    <updated>2026-09-28T08:13:32.405600+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: nltk</p>
<p>NLTK before 3.10.3 contains a regular expression denial of service vulnerability in Pl196xCorpusReader that allows attackers to cause quadratic CPU consumption by supplying malformed TEI blocks with many unmatched opening tags. Attackers can exploit lazy regex patterns in the read_block method through public APIs like words() and tagged_words() to force repeated rescans and achieve near-quadratic runtime growth.</p></div>
    </content>
    <link href="https://db.gcve.eu/vuln/pysec-2026-3752"/>
  </entry>
</feed>
