Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.
nltk.corpus.reader.pl196x.TEICorpusView.read_block and Pl196xCorpusReader public methods3.9.4 and current source v3.10.0-rc2 both reproduced..*? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and tagged_words().
Preconditions
Steps
<text> block that contains many opening tags and no matching closing tags.Pl196xCorpusReader on that corpus.words() or tagged_words() and measure elapsed time as the malformed tag count doubles.Minimal reproducible excerpt
size=1000 0.014s
size=2000 0.057s
size=4000 0.231s
size=8000 0.927s
A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.
Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.
{
"cwe_ids": [
"CWE-1333",
"CWE-400"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-08T20:28:46Z",
"nvd_published_at": null,
"severity": "MODERATE"
}