{"id":"GHSA-6mj3-qw4j-hgrw","summary":"xmldom: HTML raw-text closing-tag case mismatch causes output amplification","details":"## Summary\n\nIn HTML mode (`text/html`), a raw-text element (`script`, `style`, `textarea`, `title`) whose closing\ntag differs in case from its opening tag (e.g. `\u003c/ScRiPt\u003e` for `\u003cscript\u003e`) is mishandled by the\nparser, producing quadratic (O(n²)) output growth — a small crafted document parses and serializes\ninto output orders of magnitude larger, exhausting CPU and memory. A modest input of tens of KB can\ntherefore cause a denial of service in any service that parses untrusted HTML with xmldom. Only HTML\nmode is affected.\n\n## Details\n\nThe parser calls `parseHtmlSpecialContent` for each raw-text element in HTML mode, matched via\n`isHTMLRawTextElement` / `isHTMLEscapableRawTextElement` (so all four types — `script`, `style`,\n`textarea`, `title` — are in scope). It searches for the element's closing tag with\n`source.indexOf('\u003c/' + tagName + '\u003e', elStartEnd)`, a byte-for-byte case-sensitive match. A\nmixed-case closing tag never matches, so the search returns `-1`, and the following\n`source.substring(elStartEnd + 1, -1)` extracts text backwards from the start of the document\ninstead of the element's content. The function then returns `-1` to the parse loop, which cannot\nadvance normally and falls back to character-by-character reprocessing. Every raw-text element\nre-captures all source text preceding it, so output grows as O(n²) in the number of such elements.\n\n### Root Cause\n\n1. **Case-sensitive close-tag search** (`lib/sax.js:549`): `source.indexOf('\u003c/' + tagName + '\u003e',\n   elStartEnd)` does not fold case, contrary to the WHATWG HTML RAWTEXT end-tag-name rule.\n2. **Unguarded `-1`** (`lib/sax.js:550`): `source.substring(elStartEnd + 1, elEndStart)` runs even\n   when `elEndStart === -1`, extracting text backwards from position 0.\n3. **Unstable progression** (`lib/sax.js:556`): the function returns `elEndStart` (`-1`), driving\n   repeated character-by-character fallback in the parse loop.\n\n## Affected Versions\n\nOnly the `0.9.x` line is affected — the amplification was introduced in `0.9.0-beta.1` when\n`parseHtmlSpecialContent` was refactored, and remains through `0.9.11`. The `0.8.x` line is **not**\naffected: its older `parseHtmlSpecialContent` does not amplify, despite sharing the same\ncase-sensitive `indexOf`.\n\n## Proof of Concept\n\n```js\nconst { DOMParser, XMLSerializer } = require('@xmldom/xmldom');\n\nconst n = 1000;\nconst payload = '\u003chtml\u003e\u003cbody\u003e' + '\u003cscript\u003ex\u003c/ScRiPt\u003e'.repeat(n) + '\u003c/body\u003e\u003c/html\u003e';\nconst doc = new DOMParser().parseFromString(payload, 'text/html');\nconst out = new XMLSerializer().serializeToString(doc);\nconsole.log(payload.length, out.length, (out.length / payload.length).toFixed(1) + 'x');\n// 18026 9037063 501.3x  — an 18 KB input yields ~9 MB of output\n```\n\nOutput size grows quadratically with the number of case-mismatched raw-text elements:\n\n```\nRepeats | Input len | Output len | Ratio\n1       | 44        | 109        | 2.5x\n100     | 1826      | 93763      | 51.3x\n500     | 9026      | 2268563    | 251.3x\n1000    | 18026     | 9037063    | 501.3x\n2000    | 36026     | 36074063   | 1001.3x\n```\n\nProof of Concept from @KarimTantawey (tested with `script`); the same amplification occurs for `style`,\n`textarea`, and `title`.\n\n## Impact\n\nSmall attacker payloads can force disproportionate CPU and memory usage in services that\nparse and serialize untrusted HTML via xmldom. The quadratic growth means a modest-sized input\n(tens of kilobytes) can produce output in the tens or hundreds of megabytes, potentially\nexhausting memory or causing timeouts.\n\nThe attack only requires HTML mode (`text/html` MIME type) and mixed-case closing tags for\nany of the four raw-text element types. No special configuration or error handler setup is needed.\n\n## Severity note\n\nThe CVSS 4.0 vector scores availability only (`VA:H`, with `VC:N/VI:N`): the flaw neither discloses\nnor corrupts data, but a small untrusted HTML input (tens of KB) can force output and memory in the\ntens to hundreds of MB, enough to exhaust a service's heap or stall its event loop. It is reachable\nwith no authentication, configuration, or error-handler setup — only that the application parses\nuntrusted `text/html` and serializes the result.\n\n## Fix Applied\n\nThe raw-text closing tag is now matched case-insensitively in HTML raw-text mode (per the WHATWG HTML\n[RAWTEXT end-tag rule](https://html.spec.whatwg.org/multipage/parsing.html#rawtext-end-tag-name-state)),\nand a missing closing tag is handled explicitly, removing the quadratic output amplification. Output\nfor well-formed input is unchanged. Non-breaking; 0.9.x-only.","aliases":["CVE-2026-83612"],"modified":"2026-09-08T21:15:03.941065869Z","published":"2026-09-08T20:59:54Z","database_specific":{"severity":"HIGH","github_reviewed":true,"github_reviewed_at":"2026-09-08T20:59:54Z","nvd_published_at":"2026-09-01T15:17:39Z","cwe_ids":["CWE-178","CWE-400"]},"references":[{"type":"WEB","url":"https://github.com/xmldom/xmldom/security/advisories/GHSA-6mj3-qw4j-hgrw"},{"type":"ADVISORY","url":"https://nvd.nist.gov/vuln/detail/CVE-2026-83612"},{"type":"WEB","url":"https://github.com/xmldom/xmldom/pull/1071"},{"type":"WEB","url":"https://github.com/xmldom/xmldom/commit/7ced40c06c28d151e996a97045018c3559ae4707"},{"type":"PACKAGE","url":"https://github.com/xmldom/xmldom"},{"type":"WEB","url":"https://github.com/xmldom/xmldom/releases/tag/0.9.12"}],"affected":[{"package":{"name":"@xmldom/xmldom","ecosystem":"npm","purl":"pkg:npm/%40xmldom/xmldom"},"ranges":[{"type":"SEMVER","events":[{"introduced":"0.9.0-beta.1"},{"fixed":"0.9.12"}]}],"database_specific":{"last_known_affected_version_range":"\u003c= 0.9.11","source":"https://github.com/github/advisory-database/blob/main/advisories/github-reviewed/2026/09/GHSA-6mj3-qw4j-hgrw/GHSA-6mj3-qw4j-hgrw.json"}}],"schema_version":"1.9.0","severity":[{"type":"CVSS_V4","score":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N"}]}