{"id":"GHSA-8whx-365g-h9vv","title":"Loofah `allowed_uri?` does not detect `javascript:` URIs split by named whitespace character references","summary":"Loofah `allowed_uri?` does not detect `javascript:` URIs split by named whitespace character references","severity":"low","cwe":["CWE-184"],"vendor":"loofah","product":"loofah","ecosystem":"rubygems","affected":["loofah >= 2.25.0, < 2.25.2"],"patched":["loofah 2.25.2"],"published":"2026-07-21","updated":"2026-07-21","source":"GHSA","sourceUrl":"https://github.com/advisories/GHSA-8whx-365g-h9vv","references":[{"url":"https://github.com/flavorjones/loofah/security/advisories/GHSA-8whx-365g-h9vv"},{"url":"https://github.com/flavorjones/loofah/commit/5e91af861e3cdab47b91dd0b81f3afdfd13a5e19"},{"url":"https://github.com/flavorjones/loofah/releases/tag/v2.25.2"},{"url":"https://github.com/advisories/GHSA-8whx-365g-h9vv"}],"tags":["ghsa","rubygems"],"ingestedAt":"2026-07-21T15:51:23.763Z","slug":"GHSA-8whx-365g-h9vv","body":"## Overview\n\n## Summary\n\n`Loofah::HTML5::Scrub.allowed_uri?` does not correctly reject `javascript:` URIs when the scheme is split or prefixed by the HTML5 named character references `&Tab;` (tab) or `&NewLine;` (line feed).\n\nThis is a bypass of the fix for [GHSA-46fp-8f5p-pf2m](https://github.com/flavorjones/loofah/security/advisories/GHSA-46fp-8f5p-pf2m), which handled the equivalent numeric character references (`&#9;`, `&#10;`, `&#13;`) but did not cover the named forms.\n\n## Details\n\n`allowed_uri?` decodes HTML entities with `CGI.unescapeHTML`, which handles numeric character references but not HTML5 named character references. Payloads like `java&Tab;script:alert(1)` are therefore left intact, so the method does not recognize the `javascript:` scheme and returns `true`. A browser, however, decodes `&Tab;` and `&NewLine;` to tab and line feed and strips them from the URL during parsing, producing `javascript:alert(1)`.\n\n`&Tab;` and `&NewLine;` are the only relevant named character references: across the HTML5 named-character table, they are the only two that decode to characters the WHATWG URL parser strips from a URL (U+0009 and U+000A; there is no named reference for U+000D). `&nbsp;` / `&NonBreakingSpace;` decode to U+00A0, which browsers do not strip, so they aren't usable for this bypass.\n\nNote that Loofah's default `sanitize()` path is **not** affected, because Nokogiri decodes or entity-escapes HTML entities during parsing before Loofah evaluates the URI protocol. This issue only affects callers of the public `allowed_uri?` string-level helper that pass it HTML-encoded strings.\n\n## Impact\n\nCallers that validate a user-controlled URL with `Loofah::HTML5::Scrub.allowed_uri?` and then render the approved value into an `href` or other browser-interpreted URI attribute may be vulnerable to cross-site scripting (XSS). This includes applications that call `allowed_uri?` directly, as well as higher-level features built on top of it, such as Action Text 8.2's markdown link validation.\n\n## Mitigation\n\nUpgrade to Loofah >= 2.25.2.\n\n## Credit\n\nResponsibly reported by GitHub user @connorshea.\n\n## Affected packages\n\n- `loofah >= 2.25.0, < 2.25.2`\n\n## Remediation\n\nUpgrade to a patched release:\n\n- `loofah 2.25.2`","depth":"sunlit","depthScore":14,"depthScoreParts":{"impact":13.8,"likelihood":0,"exploitation":0,"ransomware":0},"changes":[]}