Treat %00 pct-encoding triplets in URLs as invalid (#2505) - #2509
Merged
Merged
Conversation
This AnyP::Uri::Decode() change addresses two primary sets of problems: * Since 2025 commit ad365ed, url_regex and urllogin ACLs ignored decoded URL suffix starting at the first `%00` triplet; that known-to-be-problematic (e.g., see 2003 commit 3838e29 and CVE-2004-0189) behavior was explicitly disclosed in ACLs documentation and left for the followup work to fix. With this change, those ACLs operate on an encoded URL if that URL contains a `%00` triplet. * The same 2025 commit ad365ed resulted in Squid FTP client sending truncated (at the first `%00` triplet) URL paths to FTP servers during "slash hack" attempts (see 1997 commit 2761b2b). With this change, such FTP transactions fail instead. Squid now treats `%00` triplet as invalid, preventing URL decoding. We chose this solution for several reasons, including these: 1. Matching and forwarding a truncated URL instead would break cases where those URLs were working correctly, including deployments that do not use those problematic ACLs. No workaround would be available. Covering all URL imports would also require a lot more development. 2. Either behavior may lead to incorrect access rules, but matching truncated URL while forwarding the whole URL (i.e. Squid behavior just before this change) may surprise admins more than matching and forwarding the same whole URL without decoding it (i.e. this code). 3. Treating `%00` as any other invalid triplet reduces the number of processing algorithms or special cases that developers and admins must account for: * Since 2024 commits cbb9bf1 and 226394f, url_regex and urllogin ACLs do not decode URLs that fail percent-encoding validation implemented by AnyP::Uri::Decode(). Since 2025 commit ad365ed, these ACLs match raw (i.e. undecoded) URLs instead. This commit extends that existing invalid pct-encoding triplet handling logic to `%00` triplets. * This change removes a problematic special case from Squid's FTP "slash hack" failure recovery code path: A `%00` transaction now fails without sending an `MDTM` recovery command, just like a transaction with a syntactically invalid URL pct-encoding. * As far as `%00` preservation alone is concerned, this change can be seen as restoring the symmetry with handling of the URL "userinfo@" part: The corresponding AnyP::Uri::parse() code still uses rfc1738_unescape() and, hence, preserves `%00` in userinfo. 4. RFC 3986 permits special %00 decoding treatment: 'the "%00" percent-encoding (NUL) may require special handling and should be rejected if the application is not expecting to receive raw data within a component'. This change implements the "special handling" part. Admins may define "not expecting" deployments where Squid configuration should be rejecting matching requests (after rejecting the encoding itself). 5. POSIX regex(3) API does not have a way of including a NUL character in the matching pattern. Some implementations have REG_ENHANCED mode that enables `\x00` support, but glibc and, hence, many popular Linux distributions do not support that. Squid does not use REG_ENHANCED even if it is available. Preserving %00 allows admins to match that triplet in rare special cases that require detection of such URLs. 6. Exposing unsuspecting Squid code to NUL-containing URL SBufs is still dangerous. While new code should be written to handle any SBuf bytes correctly, there are existing problematic SBuf uses, such problems are difficult for humans to spot in new code, and some external APIs (besides regex(3)) may require incompatible c-strings. 7. As far as `%00` treatment alone is concerned, this change restores how url_regex and urllogin ACLs worked before 2024 commits mentioned above. Squid v6 code still uses rfc1738_unescape() to implement those ACLs, so this helps reduce upgrade surprises and overheads for admins that have not upgraded to v7 yet. Also updated url_regex and urlpath_regex documentation to provide more examples and highlight dangers/surprises associated with those ACLs.
yadij
approved these changes
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This AnyP::Uri::Decode() change addresses two primary sets of problems:
Since 2025 commit ad365ed, url_regex and urllogin ACLs ignored
decoded URL suffix starting at the first
%00triplet; thatknown-to-be-problematic (e.g., see 2003 commit 3838e29 and
CVE-2004-0189) behavior was explicitly disclosed in ACLs documentation
and left for the followup work to fix. With this change, those ACLs
operate on an encoded URL if that URL contains a
%00triplet.The same 2025 commit ad365ed resulted in Squid FTP client sending
truncated (at the first
%00triplet) URL paths to FTP servers during"slash hack" attempts (see 1997 commit 2761b2b). With this change,
such FTP transactions fail instead.
Squid now treats
%00triplet as invalid, preventing URL decoding. Wechose this solution for several reasons, including these:
Matching and forwarding a truncated URL instead would break cases
where those URLs were working correctly, including deployments that
do not use those problematic ACLs. No workaround would be available.
Covering all URL imports would also require a lot more development.
Either behavior may lead to incorrect access rules, but matching
truncated URL while forwarding the whole URL (i.e. Squid behavior
just before this change) may surprise admins more than matching and
forwarding the same whole URL without decoding it (i.e. this code).
Treating
%00as any other invalid triplet reduces the number ofprocessing algorithms or special cases that developers and admins
must account for:
Since 2024 commits cbb9bf1 and 226394f, url_regex and urllogin
ACLs do not decode URLs that fail percent-encoding validation
implemented by AnyP::Uri::Decode(). Since 2025 commit ad365ed,
these ACLs match raw (i.e. undecoded) URLs instead. This commit
extends that existing invalid pct-encoding triplet handling logic
to
%00triplets.This change removes a problematic special case from Squid's FTP
"slash hack" failure recovery code path: A
%00transaction nowfails without sending an
MDTMrecovery command, just like atransaction with a syntactically invalid URL pct-encoding.
As far as
%00preservation alone is concerned, this change canbe seen as restoring the symmetry with handling of the URL
"userinfo@" part: The corresponding AnyP::Uri::parse() code still
uses rfc1738_unescape() and, hence, preserves
%00in userinfo.RFC 3986 permits special %00 decoding treatment: 'the "%00"
percent-encoding (NUL) may require special handling and should be
rejected if the application is not expecting to receive raw data
within a component'. This change implements the "special handling"
part. Admins may define "not expecting" deployments where Squid
configuration should be rejecting matching requests (after rejecting
the encoding itself).
POSIX regex(3) API does not have a way of including a NUL character
in the matching pattern. Some implementations have REG_ENHANCED mode
that enables
\x00support, but glibc and, hence, many popular Linuxdistributions do not support that. Squid does not use REG_ENHANCED
even if it is available. Preserving %00 allows admins to match that
triplet in rare special cases that require detection of such URLs.
Exposing unsuspecting Squid code to NUL-containing URL SBufs is still
dangerous. While new code should be written to handle any SBuf bytes
correctly, there are existing problematic SBuf uses, such problems
are difficult for humans to spot in new code, and some external APIs
(besides regex(3)) may require incompatible c-strings.
As far as
%00treatment alone is concerned, this change restoreshow url_regex and urllogin ACLs worked before 2024 commits mentioned
above. Squid v6 code still uses rfc1738_unescape() to implement those
ACLs, so this helps reduce upgrade surprises and overheads for admins
that have not upgraded to v7 yet.
Also updated url_regex and urlpath_regex documentation to provide more
examples and highlight dangers/surprises associated with those ACLs.