Skip to content

Treat %00 pct-encoding triplets in URLs as invalid (#2505) - #2509

Merged
yadij merged 1 commit into
squid-cache:v7from
squidadm:v7-backport-pr2505
Sep 28, 2026
Merged

yadij merged 1 commit into
squid-cache:v7from
squidadm:v7-backport-pr2505

Conversation

@squidadm

Copy link
Copy Markdown
Collaborator

This AnyP::Uri::Decode() change addresses two primary sets of problems:

  • Since 2025 commit ad365ed, url_regex and urllogin ACLs ignored
    decoded URL suffix starting at the first %00 triplet; that
    known-to-be-problematic (e.g., see 2003 commit 3838e29 and
    CVE-2004-0189) behavior was explicitly disclosed in ACLs documentation
    and left for the followup work to fix. With this change, those ACLs
    operate on an encoded URL if that URL contains a %00 triplet.

  • The same 2025 commit ad365ed resulted in Squid FTP client sending
    truncated (at the first %00 triplet) URL paths to FTP servers during
    "slash hack" attempts (see 1997 commit 2761b2b). With this change,
    such FTP transactions fail instead.

Squid now treats %00 triplet as invalid, preventing URL decoding. We
chose this solution for several reasons, including these:

  1. Matching and forwarding a truncated URL instead would break cases
    where those URLs were working correctly, including deployments that
    do not use those problematic ACLs. No workaround would be available.
    Covering all URL imports would also require a lot more development.

  2. Either behavior may lead to incorrect access rules, but matching
    truncated URL while forwarding the whole URL (i.e. Squid behavior
    just before this change) may surprise admins more than matching and
    forwarding the same whole URL without decoding it (i.e. this code).

  3. Treating %00 as any other invalid triplet reduces the number of
    processing algorithms or special cases that developers and admins
    must account for:

    • Since 2024 commits cbb9bf1 and 226394f, url_regex and urllogin
      ACLs do not decode URLs that fail percent-encoding validation
      implemented by AnyP::Uri::Decode(). Since 2025 commit ad365ed,
      these ACLs match raw (i.e. undecoded) URLs instead. This commit
      extends that existing invalid pct-encoding triplet handling logic
      to %00 triplets.

    • This change removes a problematic special case from Squid's FTP
      "slash hack" failure recovery code path: A %00 transaction now
      fails without sending an MDTM recovery command, just like a
      transaction with a syntactically invalid URL pct-encoding.

    • As far as %00 preservation alone is concerned, this change can
      be seen as restoring the symmetry with handling of the URL
      "userinfo@" part: The corresponding AnyP::Uri::parse() code still
      uses rfc1738_unescape() and, hence, preserves %00 in userinfo.

  4. RFC 3986 permits special %00 decoding treatment: 'the "%00"
    percent-encoding (NUL) may require special handling and should be
    rejected if the application is not expecting to receive raw data
    within a component'. This change implements the "special handling"
    part. Admins may define "not expecting" deployments where Squid
    configuration should be rejecting matching requests (after rejecting
    the encoding itself).

  5. POSIX regex(3) API does not have a way of including a NUL character
    in the matching pattern. Some implementations have REG_ENHANCED mode
    that enables \x00 support, but glibc and, hence, many popular Linux
    distributions do not support that. Squid does not use REG_ENHANCED
    even if it is available. Preserving %00 allows admins to match that
    triplet in rare special cases that require detection of such URLs.

  6. Exposing unsuspecting Squid code to NUL-containing URL SBufs is still
    dangerous. While new code should be written to handle any SBuf bytes
    correctly, there are existing problematic SBuf uses, such problems
    are difficult for humans to spot in new code, and some external APIs
    (besides regex(3)) may require incompatible c-strings.

  7. As far as %00 treatment alone is concerned, this change restores
    how url_regex and urllogin ACLs worked before 2024 commits mentioned
    above. Squid v6 code still uses rfc1738_unescape() to implement those
    ACLs, so this helps reduce upgrade surprises and overheads for admins
    that have not upgraded to v7 yet.

Also updated url_regex and urlpath_regex documentation to provide more
examples and highlight dangers/surprises associated with those ACLs.

This AnyP::Uri::Decode() change addresses two primary sets of problems:

* Since 2025 commit ad365ed, url_regex and urllogin ACLs ignored
  decoded URL suffix starting at the first `%00` triplet; that
  known-to-be-problematic (e.g., see 2003 commit 3838e29 and
  CVE-2004-0189) behavior was explicitly disclosed in ACLs documentation
  and left for the followup work to fix. With this change, those ACLs
  operate on an encoded URL if that URL contains a `%00` triplet.

* The same 2025 commit ad365ed resulted in Squid FTP client sending
  truncated (at the first `%00` triplet) URL paths to FTP servers during
  "slash hack" attempts (see 1997 commit 2761b2b). With this change,
  such FTP transactions fail instead.

Squid now treats `%00` triplet as invalid, preventing URL decoding. We
chose this solution for several reasons, including these:

1. Matching and forwarding a truncated URL instead would break cases
   where those URLs were working correctly, including deployments that
   do not use those problematic ACLs. No workaround would be available.
   Covering all URL imports would also require a lot more development.

2. Either behavior may lead to incorrect access rules, but matching
   truncated URL while forwarding the whole URL (i.e. Squid behavior
   just before this change) may surprise admins more than matching and
   forwarding the same whole URL without decoding it (i.e. this code).

3. Treating `%00` as any other invalid triplet reduces the number of
   processing algorithms or special cases that developers and admins
   must account for:

   * Since 2024 commits cbb9bf1 and 226394f, url_regex and urllogin
     ACLs do not decode URLs that fail percent-encoding validation
     implemented by AnyP::Uri::Decode(). Since 2025 commit ad365ed,
     these ACLs match raw (i.e. undecoded) URLs instead. This commit
     extends that existing invalid pct-encoding triplet handling logic
     to `%00` triplets.

   *  This change removes a problematic special case from Squid's FTP
      "slash hack" failure recovery code path: A `%00` transaction now
      fails without sending an `MDTM` recovery command, just like a
      transaction with a syntactically invalid URL pct-encoding.

    * As far as `%00` preservation alone is concerned, this change can
      be seen as restoring the symmetry with handling of the URL
      "userinfo@" part: The corresponding AnyP::Uri::parse() code still
      uses rfc1738_unescape() and, hence, preserves `%00` in userinfo.

4. RFC 3986 permits special %00 decoding treatment: 'the "%00"
   percent-encoding (NUL) may require special handling and should be
   rejected if the application is not expecting to receive raw data
   within a component'. This change implements the "special handling"
   part. Admins may define "not expecting" deployments where Squid
   configuration should be rejecting matching requests (after rejecting
   the encoding itself).

5. POSIX regex(3) API does not have a way of including a NUL character
   in the matching pattern. Some implementations have REG_ENHANCED mode
   that enables `\x00` support, but glibc and, hence, many popular Linux
   distributions do not support that. Squid does not use REG_ENHANCED
   even if it is available. Preserving %00 allows admins to match that
   triplet in rare special cases that require detection of such URLs.

6. Exposing unsuspecting Squid code to NUL-containing URL SBufs is still
   dangerous. While new code should be written to handle any SBuf bytes
   correctly, there are existing problematic SBuf uses, such problems
   are difficult for humans to spot in new code, and some external APIs
   (besides regex(3)) may require incompatible c-strings.

7. As far as `%00` treatment alone is concerned, this change restores
   how url_regex and urllogin ACLs worked before 2024 commits mentioned
   above. Squid v6 code still uses rfc1738_unescape() to implement those
   ACLs, so this helps reduce upgrade surprises and overheads for admins
   that have not upgraded to v7 yet.

Also updated url_regex and urlpath_regex documentation to provide more
examples and highlight dangers/surprises associated with those ACLs.
@yadij
yadij merged commit e337390 into squid-cache:v7 Sep 28, 2026
10 of 11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants