Repository navigation
Cache is very slow #2853
Description
Activity
- addedstate: needs triageWaiting to be triaged by a maintainer.Waiting to be triaged by a maintainer.
on May 20, 2026 @Napolitain How many files you have in
path/to/folder? Is it a big directory? How many files? Also, how long it takes to run a task?- addedstate: awaiting responseWaiting for issue author to respond.Waiting for issue author to respond.
on May 25, 2026 @Napolitain How many files you have in
path/to/folder? Is it a big directory? How many files? Also, how long it takes to run a task?I have created a benchmark PR for trying to add a baseline and then I think we can find some areas of improvements. I think there are many redundancy and allocs.
The benchmark is 20k small files and a few large files. It shows that many small files is dramatically slower than few large files. Of course, it is expected, but that doesn't mean there is no issues either. Maybe with the benchmark then we can try to improve speed against baseline.
I have tried more simple commands like
sha256sumand it is really much faster (again, to be expected, but nonetheless, there may be area of improvements).- removedstate: awaiting responseWaiting for issue author to respond.Waiting for issue author to respond.
on Jun 12, 2026 - linked a pull request that will close this issuetest: add benchmarks for glob cache performance #2881
on Jun 12, 2026 - added a commit that references this issue
on Jun 13, 2026 @Napolitain I'm developing a concurrent-streaming-zeroalloc algo for timestamping. The technique is to create a ReadDir func and inject that into the existing glob mechanism, which effectively bypasses all the allocations and allows early exit if a newer file is found. Since ReadDir loads file info, the repeated calls to os.Stat are no longer required.
I don't really trust the results, some of the boilerplate code is AI generated, but it seems like a good strategy ... but also requires more tests probably. At least one test is failing .... so there may be more effort to get the algorithm correct.
The gist is here:
https://gist.github.com/trulede/bb3f23892b116a91bbfc21a665d5353bThe results:
Benchmark Configuration Tier Original Strategy Interceptor Strategy Real World 2,000 Files / Depth 1 16.02 ms 0.049 ms ~320x Faster 2,000 Files / Depth 10 19.11 ms 0.054 ms ~350x Faster 20,000 Files / Depth 1 252.81 ms 0.117 ms ~2,160x Faster 20,000 Files / Depth 10 (0% Chg) 249.79 ms 0.138 ms ~1,800x Faster 20,000 Files / Depth 10 (1% Chg) 252.64 ms 0.062 ms ~4,000x Faster Allocations were reduced from 84Mb to 12.7Kb (20.000 files).
- addedarea: fingerprintingChanges related to checksums and caching.Changes related to checksums and caching.and removedstate: needs triageWaiting to be triaged by a maintainer.Waiting to be triaged by a maintainer.
on Jun 20, 2026 @Napolitain I updated the Gist, after some changes needed to get all tests working. Performance of the algorithm is now 1,280x faster (0% change) and 1,810x (1% change).
I will leave it up to you, if you want to take these changes into your PR and run your benchmarks.
@trulede I propose to merge #2883 and #2884 first, and then you could open another PR with further changes if you want.
The PRs from @Napolitain are simple, and this way we also avoid conflicts.
What do you think?
- added a commit that references this issue
on Jul 13, 2026
Description
Caching some very lightweight files with this syntax
path/to/folder/**/*.yamlVersion
Nightly
Operating system
Ubuntu
Experiments Enabled
No response
Example Taskfile