Repository navigation
feat(subsystembenchmarks): add per-case rapid_cache_cold and rapid_cache_warm bucket types - #1085
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces support for GCS Rapid Cache (Anywhere Cache) in the subsystem benchmarks by adding two new bucket types: rapid_cache_cold and rapid_cache_warm. It implements the lifecycle management of Anywhere Caches, including creation, polling for RUNNING state, warming for read scenarios, disabling, and cleaning up leaked caches and buckets in Cloud Build. Additionally, configuration schemas, benchmark runners, and tests have been updated to support these new cache types. There are no review comments, so I have no feedback to provide.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1085 +/- ##
==========================================
+ Coverage 90.11% 90.25% +0.14%
==========================================
Files 16 16
Lines 3408 3458 +50
==========================================
+ Hits 3071 3121 +50
Misses 337 337 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
b69dbb3 to
214f9ed
Compare
214f9ed to
3b8fb41
Compare
…to bucket and read_case
…rm in CLI and Cloud Build
…amplification scrape
… and CLI constants
…c monitoring scrape
…solution, and bucket teardown
…nd job timeout to 18000s GCS Anywhere Cache creation LROs in us-central1-b occasionally take 33-42 minutes (2,001s-2,537s) to transition from CREATING to RUNNING, exceeding the previous 1,800s timeout. Increase DEFAULT_TIMEOUT_SECONDS and _RAPID_CACHE_TIMEOUT to 3600s and raise the Cloud Build job timeout to 18000s (5h).
…e 2 warmup passes on rapid_cache_warm
…epoch Time one untouched-corpus epoch on the same VM, bucket and corpus before warming, and publish its throughput plus the warm/paired-cold speedup. Cross-build cold vs warm comparisons were dominated by VM variance. Create both Rapid Cache types with ingestOnWrite=false so the paired cold epoch is genuinely uncached, scrape per-window cache hit/miss bytes, and flag warm rows whose hit ratio is below --min-warm-hit-ratio.
Both start a fresh DataLoader and pay worker spawn plus connection setup; later warm rounds reuse persistent workers, so dividing by the warm mean would credit the cache with setup savings it did not earn.
Round 1 of a persistent-worker DataLoader carried worker spawn, imports and gcsfs auth/connection set-up (time to first batch 3.4-7 s on n2-standard-64) that later rounds do not, and that varies by seconds between runs. Workers now prime gcsfs with a metadata-only request and park at a gate before the round barrier; round 1's clock and the gate open together. Start-up is reported in dataset_build_time instead.
… hit-ratio column Rapid Cache hits are reliable on n2 VMs (b/570794310 tracks the c4/c3 campus-placement issue), so drop the workarounds added for c4 and cross-VM noise: - Remove the paired same-VM cold epoch on rapid_cache_warm and its six rapid_cache_paired_cold_* / warm-first-round schema columns; warm runs again use ingestOnWrite=true, untimed warmup passes, then timed rounds. - Remove the WebDataset worker start gate and gcsfs priming. - Replace rapid_cache_hits.py, --min-warm-hit-ratio and the hit/miss byte columns with a best-effort rapid_cache_hit_ratio column filled by the existing amplification scrape, with no build gate. - Hard-code the 3600 s Rapid Cache creation timeout instead of plumbing --rapid-cache-timeout through Cloud Build, the runner and the env. - Simplify warm_if_needed to a single url_to_fs/find/open path.
8c594fa to
e2c5197
Compare
Summary
Adds
rapid_cache_coldandrapid_cache_warmbucket types to the subsystem benchmark harness so data-loading and checkpointing benchmarks can run against per-case GCS Rapid Cache (Anywhere Cache) instances.dataloading/rapid_cache.py,dataloading/bucket.py,dataloading/configurator.py,checkpointing/configurator.py):BUCKET_TYPESwithrapid_cache_cold(rccold) andrapid_cache_warm(rcwarm). Both create a regional bucket inspec.locationand provision an Anywhere Cache inspec.zoneviaPOST b/{bucket}/anywhereCaches(ingestOnWrite=Falseforrapid_cache_cold,ingestOnWrite=Trueforrapid_cache_warm), pollingGET b/{bucket}/anywhereCaches/{zone}every 10s untilstate == "RUNNING"(default timeout3600s)._deletecallsPOST b/{bucket}/anywhereCaches/{zone}/disable, removes all objects under the bucket, and suppresses errors onfs.rmdir(name)because GCS blocks bucket deletion during the 1-hour grace period after disabling an Anywhere Cache.dataloading/read_case.py,dataloading/rapid_cache.py,checkpointing/checkpoint_case.py):run_read_caseoverridesrounds=1whenparams.bucket_type == "rapid_cache_cold"so subsequent rounds on the same bucket (which hit the cache after Round 1'sadmit-on-first-miss) are not averaged into the cold measurement.warm_if_needed(invoked before timed rounds inrun_read_caseand inrun_checkpoint_casewhen"read" in params.scenario) runs whenbucket_type == "rapid_cache_warm"on ags://prefix. It reads every object under the prefix in 16 MiB chunks across up to 16 threads for 2 passes (DEFAULT_WARMUP_PASSES = 2) with a 65s settle delay (DEFAULT_WARMUP_SETTLE_SECONDS = 65) after each pass, then invalidates the filesystem cache. This allows asynchronousadmit-on-first-missadmission to finish for any shards not admitted during concurrentingest()writes and separates warmup reads from the timed 60s Cloud Monitoring alignment window.dataloading/amplification.py):bucket_typeisrapid_cache_coldorrapid_cache_warmand bucket-levelnetwork/sent_bytes_countandapi/request_countboth returnNoneor0.0(when reads are served from the cache rather than origin bucket egress),enrich_csvnormalizes both to0.0so--require-amplificationsucceeds.gcs_read_bytes,gcs_read_request_count, andgcs_read_amplification_ratiotogether only when both egress bytes and request count are non-Noneandideal > 0.run.py,cloudbuild/subsystembenchmarks/*):--zonewhen--bucket-typeiszonal,rapid_cache_cold, orrapid_cache_warm, and adds--rapid-cache-timeout(GCSFS_SUBSYSTEM_RAPID_CACHE_TIMEOUT, default3600).subsystembenchmarks-cloudbuild.yamlto acceptrapid_cache_coldandrapid_cache_warm, pass_RAPID_CACHE_TIMEOUT, increase the job timeout to18000s(and SSH key TTL to5h), and disable active Anywhere Caches before cleaning up case buckets indelete-bucketsandcleanup-leaked-resources(scoped by regex to<prefix>-(regional|zonal|hns|rapid_cache_cold|rapid_cache_warm)-<8hex>-).dataloading/webdataset/configs.yaml):{axis: "shard_size", file_count: 256, rows_per_file: 196, enabled: false}(the unbuffered 256-shard variant), changing the default enabled WebDataset sweep from 11 to 10 cases.Tests
rapid_cache(test_rapid_cache.py), bucket lifecycle (test_bucket.py),run_read_casesingle-round cold override and warmup ordering (test_read_case.py),run_checkpoint_casewarmup hook (test_checkpoint_case.py), amplification enrichment (test_amplification.py), configurator abbreviations (test_configurator.py,checkpointing/tests/test_configs.py,ray_pytorch/tests/test_configs.py), CLI/Cloud Build wiring (test_run_groups.py), and WebDataset default sweep count (webdataset/tests/test_configs.py).