Skip to content

CI performance dashboard #1182

Description

Maintained by the CI profiler routine, which runs every 4h. Each run replaces this body and carries the history rows over.

Headline (run 2026-10-09 20:16 UTC)

Window: 2026-10-08 20:16 → 2026-10-09 20:16. Now = since the previous profile (16:16).

Sample:

  • Job detail for 46 runs:
    • all 14 successful or in-progress CI merge_group runs since 16:16
    • 9 of 36 cancelled merge_group runs
    • 23 pull_request runs, sampled evenly from completed runs (12 full, 11 draft/skip)
  • All 299 PR runs since 16:16 at run level.
  • 7 PR timelines.
  • About 250 REST requests. No job logs were read this run: queue wait dominates, not step time.

⚠️ The ubuntu-latest pool has been backlogged since ~18:00. Merge-group runs take 77–81 min instead of ~22.

metric 24h now (since 16:16)
PRs merged ≥99 (carried) 8: #1301 at 18:02 (direct), #1329 at 18:36, then 6 at 19:35–19:38
Merge queue last enqueue → merge – #1284, #1286, #1288, #1247: 111–112 min; #1274, #1290: 78 min
merge_group CI runs 209: 92 success, 103 cancelled, 6 failed 50: 6 success, 36 cancelled, 8 in progress
merge_group created → ci-ok – 77.3–80.7 min (n=6). Last test job done at 59–68 min.
PR run created → ci-ok (full runs) – p50 23.3 / p90 65.5 (n=12). Runs created before 17:30 took 20–23 min; after 17:45, 51–79 min.
PR runs 570 (299 since 16:16) 202 full + 97 draft/skip. 110 cancelled by concurrency, 44 failed (mostly #1293 before 18:02).
Linux queue wait p50 / p90 – merge_group 5.9 / 18.3, PR 0.5 / 9.0, max 22.9
macOS queue wait p50 / p90 (max) – 0.87 / 4.08 (5.4), n=488
Windows queue wait p50 / p90 (max) – merge_group 1.03 / 2.58 (3.3); PR 0.08 / 1.05

Merge-queue causes since 16:16:

Job-minutes per run (unchanged):

  • merge_group success: 312 Linux / 69 Windows / 60 macOS.
  • Full PR run: 353 Linux / 65 Windows.
  • Daily totals are carried from 16:16, about 150k Linux / 35k Windows / 13.6k macOS. PR volume in the last 4h ran at about 1.8× the 24h average.

Merge-queue settings (ruleset 24668460, read 20:2x): max_entries_to_build 8 (was 5), max_entries_to_merge 5, min_entries_to_merge_wait 5 min, ALLGREEN, timeout 90 min (was 120). Commented on #1248.

History

run (UTC) merges/24h MQ enq→merge p50/p90 (min) mg CI success p50 PR push→ci-ok p50/p90 mac queue p50/p90 linux queue p50/p90 job-min/day L/W/M
2026-10-09 20:16 ≥99 (8 since 16:16) last 78 / 112 (n=6, Linux backlog) 79.6 (n=6, 77–81) 23.3 / 65.5 (n=12; 51–79 after 17:45) 0.87 / 4.08 5.9 / 18.3 mg (16.6 p50 at 19h) ~150k / ~35k / 13.6k (carried over; CI per-run unchanged: mg 312 / 69 / 60)
2026-10-09 16:16 ≥99 first 84 / 509 · last 76 / 94 (n=13, burst drain) 22.3 (n=30) 21.8 / 24.6 (n=23) 0.67 / 2.73 0.07 / 2.12 ~150k / ~35k / 13.6k (CI 131k / 24k / 13.6k)
2026-10-09 12:24 72 first 107 / 168 · last 69 / 107 (n=8, burst drain) 23.1 (22.6 since 08:17) 21.5 / 22.3 (n=11, after #1166) 0.1 / 0.4 0.0 / 3.9 ~150k / ~40k / 9.4k (CI 106k / 19k / 9.4k, compat carried over)
2026-10-09 08:17 69 first 65 / 82 · last 64 / 69 (burst) 23.4 (21.5 since 04:17) 31.7 / 34.2 (n=58) 0.2 / 1.8 0.1 / 1.4 ~180k / ~40k / 10.1k (Gradle compat not re-measured)
2026-10-09 04:16 ~66 first 100 / 380 · last 86 / 142 (47 since 21:35) 25 (22 since 21:35) 31 / 33 (n=7) 0.2 / 1.9 0.1 / 3.3 ~263k / ~89k / 13.4k (Gradle carried over)
2026-10-09 00:17 72 first 48 / 290 · last 26 / 78 39.1 (22.6 after #1143) 47.2 / 56.9 (31.7 / 33.8 now) 0.2 / 1.7 0.4 / 3.7 295k / 92k / 10k (all workflows, 24h)
2026-10-08 21:50 49 80 / 157 44.5 (23.9 post-#1133) 43.2 / 71.9 (29.0 / 30.3 post-#1133) 22.0 / 55.5 1.2 / 12.8 97.5k / 28.3k / 12.0k (7-day avg)
2026-10-08 19:36 43 83 / 177 42.6 (22–25 after #1133) 41.6 / 78.5 1.1 / 69.8 0.1 / 7.6 139k / 21k / 21k (CI only, 24h; L/M/W order)

Open ci-perf issues by ROI

ROI on one scale: (thousands of job-min/day weighted Linux ×1, Windows ×2, macOS ×3, plus 1 per critical-path or PR-latency minute) × confidence ÷ effort (S=1, M=2, L=3).

ROI issue saving status
~20 #1170 CI on push to main re-tests the SHA the queue passed ~24k Linux + 3.8k Windows + 3.4k macOS PR #1355 open (identical-SHA reuse)
8.8 #1373 PR CI re-tests unchanged diffs after merge-main pushes (new) ~20–30k Linux + 4–6k Windows; shrinks the Linux backlog
7.8 #1372 ci-ok waits 9–18 min for ubuntu-latest (new) −9 to −18 min merge-queue critical path under backlog S fix
7 #1248 (settings) merge-queue build concurrency Build raised to 8; effect masked by the backlog human decision
6.0 #1300 Gradle compat: 4 Windows hosted cells on every matching PR ~5k Windows PR #1353 open
3.9 #1267 Gradle e2e PR tier: all 4 Gradle lines on every PR ~9.8k Linux partly in PR #1355
2.6 #1225 coverage-docker (sbt): image rebuild on the critical path 1–2 min partly in PR #1355
2.1 #1176 e2e-macos: one-minute macOS jobs ~1k macOS
2.1 #1171 Gradle e2e shards unbalanced ~1–2 min per merge_group run Gradle e2e was the last test job in 2 / 6 groups today; coverage-merge in 4
1.9 #1172 e2e fan-out see issue
1.3 #1173 test (windows/macos): redundant Build step ~2 min per leg
0.42 #1174 coverage runs llvm-cov report twice ~530 Linux

Open ci-perf PRs:

Merged ci-perf / ci-janitor PRs (last 7 days) and measured effect

PR merged before → after
#1284 skip vlt compat on ci.yml-only changes 10-09 19:35 Too early to measure.
#1286 retry rustup installs (ci-janitor) 10-09 19:35 No rustup failures in the sampled runs.
#1275 path-filter compat on push to main 10-09 15:46 Only 3 main pushes since 16:16. Not re-measured.
#1255 skip duplicate Gradle compat legs on PRs 10-09 14:09 Gradle compat PR run 536 → 209 job-min. Delivered for Linux.
#1261 build coverage-docker's binary once per leg 10-09 13:43 coverage-docker 59 → 45 (merge_group), 69 → 42 (PR) job-min per run. Delivered.
#1206 compat matrices on PRs only for their own files 10-09 13:25 Compat PR runs per CI PR run ~8 → 1.5. Delivered.
#1166 shard test-release into 3 10-09 08:40 PR push → ci-ok 31.7 → 20–23 min when unloaded (16:17–17:15). Holding.
#1234 merge e2e rows that share a toolchain 10-09 09:10 Legs 125 → 103. Did not deliver in minutes.
#1222, #1251 retries 10-09 09:07 / 13:45 No Docker-fetch or uv eviction since.
#1143 decouple CI platform builds 10-08 21:34 Unloaded merge_group p50 ~22. Today masked by the backlog.
#1146 fail-fast eviction 10-08 21:05 Working.
#1208, #1189, #1226, #1229, #1215, #1201 (retry and flake fixes) 10-09 00:13 – 08:01 No recurrence.
#1133, #1093, #1179, #1139, #1102, #1087, #1080, #1059, #1018, #619, #897, #892, #874 10-02 → 10-08 See earlier rows.

Method notes

  • GraphQL is blocked in this environment, so merge-queue state comes from REST timelines of merged PRs.
  • Job listings are fetched with filter=latest, page 2 explicitly (merge_group runs have 148 jobs).
  • Queue wait = job started_at − created_at. Skipped jobs are excluded, and so are cancelled jobs that never got a runner.
  • A "full" PR run is one still running or longer than 6 min. Draft/skip runs have 30 jobs, 29 of them skipped.
  • 24h run counts use workflows/ci.yml/runs?event=…&status=… total_count.

Generated by Claude Code

Activity

  1. mikolalysenko commented on Oct 8, 2026

    @mikolalysenko
    CollaboratorAuthor

    [agent] There are now two open dashboard issues: this one and #1175, both titled "CI performance dashboard" and labeled ci-perf-dashboard. The body says it is maintained in place, so the profiler routine should update one issue rather than open a new one each run. I'm not closing either one, because I can't tell which issue the routine will update next. Whoever owns the profiler should close the copy it no longer updates.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions