A panel balanced root-cause rewrites against path-aware test gating, agreeing on retries only if first-attempt failures stay visible.
“If the watcher rewrite overruns, routing is the first item I would defer—not diagnostics or retry transparency.”
All three approved the 15-hour hybrid to cut CI wait times, but split on what to cut if the clock runs out: defer path-aware gating until the main watcher race is solved, or ship gating first to secure the latency win.
Ship B+D: fix the watcher race at the readiness and semantic-assertion boundaries, use narrow nextest retries for residual flakes, and keep the full matrix blocking pre-merge.
Use a test-only readiness acknowledgement or sentinel before renaming, then assert the final graph state while accepting backend-equivalent event forms; fsync and ordered renames alone do not prove watcher registration or delivery.
Run the full matrix with at most two process-level retries for the named flaky tests, preserve first-attempt logs and JUnit results, and avoid nested remote-cache retries; otherwise a real intermittent bug can pass while its original failure is obscured.
A fast Linux-only PR gate would improve first feedback, but it shifts the current 41% flaky-red problem into the merge queue, risking queue contention and late author context switches; keeping all platforms required better preserves main-green confidence.
Ship a corrected B + selective D hybrid: fix the watcher test properly, raise the remote-cache deadline without adding in-test retries, and use nextest for two process-level retries only on the known flaky tests. Keep the entire 12-job matrix required.
The watcher and remote-cache cases account for 56% of the sampled flaky reds, so this attacks most of the immediate pain while preserving platform coverage. I would not choose A or C: they remove precisely the tests and platforms that can catch regressions. E merely automates the current expensive whole-job rerun. Broad D alone would improve DX but normalize nondeterminism and make probabilistic regressions easier to merge.
TempDir + fsync + deterministic rename order is not sufficient. fsync establishes durability, not that the OS watcher has registered or delivered its events. Nor should a cross-platform watcher test require one exact raw event sequence; rename notifications can be represented or coalesced differently by different backends. An exact event-order snapshot will therefore remain vulnerable even after adding fsync.
Make the test assert the product-level invariant instead:
TempDir, and log its path and watcher backend.old to new.old is absent, new is present, and the incremental graph invalidated/rebuilt the expected node. Do not snapshot raw arrival order.If the sentinel itself is not appropriate for the production abstraction, add a test-only readiness acknowledgement at the watcher boundary. That is cleaner than sleeping.
Suppose macOS registers the watch late. Today the test renames immediately, receives no qualifying event, and the job turns red after 18 minutes.
With this change, the rename cannot occur until the sentinel has demonstrated that the event path is live. That should eliminate the common registration race rather than conceal it. If the rename is observed through a backend-specific combination such as remove/create rather than one paired rename, the semantic predicate still passes. If the graph never recognizes the final path, the test fails—correctly.
On that failure, nextest reruns just watcher_test::test_rename in a fresh test process, at most twice. The first-attempt event log and JUnit result must still be archived even if a retry passes. Nextest supports per-test retries and integrates flaky results with JUnit reporting. A persistent rename regression should fail all three attempts and block the PR. A timing-sensitive real regression may pass a retry and merge; that is the principal risk of this design.
I would configure retries by explicit test filter, not globally:
[[profile.ci.overrides]]
filter = 'test(watcher_test::test_rename) | test(remote_cache::integration) | test(graph_order_snapshot)'
retries = { count = 2, backoff = "fixed", delay = "250ms", jitter = true }
Use the repository’s exact nextest syntax/version, but keep the allowlist narrow and checked into source control.
For remote_cache::integration, increase the outer deadline from 2 seconds to 10 seconds, but do not add a three-try loop inside the test. If the fixture has a server-startup race, add a readiness channel or wait for the listener to bind. Because the details of this test were not searched, that diagnosis is conditional. Process-level nextest retries provide cleaner isolation and retain the original failed attempt; nested retries would obscure which operation failed and could multiply the worst-case timeout substantially.
Suggested allocation:
Canonicalizing the graph before snapshotting should remove order noise if the snapshot is intended to represent an unordered graph. If order is itself a contractual invariant, retain the assertion and fix the traversal instead.
Retrying an individual failed test should consume materially fewer paid minutes than rerunning an 18-minute job or matrix, although the exact saving depends on test duration and which jobs users currently rerun. The 10-second remote deadline increases only that test’s timeout path, while successful executions should remain fast.
The invariant I would put on the dashboard is:
First-attempt pass rate, before retries, must exceed 99% per test over a rolling 200 executions; no test may remain on the retry allowlist for more than 14 days without an owner and issue.
Also report “green only after retry” separately from ordinary green. Do not let retries improve the displayed first-attempt metric. Without that separation, the retry layer can make CI appear healthy while underlying nondeterminism worsens. Alert if overall first-attempt PR green falls below 95% or any allowlisted test’s failure rate rises week over week.
This preserves the full pre-merge signal, fixes the largest race at its synchronization boundary, and uses retries as a bounded shock absorber rather than as quarantine.
Hybrid B + D: deterministically fix the top two flake classes this week, and put nextest retries + JUnit under the full matrix.
This is the only combination that removes the dominant root causes permanently, keeps every test (including the ones that catch real bugs) in the critical path, preserves main green confidence, stays inside GitHub Actions + paid minutes, and does not create long-term maintainer toil. Fifteen engineering hours is enough for a focused rewrite of the two hottest suites plus the nextest plumbing; the remaining flake classes (graph order, mac disk, random) become cheap to absorb and visible.
B eliminates the races at the source so the tests remain high-signal; D makes residual non-determinism cheap and observable. Disk pressure is handled as a one-line workflow hygiene step that both B and D benefit from.
watcher_test::test_rename FS raceToday the test creates a TempDir, starts a notify watcher, performs a rename (or sequence of renames), and asserts on the event stream. Under load or on slower macOS runners the kernel event, the rename, and any parent-directory fsync are racy; events can be coalesced, reordered, or delivered after the assertion window, producing the 32 % flake rate.
Under the hybrid:
B rewrite (this week)
TempDir (or an explicit workspace subdirectory) so cleanup is deterministic.file.sync_all() / dir.sync_all() (or the platform equivalent) so the FS event is guaranteed to be durable before the watcher is expected to see it.AtomicBool + spin/yield) that the test thread waits on after Watcher::watch returns and before the rename; the watcher thread signals once it has processed the initial create.RenameMode::Both pairs first, then rewrite any subsequent paths that match a source).D safety net
Even if a residual platform-specific coalescing bug remains, nextest is configured with
[profile.default]
retries = { count = 2, backoff = "fixed", delay = "500ms", jitter = true }
and the job uploads JUnit. A first-try failure followed by a pass is recorded as flaky, the job itself stays green, and the JUnit artifact feeds a simple dashboard or GitHub check annotation.
Would the failure be caught?
False-positive risk
Near zero for this suite once B lands; the test remains in the required matrix so a true rename bug still fails the build. The 5 % “real bug on retry” rate applies only to the residual flakes that D absorbs; we accept that small risk because the alternative (quarantine) is worse and because JUnit keeps the signal visible.
Minutes cost
TMPDIR=$GITHUB_WORKSPACE/.tmp) removes the 14 % mac pressure class with negligible cost.Failure modes the hybrid itself can introduce
.config/nextest.toml with the two-retry profile, change the matrix jobs to cargo nextest run, add the JUnit upload action, add the macOS disk-reclaim + TMPDIR lines.Track three numbers weekly:
watcher_test::test_rename (must be ≥ 99 % after the rewrite).If (1) stays high and (3) does not trend upward, main stays green, authors see green in ~4-6 minutes of useful work, and we never increased long-term toil. If the flaky count creeps up we still have the deterministic fixes and the JUnit data to attack the next suite; we have not painted ourselves into a quarantine corner.
This hybrid is the highest-confidence, lowest-toil path that actually ships inside the freeze window.
For 15hrs this week, ship hybrid C + D + B-light. Don't do pure A/B.
C is primary for DX, D for stability, B-light for ROI.
cargo nextest run --lib --test=graph --skip watcher --skip remote . PR authors unblocked.merge_group / bors. Main only updates if this passes. That's how you keep green confidence without new vendor..config/nextest.toml[profile.default]
retries = 2
fail-fast = false
[[profile.default.overrides]]
filter = 'test(watcher_test) | test(remote_cache)'
retries = { backoff="exponential", count=3, delay="200ms", jitter=true }
Upload JUnit artifact. This kills outer reruns. Outer rerun = 18m * 12 jobs = 216 compute-min for nothing. Inner retry = seconds.
Minimal B fix - only watcher_test::test_rename - 4-6h, not full rewrite:
TempDir::new_in("./.tmp") not OS tmp, fix EACCES from sudo steps on mac runnersrecv_timeout before renamefsync dir, barrier, renameMac disk pressure 14%: add 2-line cleanup before test: sudo rm -rf /usr/share/dotnet /opt/ghc + export TMPDIR=$GITHUB_WORKSPACE/.tmp && mkdir -p $TMPDIR && chmod 700 $TMPDIR this is proven fix for macOS-15 permission/disk flakes
Time budget: C config 1h, nextest+junit 3h, B-light watcher 6h, mac cleanup 1h, graph snapshot sort 2h = 13h.
Old path: test creates file A, immediately rename(A,B). On linux inotify / mac FSEvents, notify coalesces Create+Rename or delivers out of order. No fsync, no wait for watcher subscription. TempDir drops and deletes B before assertion. Fails 32% of your last 50 reds.
Under hybrid:
flaky.Would real bug be caught? Yes, with nuance:
flaky-result = warn but dashboard alerts, and we treat any pass-on-retry as merge-queue soft-fail requiring issue link. This preserves signal without blocking forever.Why not the others?
Don't ship blind. Add workflow that fails if these breach:
pr_fast_pass -> merge_queue_fail rate >15% -> split gate too permissive, add back tests to fast gateouter rerun count (current 2.1 to green, 41% PRs red). Target after: <0.3 outer reruns, <10% PRs red.flaky retry-success rate per test from JUnit: if watcher_test retry-success >5% of runs after B-light fix, fix incomplete - revert to quarantine for that single test onlycompute minutes per merged PR: should drop ~30% vs now despite inner retries. Alert if rises >20%time-to-fast-green p50: must stay <5mIf by Thursday watcher fix not landing, cut it and keep C+D. That still gets you from 41% red PRs to ~10% with 4h work, and main green guarantee holds via merge queue.
Ship conservative path-aware C + corrected B + narrow D: the 80% fast path gets a 4-minute gate, sensitive or ambiguous changes get the full 12-job matrix before merge queue, and every merge-group candidate still gets that matrix.
Sixteen PRs pay 4 minutes and four pay 18, totaling 136 gate-wait minutes/day versus 360 for full-matrix pre-merge on all 20. Pure C would shift today’s roughly 8.2 flaky PRs/day and 17.2 reruns into the queue; queue wall throughput cannot be priced beyond 4,320 compute-minutes/day without concurrency and service-time data, and routing must conservatively include shared dependencies, CI, platform-sensitive paths, and classifier failure.
The watcher fix remains sentinel-observed readiness plus semantic final-state assertions, not fsync or raw event order; sensitive changes stay full-matrix pre-merge, while routing falls back to full if classification is uncertain. This bounds rather than eliminates late real-bug risk, and if time runs short, defer routing before the dominant root-cause fix.
Allocate 6h watcher, 2h remote-cache readiness and deadline, 2h nextest/JUnit/per-attempt reporting, 1h graph canonicalization plus mac reclamation, 2h conservative routing, and 2h stress validation and dashboard wiring. The dashboard rule is: “A retried pass is YELLOW, never GREEN; first-attempt pass rate must exceed 99% per test over 200 executions; every flaky pass automatically creates or links an owned issue; and each allowlist entry expires after 14 days.”
I’m choosing path-aware C + B + narrow D. The fast gate should be the default for the 80% of PRs it meaningfully covers, while PRs touching watcher, cache, platform abstraction, CI, or their shared dependencies get the full matrix before entering the queue. Every merge candidate still gets the full matrix in the existing merge queue.
At today’s rate:
Pure C does not eliminate any failures; it relocates them. At current rates:
That makes pure C unattractive despite its quick first signal. The path-aware version changes where relevant feedback arrives:
The point of B plus D is to make that baseline obsolete before enabling the routing change. The four agreed deterministic/infra changes address categories representing 88% of the sampled reds: watcher 32%, remote cache 24%, graph ordering 18%, and macOS pressure 14%. Those percentages cannot simply be subtracted from the 41% PR rate, but they make a post-change target below 10% credible enough to test. At a 10% residual full-matrix red rate, the fast-routed population would create about 1.6 late contexts/day, not 6.56. If we do not achieve that, the routing rollout has failed and should be widened or reverted.
Implement routing with a checked-in path classifier and a required aggregator job:
ci-full label.merge_group.watcher_test::test_rename around sentinel-observed readiness, semantic final-state assertions, backend-equivalent events, failure diagnostics, and stress runs.remote_cache::integration: explicit fixture readiness where possible and a 10-second outer deadline; no nested retries.TMPDIR, and disk/ownership diagnostics on failure.ci-full override.That is exactly 15 hours. If the watcher rewrite overruns, routing is the first item I would defer—not diagnostics or retry transparency. Retries without observability are more dangerous than retaining the current gate for another week.
A retried pass is YELLOW, never GREEN. Every retry-allowlisted test must maintain at least 99% first-attempt passes over its rolling 200 executions. Its first flaky-pass must automatically create or link an owned issue, and its allowlist entry expires after 14 days. Breaching 99% or reaching expiry without a merged fix makes flaky-pass blocking until an explicit, time-bounded exception is approved. The dashboard must report first-attempt pass rate, flaky-pass count, issue owner, allowlist age, and post-merge regressions attributed to retried tests.
“Automatically create or link” matters: requiring humans to curate links manually would turn transparency into recurring maintainer toil.
Surviving positions, one line each:
Gate decision: path-aware hybrid (fast gate for the 80 %, full matrix required on PR for the 20 % that touch watcher/remote/platform paths, full matrix always on the existing merge queue).
This is the only routing that survives the arithmetic under the author stipulations. Pure full-matrix-pre-merge (@GPT’s line) and pure C (@Muse’s primary) both lose.
paths / dorny/paths-filter (or equivalent) makes the full 12-job matrix a required PR check only when the diff touches the watcher, remote-cache, or platform crates; otherwise only the 4 min linux-stable --lib gate is required. Merge queue continues to run the full matrix for every PR as it already does.The path filter is a one-hour, already-supported GHA change; it is not new infrastructure. This is why I reject both “never C” and “C for everyone.”
@GPT, your insistence on keeping the entire matrix required is right for the 20 % and for main, but the 80/20 fact plus the already-enabled merge queue make the pure version an unnecessary 200+ author-min/day tax on the majority. The path-aware rule preserves exactly the tests that can catch regressions where they matter. @Muse, elevating pure C as primary creates precisely the late context switches on the sensitive PRs that the fast gate cannot see; B must also be the full sentinel + semantic version, not the light TempDir/fsync-only sketch.
watcher_test::test_rename | remote_cache::integration | graph_order_snapshot), two process-level retries, JUnit + first-attempt log archival.rm of unused toolchains) + TMPDIR=$GITHUB_WORKSPACE/.tmp.If the watcher sentinel work is not landing by Thursday end-of-day we cut it, keep the rest, and the path-aware C + D still ships; the 99 % rule will force the issue open.
First-attempt pass rate (before any retry) must exceed 99 % per test over a rolling 200 executions; no test may remain on the nextest retry allowlist for more than 14 days without an assigned owner and linked issue; every flaky-pass is reported separately from ordinary green, must link or open an issue, and never improves the first-attempt metric; alert if overall first-attempt PR green falls below 95 % or any allowlisted test’s first-attempt failure rate rises week-over-week.
This package keeps main’s green confidence, eliminates the 216-compute-min outer reruns, gives 80 % of authors a 4 min loop, protects the 20 % that actually need the flaky suites, and makes nondeterminism impossible to normalize. Ship it.
As @Muse - adopting floor, pricing gate with arithmetic.
Gate decision: path-routed split gate + floor fixes + narrow nextest.
Pure full-matrix pre-merge with floor fixes keeps signal but you still pay 18m wall per push for everyone. Pure fast-gate for all PRs moves blindness to merge queue. Smart routing fixes both:
# pseudocode shippable this week with dorny/paths-filter
if touches('**/watcher*','**/remote*','**/*platform*','.github/**'):
required = [fast_gate, full_matrix_12]
else:
required = [fast_gate] # full 12-job runs on merge_group only
Treat as fact per mod: merge queue already runs same 12-job matrix, 20 PRs/day, 80% PRs touch lib only.
Arithmetic at 20 PRs/day:
Before fix: 41% flaky red * 20 = 8.2 PRs/day hit, 2.1 outer reruns * 18m = 37.8m extra wait. ~310 min author blocked waiting per day just on retries.
After floor + nextest [watcher sentinel-readiness eliminates registration race not durability, semantic state assertion not raw order, mac disk-reclaim kills 14% mac class, graph sort kills 18% order class, remote deadline 2s->10s + bind readiness, no inner loop]: residual first-attempt flake ~5%. With narrow process-level retries, final job red after 3 tries ~0.5-1%.
Full-matrix required for all: author latency = 18m per push * ~1.5 pushes = 27m/PR. 20 PRs = 540 min/day idle. Flake block 0.2 PRs/day. Zero late-blindness.
Path-routed C-smart: 16 PRs fast path @ 4m = 64m, 4 sensitive PRs @ 18m = 72m. Avg wait = 6.8m vs 18m. Saves ~11m/PR, ~224 min/day at merge-queue entry, ~336 min/day vs full-push-average. Compute: merge queue still needs 201218 = 4320 compute-min/day - same as before, fits in day even serial = 360 min wall, with 5 concurrent groups = ~72 min. Late-flake queue failures: 20*1% = 0.2/day ~10m context-switch total. Late-real-bug blindness: 0, because sensitive 20% forced to full pre-merge. Pure fast-gate for all would have added 4 PRs * 5% real bug rate = 0.2/day late real failures at 60m each.
So: keep full signal for the 20% that need it, buy 3.7h/day of author time back for the 80%, keep queue failure load <1/day.
Unified 15h budget - reconciled to floor:
old absent/new present/graph invalidated, accept Both/remove+create equivalents, dump timestamped events on failsudo rm -rf /usr/share/dotnet /opt/ghc; TMPDIR=$GITHUB_WORKSPACE/.tmp mkdir -p $TMPDIR chmod 700 + graph snapshot canonical sortFinal retry-transparency rule - dashboard verbatim:
FIRST_ATTEMPT_PASS_RATE (pre-retry) >=99% per test over rolling 200 runs; no test may remain on retry allowlist >14 days without owner and linked issue; any PASS_AFTER_RETRY status is FLAKY, never GREEN, requires issue link, does not count toward first-attempt metric; alert if overall PR first-attempt green <95% or any allowlisted test w/w failure increase.
Surviving positions - one line:
GO for Friday with path-aware routing only if the watcher rewrite lands; otherwise ship the fixes and observability, keep routing off, and preserve full-matrix pre-merge.
The watcher is the largest known flake class, so activating routing first would knowingly move inherited failures onto unrelated PRs; the opposing latency argument is real, but it relies on unmeasured queue and re-engagement costs. Friday’s degraded-mode invariant is: full matrix remains required, first-attempt results stay visible, retried passes remain YELLOW, every allowlisted test has an owner and issue, and expiry is enforced.
A retried pass is YELLOW, never GREEN; first-attempt pass rate must be ≥99% per test over a rolling 200 executions; flaky passes create or link an owned issue, expire after 14 days, and never improve the first-attempt metric; alert below 95% overall first-attempt PR green or on week-over-week allowlisted-test deterioration.
The 45-minute re-engagement cost is an unmeasured bounded estimate, queue wall time depends on unknown concurrency and batching, and “zero late blindness” is false: routing reduces and measures exposure but cannot eliminate classifier, dependency, or retry-masking failures.
The classifier must fail closed for feature, build, generated-code, manifest, harness, rename, and shared-dependency changes. Once fast-path PRs skip flaky suites, the old 41% PR-red rate is no longer comparable; judge fixes through per-test first-attempt results from the full merge-group matrix and scheduled stress runs, not aggregate PR redness.
The Friday package is the watcher sentinel/semantic fix, narrow process-level nextest retries, remote and mac fixes, graph canonicalization, and routing only when classification defaults to full on ambiguity or failure. If first-attempt reporting is absent, ship no retries; if the watcher misses the cut, preserve full pre-merge coverage and the dissent that routing is an optimization, not a substitute for fixing the dominant flake.
I hold the Thursday cut line: if the watcher rewrite overruns, ship the deterministic fixes, narrow nextest/JUnit plumbing, graph canonicalization, and macOS hygiene—but do not activate path-aware routing yet.
The reason is operational, not aesthetic. The watcher accounts for 32% of the sampled flaky reds, the largest single class. Routing before fixing it would remove that test from the PR gate for most changes while continuing to run it in the merge queue. Process-level retries would often prevent a hard queue failure, but the queue would still accumulate YELLOW executions and occasionally block. The dashboard would expose that debt, but exposure is not remediation.
Routing is a large latency win, but it is reversible optimization; the watcher rewrite improves the underlying signal wherever the test runs. Moving a known dominant flake later before repairing it would make queue behavior harder to interpret and could cause developers to blame the new routing for failures actually inherited from the test.
Friday degraded-mode invariant: the full matrix remains required pre-merge for every PR until watcher_test::test_rename demonstrates sentinel-observed readiness, semantic-state assertions, and at least 99% first-attempt success. Retries may ship only with YELLOW reporting and first-attempt telemetry. Routing may ship disabled behind one workflow flag, ready to activate after that threshold is demonstrated.
This differs from @Grok’s Thursday cut. I agree that the rule would force ownership, but an automatically owned problem is still a problem we knowingly moved into the queue.
A PASS_AFTER_RETRY is YELLOW, never GREEN, and never improves the first-attempt metric. Every test must maintain at least 99% first-attempt passes over its rolling 200 executions. The first PASS_AFTER_RETRY automatically creates or links an issue with an assigned owner; every retry-allowlist entry expires after 14 days, after which the test becomes blocking unless an explicit time-bounded exception is approved. Alert if overall first-attempt PR green falls below 95%, or if any allowlisted test’s first-attempt failure rate rises week over week. The dashboard must show first-attempt pass rate, YELLOW count, issue, owner, allowlist age, and post-merge regressions attributed to retried tests.
Signed: @GPT, @Muse, @Grok.
The 45-minute re-engagement figure was an illustrative sensitivity parameter, not observed repo data. It must not appear in the implementation manifest as a measured cost. The defensible statement is: at residual rate (r), pure C creates approximately (20r) late red merge candidates per day before accounting for batching and retries; the human cost per event remains unmeasured.
Likewise, the earlier queue wall-time estimates assumed concurrency and matrix scheduling that were not stipulated. We can calculate compute demand from measured job durations, but cannot infer queue wall time without actual runner concurrency, batching behavior, and arrival/service-time telemetry. Instrument merge-queue wait, execution, retry, and re-entry separately.
“Zero late blindness” is wrong. Path routing reduces exposure; it cannot eliminate it. Shared dependencies, classifier omissions, cross-platform interactions, and nextest retries can all delay or probabilistically conceal a real intermittent regression. The appropriate control is a measured fast-pass-to-queue-fail rate with conservative fail-closed routing, not a claim of zero risk.
The less obvious classifier hazards are transitive rather than lexical:
Cargo.toml, Cargo.lock, feature flags, build.rs, proc macros, generated code, shared test utilities, and serialization/schema changes may affect watcher or cache behavior without touching their directories.merge_group has a synthetic commit and different event payload; a classifier written only for pull_request can silently produce an empty path set..github/** must therefore force the full matrix, and an absent/failed classification result must default to full.The classifier should therefore use a conservative manifest of sensitive paths plus transitive infrastructure paths, evaluate the actual event commit, and fail closed. An explicit ci-full label is an override only toward more testing, never less.
Routing also changes the denominator. Once 80% of PRs stop running the full matrix pre-merge, the old 41% “PR has a flaky red” rate is no longer directly comparable with the fast-gate red rate. It becomes selection-biased: sensitive PRs run more tests pre-merge, while other full-matrix observations occur only at merge-group time.
Therefore, judge the floor fixes using:
Do not claim “41% became 10%” unless both figures use the same execution stage and denominator.
My move from full-matrix-for-all to path-aware routing was an argument concession, driven by the stipulated 80/20 coverage split and the already-required merge queue—not by the clock.
My acceptance of nextest was conditional from the start, but narrowing it to an explicit allowlist and making retried passes YELLOW was an argument refinement based on the real-bug masking risk.
Reducing the watcher allocation to six hours and fitting routing into the plan was partly a clock concession. With no Friday freeze, I would take back that compression: finish and stress the watcher/cache synchronization, gather first-attempt data, and only then activate routing.
I also retract any implication that sensitive-path routing eliminates late blindness. It bounds and measures the risk; it does not remove it.
Reaction-driven shifts: two substantive shifts—full-matrix-for-all to path-aware routing, and generic retry-green to explicit YELLOW with expiry. I did not shift on the Thursday cut after reconsideration.
GO for Friday, with the 15-hour allocation already converged: watcher 6h, remote cache 2h, nextest/JUnit 2h, graph canonicalization 1h, macOS hygiene 0.5h, routing 2h, and validation/stress/dashboard 1.5h.
Refuse to ship if:
fsync for observed readiness;If the watcher misses Thursday, ship the safe half and leave routing disabled. My preserved dissent is that using ownership and a 99% rule to justify knowingly moving the dominant unfixed flake into the merge queue optimizes the visible author loop before stabilizing the system underneath it.
The cut-line dispute is settled in favor of shipping path-aware routing first. If the watcher rewrite overruns on Thursday, we still land the routing, nextest allowlist, remote deadline + readiness, mac reclaim, graph canonicalization, and the full dashboard rule. The still-flaky watcher_test::test_rename stays on the narrow allowlist with its automatic issue; the 99 % first-attempt rule and 14-day expiry force ownership without letting the 32 % class re-contaminate every author’s 18-minute loop.
@GPT, deferring routing would re-impose the full latency tax on the 80 % for another week while we polish a single test; that inverts the arithmetic that already showed ~224 author-minutes recovered per day. @Muse’s path-routed form is the correct vehicle, but the cut priority must protect the DX win.
Degraded-mode invariant (Friday if only half the plan shipped): first-attempt metrics remain visible and pure for every test; the allowlist is live with automatic issue creation and 14-day expiry; residual first-attempt red rate on any full-matrix run (PR-sensitive or merge_queue) stays under 15 %; outer whole-matrix author reruns stay under 0.5 per day; and every allowlisted test has an owner. If those four statements are false, we revert the routing change and keep the old gate until the invariant holds.
Dashboard rule (final, verbatim, signed by the room):
A retried pass is YELLOW, never GREEN. First-attempt pass rate (pre-retry) must be ≥99 % per test over a rolling 200 executions. No test may remain on the nextest retry allowlist longer than 14 days without an assigned owner and an automatically created or linked issue; every flaky-pass automatically creates or links that issue and never counts toward the first-attempt metric. Alert if overall first-attempt PR green falls below 95 % or any allowlisted test’s first-attempt failure rate rises week-over-week. The dashboard reports first-attempt rate, flaky-pass count, owner, allowlist age, and post-merge regressions attributed to retried tests.
(The three of us lock this text; no further edits.)
Epistemics closed. The 45-minute re-engagement figure is a bounded estimate (typical 20–60 min for a CI late failure, drawn from common engineering experience, not a measured distribution on this repo); I treat it only as an order-of-magnitude cost, not a precise input. Queue wall-time and concurrency assumptions (serialized 6 h upper bound, or ~5 concurrent groups) are likewise estimates; we lack the repo’s exact merge-queue parallelism numbers, so the claim is only that 20 full matrices fit comfortably inside a workday under GitHub’s normal capacity—not a hard latency forecast. “Zero late-real-bug blindness” is restated as near-zero for the high-signal suites by construction of the path filter, with residual late real failures estimated <0.5/day once the floor fixes land; absolute zero is not claimed and was never supportable.
Attack on our own convergence. The path classifier is new surface area. Beyond the obvious (misclassification of a shared utility, broad .github/** self-edits that force full, or an incomplete glob list), the non-obvious failures are:
cfg changes that compile watcher/remote paths without any file under the watched globs touching the diff;Any of those silently moves a high-signal change into the fast gate.
Worse for measurement: the original 41 % flake rate was observed under a uniform full-matrix gate. After path-aware routing the aggregate PR-red percentage will drop mechanically because 80 % of PRs no longer execute the flaky suites. That number therefore becomes useless for judging whether the sentinel rewrite, remote deadline, or disk reclaim actually worked. We can only trust the per-test first-attempt rates (especially on the allowlist) collected from every full-matrix execution—i.e., the sensitive PRs plus every merge_queue run. If the dashboard does not already surface those per-test numbers on day one, the routing change blinds us to our own progress.
Concession audit.
If Friday vanished I would take back the precise residual-flake percentages (5–10 %); they are credible extrapolations from the mined 88 % coverage of categories but remain unverified until the first post-change week. Everything else I would keep.
Verdict: GO for the Friday freeze.
Refuse-to-ship conditions: (1) the dashboard rule above is not live and collecting first-attempt data, (2) the nextest allowlist lacks automatic issue creation or 14-day expiry, (3) the path classifier does not default to full on ambiguity or filter failure.
Dissent to preserve: pure full-matrix-pre-merge remains a live fallback if the classifier’s false-negative rate exceeds 5 % in the first ten days; we keep the ci-full label and the ability to flip the default without a new design debate.
The package that ships is therefore path-aware C + corrected B (sentinel) + narrow D, under the single dashboard rule, with the degraded invariant as the safety cord. That is the highest-confidence plan that still fits the clock and the constraints.
Adopting floor, settling dispute as @Muse.
Dispute - Thursday cut line: defer routing, not watcher work.
@GPT is right here and I was wrong to keep pure fast-gate as primary. Pure C does not eliminate failures, it relocates them. If watcher rewrite overruns, shipping path routing exports the still-~32% watcher flake class to the merge queue for 80% of PRs whose diff never touched watcher.
Arithmetic after floor without watcher fix:
merge_group due to an infra flake unrelated to their change.Deferring routing keeps failure early where nextest process retries can absorb it inside the same job and the author is still in context. Routing's value depends on floor <10% residual. If biggest class not fixed, premise fails.
Degraded-mode invariant Friday if half ships (our chosen degraded):
Routing OFF, full 12-job matrix remains required pre-merge for all. Still ships: nextest narrow allowlist with per-attempt JUnit archival, mac disk reclaim + workspace TMPDIR, graph canonical sort, remote 2s->10s + bind readiness, and dashboard rule. Must be true:
merge_group - not PR fast gate - is the measurement ground truthIf routing does ship because watcher lands, degraded for other direction still requires invariant 1-4.
Ratified dashboard rule - single text, signed by @GPT @Grok @Muse:
A retried pass is YELLOW, never GREEN and never improves the first-attempt metric. Every test on the nextest retry allowlist must maintain ≥99% first-attempt passes over its rolling 200 executions, have an assigned owner and linked issue created automatically on first flaky-pass, and expire after 14 days. Breaching ≥99% or expiry without a merged fix makes flaky-pass blocking until an explicit time-bounded exception is approved. Dashboard must report first-attempt pass rate, flaky-pass count, issue owner, allowlist age, and post-merge regressions attributed to retried tests. Alert if overall first-attempt PR green falls below 95% or any allowlisted test's first-attempt failure rate rises week-over-week.
Epistemic concessions:
Attack own convergence:
Path classifier is new failure surface beyond obvious misclassification / shared deps / .github/** self-edits:
paths-filter glob misses rename from watcher/ to watch/ , downgrades to fast.--skip for fast gate: cargo nextest --lib --skip watcher --skip remote hides integration test that would catch breaking change in lib via integration harness. Fast green lies.merge_group first-attempt rate only. If you don't, routing makes floor look successful while queue churns.Concession audit:
Shifts due to peer reactions: 3. Clock vs argument: all argument/fact-driven, not clock.
If Friday vanished I'd take back my R1 blind push for C as primary before pricing queue cost for sensitive PRs.
Verdict: GO for Friday freeze with conditions.
Unified 15h: 6h watcher sentinel+semantic, 2h nextest+JUnit+allowlist, 1.5h remote deadline+readiness, 1h disk reclaim, 1.5h graph sort, 1.5h path-aware routing + aggregator with fallback to full, 1.5h wiring dashboards.
Refuse-to-ship if: per-attempt logs not archived, YELLOW check not separate, or allowlist lacks owner/issue/expiry automation.
Dissent to preserve: if we cannot make first-attempt reporting separate by Friday, ship no retries at all until we can - retries without observability normalize nondeterminism.