After rejecting stream revocations as safety, the panel debated broker epochs versus consumer high-water marks for a tight Friday deploy.
“There is no atomic transaction spanning the Postgres lease CAS and Kafka producer initialization.”
With shard topics restricted to an internal materializer and one analytics batch, treating Kafka as an untrusted ingress log provides state safety. But high-water checks leave stale records on disk and require strict ACLs; all three models would adopt broker epochs if not for the Friday cutoff.
No-go for Friday as written: ship only after the sink atomically rejects stale generations, old clients cannot acquire affected leases, and acquire/renew semantics are proven under concurrency and failover.
A frozen worker, Redis partition, lagging subscriber, missed publication, or incorrect consumer-group topology can delay or omit the abort, so the 120 ms figure applies only to healthy subscribers. Use an outbox if you need reliable notification, but never treat the stream as safety.
The worker can pass a live check, freeze, lose the lease, and later flush the already-checked record. A signed token proves authenticity only; safety requires the broker, gateway, or downstream resource to enforce a current/high-water generation. Equal-token retries need separate idempotency handling rather than blanket rejection.
Old workers ignore both the stream and new producer enforcement, so they can reproduce Tuesday’s split brain during rollout. Enforce a capability/version gate or drain old workers before affected leases can be reclaimed; otherwise added complexity creates false confidence without removing the original failure.
Use authoritative Postgres time for the CAS and renew, never renew an expired lease, discard every losing token, and audit sequence caching because cached values can issue a lower token after a higher one. Forced aborts will end legitimate long sections unless work is checkpointed, restartable, and side effects are fenced at each durable boundary.
No-go on the proposal as currently worded. The revocation stream is a useful acceleration mechanism, but the stated “correctness boundary” still contains a fatal TOCTOU hole: checking the live lock service immediately before a Kafka produce is not fencing. The lease can expire and token 403 can be granted after middleware validates token 402 but before the broker accepts 402’s record. Likewise, a signed token proves authenticity, not freshness or current ownership.
I would approve Friday only if the actual protected resource—or a gateway that exclusively controls it—atomically enforces fencing, and old clients are prevented from acquiring the affected leases.
For databases, fencing can usually be placed in the mutation itself: update the resource and its remembered fencing generation transactionally, rejecting a token below the resource’s high-water mark or different from the current generation. A separate GET /lease, followed by a database write, is not equivalent because ownership can change between those operations.
Kafka is harder. Merely putting fencing_token=402 in a header leaves Kafka accepting the record. Kafka does not natively compare an arbitrary lease token against your Postgres lease row. You need one of these real enforcement designs:
transactional.id derived from the lease name, with takeover initializing a new producer epoch before the new owner writes. Kafka then broker-fences the prior producer session rather than trusting application middleware. This needs careful partitioning, transaction, retry, and throughput validation.Do not describe “signed token embedded in the produce path” as sufficient unless some downstream enforcement point maintains current/highest generation and rejects stale tokens. Also, middleware must not cache successful live checks beyond a rigorously defined grant, because caching recreates the same expiry window. Even an uncached check still has the check-to-produce race.
Every failure you named can make revocation arbitrarily late or absent:
There is also a Redis Streams topology trap: consumers in one consumer group load-balance entries rather than broadcasting every entry to every worker. If all 400 sidecars share a group, only one normally gets a given revocation. Independent cursors or per-worker groups provide fan-out but create cursor, retention, pending-entry, and inactive-consumer management burdens.
If you care about reliable best-effort notification, insert a transactional outbox record in the same Postgres transaction as takeover and relay it to Redis. Direct “commit, then publish” necessarily has a crash window in which the event is permanently missing. Even with an outbox, notification remains UX, not safety.
A delayed or duplicated revocation must match (name, old_owner_id, old_token) exactly. Do not revoke by name or owner alone. Otherwise a delayed revocation for token 402 can kill the same worker’s later legitimate token 404 lease.
The “p99 ~120 ms” claim should be scoped to something like “database takeover commit to callback in a healthy, subscribed, non-paused process.” It says nothing useful about the pathological processes fencing exists to handle. The safety-relevant tail is unbounded under partition or process suspension.
“Renew fails on token mismatch” is insufficient. Renewal must atomically require all of:
WHERE name = $name
AND owner_id = $owner
AND fencing_token = $token
AND expires_at > clock_timestamp()
and extend expiry from database time. If an expired-but-not-yet-superseded lease can be renewed solely because its token still matches, the old owner can resurrect ownership after the contractual TTL boundary. Decide explicitly whether equality at the boundary is valid; consistency matters more than choosing < versus <=.
The acquire CAS is sound only if each attempted token is permanently discarded whenever that specific CAS loses. A safe shape is one atomic statement/transaction that obtains nextval, performs INSERT … ON CONFLICT … DO UPDATE … WHERE expires_at < now(), and returns ownership only when a row was actually inserted or updated. Postgres sequence values are not rolled back, so gaps are harmless for fencing. But never retain an allocated token and reuse it on a later retry.
Check the sequence configuration:
BIGINT, NO CYCLECACHE 1setval, restart, or alter itCACHE 1 matters more than it may appear. With cached blocks, session A can retain lower unused values while session B obtains and successfully uses higher values; A can later issue a lower token for a later acquisition. Uniqueness is not enough—successful acquisitions of the same resource must increase in takeover order.
Also scrutinize database disaster recovery. If an asynchronously replicated Postgres standby can be promoted after externally visible token allocations or fenced writes are lost from the database history, the system can reuse or regress generations relative to Kafka or another external sink. You need synchronous durability appropriate to the guarantee, or a failover-incarnation epoch incorporated into tokens and enforced downstream.
Release must predicate atomically on name, owner, and token. Deleting the row is not itself unsafe with a truly monotonic external allocator, although retaining a tombstone/high-water row can make auditing and per-resource fencing easier.
Yes, it will terminate legitimate long critical sections during control-plane failures. That is not a bug; it is the availability cost of choosing safety. A worker that cannot prove its lease remains current must stop initiating side effects before its known lease deadline. “Abort on mismatch” is too late: timeouts, inability to renew, connection-pool stalls, and missed renewal deadlines must all close the gate.
Long-running work therefore needs:
Local cancellation alone cannot make an arbitrary critical section safe: it may leave partial DB state, open Kafka transactions, or an external request whose outcome is unknown. Takeover and retry behavior must be designed around those ambiguous outcomes.
During any period when old clients can acquire these leases and write through unfenced paths, the original production failure remains possible. Deploy order must be:
If the lock API cannot identify client capability strongly enough to deny old acquisition, drain old workers before enabling takeover—not a normal mixed rolling deploy. A user-agent string is not a safety control.
The dangerous outcome is declaring victory because sidecars make normal incidents look fast while the actual Kafka path remains check-then-write. That would convert a known best-effort system into one operators incorrectly believe is fenced, encouraging shorter TTLs and more aggressive takeover while stale writes remain possible.
Other regressions include making every write synchronously dependent on the lock service, causing a lock-service or database blip to halt all stateful workers and potentially trigger retry storms; aborting valid new leases because of replayed revocations; using a shared Redis consumer group and silently delivering each revocation to only one sidecar; and repeatedly killing long jobs that lack checkpoints. Aggressive TTL reclaim combined with non-checkpointable work can increase duplicate attempts and partially completed side effects even if local dual-writer duration improves.
My deploy gate is simple: no-go Friday if “live confirmation or signed token” remains the Kafka safety mechanism. Go only if you demonstrate an atomic broker/gateway/consumer enforcement point, expired leases cannot renew, the sequence is monotonic across concurrency and failover, and old clients are denied acquisition before mixed-mode operation. The Redis revocation work can ship independently as an explicitly non-authoritative optimization.
The dual-writer window is the original bug, and the proposed fix still leaves it load-bearing on incomplete client discipline plus a best-effort notification path. That is the core problem. Stream revocation plus sidecar cannot be the safety mechanism; it can only be a latency optimization for processes that are still running, still subscribed, and still able to observe the message. Anything that freezes the worker (the exact GC thrash that already happened), partitions it from Redis, or leaves the subscriber lagging or not yet started keeps the stale owner writing until the mandatory fencing path actually rejects it. Redis streams have no delivery guarantee that bounds the dual-writer interval under the failure modes that already occurred; lag, AOF/RDB recovery, multi-region cross-AZ partitions, and subscriber restart all produce windows measured in seconds to minutes, not a reliable p99 of 120 ms. Treating the stream as collapsing the window is therefore only true for the well-behaved subset of processes that never needed the fix in the first place. Fencing must stand alone: every produce and every DB write for the non-compliant classes has to reject on a stale or missing token without consulting the stream at all. Live lock-service confirmation on every write is one way to do that, but it turns the lock service into a synchronous dependency for the entire data plane; if the lock API is slow or partitioned, writers fail closed and you have turned a correctness bug into an availability outage. A signed fencing token that the Kafka middleware and DB helpers can validate locally (or, better, a token that the protected resource itself remembers as a high-water mark and rejects below) is safer; the classic fencing pattern does not require a round-trip on every write once the resource has seen a higher token.
The acquire path has a quieter but real hole around the sequence and the CAS. Allocating nextval('lease_token_seq') before the compare-and-set is fine for uniqueness and monotonicity—Postgres sequences never roll back and never reuse—so you get gaps on failed CAS, which is harmless. The CAS itself (expires_at < now() or missing name) is subject to clock skew between the lock service nodes and between the service and clients; a node whose clock is ahead can reclaim a still-valid lease, hand out a higher token, and leave the previous owner with a token that is already stale from the new owner’s perspective. If two reclaimers race, the later nextval wins the higher token and the earlier one loses the CAS, which is correct, but only if the update is a true single-row atomic conditional write that also stamps the new owner and new token together. Any read-modify-write that is not a single UPDATE \ldots WHERE expires_at < now() RETURNING (or equivalent) re-introduces lost updates. Token reuse is not a sequence problem, but a reclaim-after-delete or a “soft release that zeroes the token” problem; keep the row and only ever advance the token, never recycle. The same-owner renew path must also be fail-closed on token mismatch, which you already state; that is necessary but not sufficient if the data-plane path still accepts the old token.
Fail-closed renew plus forced abort does kill legitimate long critical sections that outlast a single TTL without intermediate checkpoints. If a worker holds the lease for a multi-minute indexing or compaction pass, a single missed renew (GC, network blip to the lock service, or a brief Redis hiccup that is unrelated to ownership) will revoke it, the sidecar will abort, and the work is lost unless the critical section is written to be restartable from a durable checkpoint that itself carries the fencing token. You therefore need an explicit contract: every long section must checkpoint under the current token, re-validate ownership before the next side-effect, and treat abort as a normal control-flow path rather than a crash. Without that, the “fix” simply trades split-brain writes for systematic incomplete work and thundering-herd re-acquisition. Auto-renew in a sidecar helps only while the process is healthy; the moment it freezes past the renew interval the lease is gone and the abort fires (or, worse, does not fire because the sidecar froze too).
Rolling deploy of ~400 workers over six hours is a prolonged mixed-mode window in which old clients still perform the exact Kafka writes that caused the outage. New clients may enforce middleware, old ones will not; a reclaim that happens while both generations are live still produces dual writers for as long as the old process remains thawed and unaware. The stream revoke message is useless to an old binary that does not subscribe, and the lock service itself cannot reach into the old process’s producer to stop it. During that window the system is strictly more complex than today (new code paths, Redis dependency, sidecar lifecycle) while remaining equally open to the original failure mode. Any partial rollout that lets operators believe “fencing is now mandatory” when half the fleet still ignores it is worse than the status quo, because it creates a false sense of safety and delays the real hardening.
What makes the whole package worse than today’s “hope they check fencing tokens” is any design that lets the stream or the sidecar become load-bearing in people’s mental model, or that adds failure modes that encourage bypasses. If the live lock check is expensive, teams will add “temporary” escape hatches. If the sidecar is crash-only and shares fate with the worker, freezes still produce silent dual writers. If revocations are emitted on every supersede but the consumer is best-effort, on-call will start treating missing aborts as “Redis being flaky again” instead of as a fencing violation. The net result is more moving parts, a new Redis operational surface, and the same correctness hole until every produce path and every DB helper for the three non-compliant classes is proven to reject stale tokens under partition, under load, and under partial deploy.
Highest-risk failure modes in order: (1) incomplete or bypassable data-plane fencing, especially on Kafka produce and any async/batch write path; (2) mixed old/new clients during the six-hour rollout; (3) process freeze or Redis lag that leaves the dual-writer window open far beyond the aspirational 120 ms; (4) acquire CAS races or clock-skew premature reclaim that hands out a new token while the previous owner is still live; (5) long critical sections that cannot survive forced abort without checkpoints. The stream and sidecar are fine as UX, but they must be documented and tested as non-safety. Before Friday I would not ship this as-is. Make the fencing enforcement complete and mandatory first (ideally with resource-side high-water marks rather than a live lock round-trip on every write), prove the acquire path under concurrent reclaim and skew, require checkpoint/restart contracts for long sections, and stage the rollout so that old binaries are either upgraded or explicitly fenced out of the dangerous produce paths before any reclaim can create a mixed dual-writer. Otherwise the next regional blip will look exactly like Tuesday, only with more log lines about revocations that arrived too late.
This fix has 2 systems: a best-effort UX optimization and a safety mechanism. You must ship them with opposite priorities.
1. Where stream revocation fails - and it will:
"collapses dual-writer window to p99 ~120ms" is false for the exact bug you hit. A GC-frozen process cannot process a Redis message. Your 11 minute dual-writer is the p99 case.
Failure modes that MUST be assumed:
indexer-17 freezes, misses publish, unfreezes after indexer-04 already wrote. No abort. Redis Streams do not buffer indefinitely for consumers that weren't connected; lagging / trimmed consumer = silent drop.$ you skip history; if from 0 you abort on stale revocations after restart.COMMIT lease row in PG, then publish to Redis. If process crashes between those two, new owner holds token 403 and no revoke is ever emitted. You need transactional outbox or publish-then-verify loop with background reconciler. Without it, revoke is best-effort.Conclusion: Revocation can never be safety. It improves abort latency for well-behaved processes only.
2. Mandatory fencing is the only safety - but you changed its fault domain
Proposed: Kafka producer middleware refuse writes unless live lock service confirms (name, owner, token) or signed token embedded.
This is correct direction, but:
fencing_token in the record header and have a broker-side or downstream consumer check that rejects old_token < max_seen. If you only check in middleware, then indexer-17 check passes, freezes, lease expires, 04 acquires, 17 unfreezes and flushes its already-checked buffer. Fencing must be validated at the data-plane sink, not the producer memory.3. Acquire CAS + lease_token_seq — there IS a hole in your sketch
nextval('lease_token_seq') is strictly monotonic and never rolls back even if txn aborts. Gaps are fine, reuse is not.
Two bugs to kill before Friday:
a) CACHE breaking monotonicity per key: If sequence is defined with CACHE >1 default maybe 50, each PG backend preallocates a block. Session A gets [1..50], B gets [51..100]. A acquires with 1, B supersedes with 51, then A retries and uses cached 2 which is < 51. You just issued a lower token after a higher one for same name. Fencing collapses. You must have CACHE 1 per the fencepost notes, or allocate with nextval inside the critical section ordered by row lock.
b) Lost-update / token burn:
Wrong pattern:
token = nextval()
UPDATE leases SET owner=.., token=token WHERE name=$ AND expires_at < now()
If two workers race, both call nextval, both see expired, one UPDATE wins, the other gets 0 rows but burned token 404. Still safe monotonically, but if loser then does INSERT with old token you have race.
Also INSERT ... ON CONFLICT DO UPDATE SET token=nextval() WHERE expires_at < now() — in Postgres, nextval() in DO UPDATE is evaluated even when WHERE fails, burning tokens and potentially violating ordering if you rely on RETURNING to know winner.
Safe pattern — loop in app, single statement winner:
-- attempt reclaim
UPDATE leases
SET owner_id=$2, fencing_token=nextval('lease_token_seq'), expires_at=now()+$3
WHERE name=$1 AND expires_at < now()
RETURNING fencing_token;
-- if 0 rows, try insert
INSERT INTO leases(name, owner_id, fencing_token, expires_at)
VALUES ($1,$2,nextval('lease_token_seq'), now()+$3)
ON CONFLICT (name) DO NOTHING
RETURNING fencing_token;
-- if conflict, retry from top
Must run with now() from DB clock. Use FOR UPDATE or the WHERE expires_at < now() as the CAS. Renew must also be token-gated: WHERE name=$ AND owner_id=$ AND fencing_token=$ AND expires_at > now().
4. Fail-closed renew + forced abort vs long critical sections
Fail-closed renew kills legitimate work if GC > TTL, even without contention. If you set renew interval 10s / TTL 30s, a 35s STW pause = lease loss + abort. That's desired for safety, but your workers must checkpoint.
Requirement: any shard build > renewal interval must be idempotent or resume from checkpoint after abort. If indexer-17's 11 min write was one long non-atomic Kafka transaction, forced abort halfway leaves partial offsets. You need either atomic chunk commits gated by fencing check per chunk, or ability to rollback.
Also bound detection: renew time must include pool checkout time. Hikari connectionTimeout > renew interval means you won't detect loss in time. Keep checkout <= refresh interval.
5. Rolling deploy of 400 workers over 6h — mixed mode is dangerous
During rollout, old workers ignore both revocation and new middleware enforcement. New workers will enforce and abort correctly, old ones remain split-brain writers.
If you flip DB/producer enforcement to mandatory on day 1, old workers start hard-failing writes they previously would have done (good, but looks like outage). If you flip to audit-only first, you have 6h where old+new can still dual-write.
Mitigation: ship middleware in two flags: 1) log on fencing fail, 2) enforce. Run 1) for 1h, then 2). Deploy new workers to non-critical shards first. At the sink (Kafka consumer / shard writer), still enforce max-token check regardless of producer version — that's your backstop.
What makes this worse than today?
name not (name, old_token) → you abort current valid holder after a delayed redelivery.Go / No-Go
No-go as a single bundle for Friday if mandatory fencing at Kafka sink + DB helpers is not live and load-tested.
Conditional Go:
CACHE 1, correct CAS loop, use DB now(). Unit test concurrent supersede — token must always increase.incoming_token <= max_token_for_shard stored in PG/shard metadata. No live PG check per message in hot path if you can verify signed token + monotonic check at sink.If you can't enforce at sink before Friday, delay. Revocation alone re-introduces the Martin Kleppmann failure mode you cited.
Ship sink-side high-water enforcement Friday for materialized state, and keep Kafka broker epochs out of the cut unless Thursday proves the full handover and bypass story; if Kafka’s log itself is authoritative, neither consumer high-water nor unproven epochs is sufficient.
Store per-shard fence state with the materialized data, and atomically apply the business mutation, reject lower tokens, and deduplicate equal-token retries with a stable operation key. This is not Kafka-log fencing: if any correctness-relevant reader can consume the raw log, use an exclusive gateway or delay the rollout.
Thursday must prove sustained throughput and latency at projected load, freeze-and-takeover fencing with zero stale accepted output, partition and coordinator-failover behavior, and that every producer and bypass path is blocked. The Kafka handover boundary is epoch initialization—not the Postgres lease CAS—so a visible post-CAS/pre-epoch commit means the promised guarantee is narrower or the design fails; fallback is the guarded materializer, or delay Kafka enforcement if the log itself must be clean.
I choose broker-native fencing with a stable transactional.id per lease, not the generic high-water-mark proposal. The decisive reason is architectural: Kafka is the protected sink here, and a downstream consumer’s high-water mark does not prevent Kafka from accepting stale records. It can protect a later materialization step, but it cannot make the Kafka log itself free of stale or duplicate output. If the Kafka records and offsets are the externally relevant side effect, that distinction is fatal.
I also agree with the moderator’s skepticism of incoming_token <= max_token. The normal fencing predicate rejects lower generations, not equal ones. Multiple legitimate records and retries from the current lease necessarily carry the same fencing token, so rejecting equality conflates lease fencing with record deduplication. Equal-token retries require a separate idempotency key or monotonic per-shard operation/offset sequence. A high-water design could work only if fence state lived at the actual materializing sink and the fence update, idempotency decision, and business mutation occurred in one transaction. That is not a generic Kafka acceptance mechanism.
A stable transactional ID is not automatically equivalent to lease fencing. Kafka fences an old producer when a new producer using the same transactional.id initializes and obtains a newer producer epoch; Postgres granting token 403 does not itself advance Kafka’s epoch. Therefore this naïve sequence is still unsafe:
The stale transaction in step 2 may already be committed. This is the broker equivalent of the middleware TOCTOU gap.
Accordingly, the shippable contract cannot be merely “initialize the producer after acquiring the lease.” The new lease must not be reported active for Kafka output until the broker epoch has been advanced, and the handover semantics must explicitly define the broker epoch transition as the Kafka ownership linearization point. Even that prevents simultaneous post-handover writers rather than magically making Postgres TTL expiry and Kafka epoch advancement atomic. There is no atomic transaction spanning the Postgres lease CAS and Kafka producer initialization.
That leaves a hard decision:
For Friday, I would accept the first guarantee only if it is documented honestly and the acquiring worker cannot produce generation-403 data before epoch initialization succeeds. I would not claim strict token-based Kafka fencing at the Postgres takeover instant.
The team must demonstrate all of the following against a real staging Kafka cluster, not mocks:
Exact identity mapping. Every lease name maps deterministically to one stable transactional.id; names cannot collide through truncation, hashing, escaping, tenant omission, or environment omission. ACLs must prevent an unrelated principal from initializing that ID and fencing the legitimate owner.
Actual transactional use. All output from these three worker classes goes through initialized transactional producers. No plain producer, emergency producer, admin replay path, direct library construction, or “transactions disabled” feature flag may bypass the wrapper.
Freeze/takeover test. Producer 402 begins a transaction and sends records, is frozen beyond TTL, producer 403 completes handover using the same transactional ID, and 402 then resumes. Its send or commit must fail with fencing, and read_committed consumers must not observe its aborted transaction.
The dangerous gap test. Producer 402 attempts to commit after the Postgres takeover but before 403 initializes. The observed result must be compared with the promised handover semantics. If the transaction is visible—and I expect it can be—the team must not describe Postgres takeover as the Kafka fencing point.
Consumer isolation. Every correctness-relevant consumer uses read_committed. Consumers using read_uncommitted can observe records from transactions later aborted by takeover.
Partition behavior. Test old owner connected to Kafka but partitioned from Postgres, new owner connected to both, and the inverse. Also test coordinator failover, producer restart, transaction timeout, retries after unknown commit outcome, and two simultaneous reclaimers.
Throughput and cardinality. Measure transaction-coordinator load, transaction-init latency, transaction size, commit rate, producer count, connection count, and recovery time at projected peak plus headroom. A producer per actively held shard may be operationally expensive; pooling cannot merge different lease identities without defeating the fencing mapping.
Failure containment. ProducerFencedException must permanently poison that producer instance and abort its worker’s local ownership. It must never trigger transparent producer recreation under the stale lease, fall back to a nontransactional producer, or retry with another transactional ID.
Ordering and idempotency. Records must retain a separate operation ID or shard-offset identity. Broker epoch fencing prevents the old producer session from continuing after takeover; it does not replace application-level handling of retries whose commit result is unknown.
If any bypass remains, or the freeze/takeover test works only because the old producer voluntarily notices revocation, validation has failed.
Inventory and block bypasses. Enumerate every Kafka producer construction and replay path for the three classes. Centralize producer creation and add telemetry for transactional ID, lease name, token, and producer epoch.
Deploy transactional production without lease takeover changes. Run the exact stable-ID mapping and read_committed configuration in staging, then canary it on selected production shards. Do not run two live owners merely to test in production.
Run the Thursday gate. Especially the post-Postgres/pre-Kafka gap test, coordinator-failover tests, and peak-load test.
Prevent old acquisition. Add a capability/version gate at lease acquisition and drain old owners. No old binary may reacquire an affected lease after enforcement begins.
Enable broker-fenced handover. A new owner may begin Kafka work only after successful transactional producer initialization for the stable lease ID. Initialization timeout or ambiguity is fail-closed.
Enable forced abort and checkpoint recovery. A fenced producer terminates the current attempt; checkpoint/retry logic starts a new attempt under a newly acquired lease rather than reconstructing a producer under the stale token.
Ship Redis revocation last and feature-flagged. It remains an abort-latency optimization, not an ownership signal.
Do not add the transactional outbox to Friday’s correctness bundle. Ship revocation fire-and-forget only if its API, dashboards, and runbooks label it best-effort and the feature can be disabled independently. Without an outbox there is an unavoidable commit-to-publish crash gap, so do not publish an SLO implying every takeover produces a revocation.
If reliable operational notification matters, add the outbox next: write the takeover and outbox event in the same Postgres transaction, relay idempotently, and key deduplication by takeover token. An outbox removes permanent publication loss after a committed takeover, but it still cannot bound notification latency during Redis failure, relay failure, partition, trimming, or a frozen subscriber.
I concede the original service/client clock-skew objection. If expires_at is created and compared using the authoritative Postgres clock in the atomic CAS, worker-node and client clocks do not decide expiry. Database clock jumps and asynchronous failover are separate operational concerns, but they are not an argument that ordinary service-node skew breaks this CAS.
I personally refuse Friday’s deployment if any of these is true:
So my recommendation is conditional go for broker epochs, with a deliberately narrower guarantee than strict Postgres-token fencing. If Thursday validation fails, the fallback is not the underspecified high-water consumer check; it is to delay the Kafka enforcement rollout, deploy only the safe acquire/renew hardening and optional UX revocation, and keep old clients barred from any new “fenced” mode until an acceptance gateway or properly transactional sink design is ready.
@GPT correctly isolates the fatal TOCTOU: any check that lives before the broker or DB accepts the write can be invalidated by a freeze, a lease reclaim, and a flush of already-authorized buffers. That is exactly how indexer-17 produced for eleven minutes after token 402 was superseded. @Muse’s high-water sketch is directionally right for the DB helpers, but the moderator’s skepticism is warranted for Kafka: a pure “max_token_for_shard in PG” without an exclusive produce path or consumer-side generation barrier still leaves a window in which records are accepted before the high-water advances, and equal-token retries become ambiguous if the same generation is re-used after a partial failure.
I therefore choose broker-native fencing via a stable Kafka transactional.id derived from the lease name as the primary sink mechanism for the three non-compliant worker classes’ produce paths, with classic high-water-mark enforcement inside the mutation transaction for their DB write helpers. The broker path is the one that actually closes the incident that already happened; the high-water path is the one that is already natural for Postgres.
For Kafka, every producer that can touch a leased shard must use a transactional producer whose transactional.id is a pure function of the lease name (for example lease-shard:8841). On successful acquire the new owner initializes a new producer epoch (by beginning a transaction or calling the fencing init) before any data records are sent. Kafka’s broker then rejects further records from any prior epoch for that transactional identity. This removes the dual-writer window at the acceptance point rather than after the fact.
Tradeoffs that must be accepted: transactional producers carry higher latency and memory overhead than fire-and-forget producers; the transactional.id is sticky to a producer instance and therefore forces a clean hand-off (exactly what we want for mutual exclusion, but it will surface any code path that still creates an ordinary KafkaProducer); partition assignment and exactly-once configuration must be re-validated because a single transactional.id cannot be used concurrently by two processes; and any remaining non-transactional produce path (admin tools, debug scripts, old libraries) becomes a permanent bypass. If the workload cannot tolerate the throughput cost or if any bypass survives, the design collapses back to an application-level high-water.
For DB helpers the high-water lives in the same row (or a tightly coupled generation column) that is mutated: the UPDATE or INSERT carries AND fencing_token = $token or AND $token >= high_water and advances the high-water atomically with the business change. Equal-token retries from the same owner are accepted (they are the natural idempotent retry of a generation); strictly lower tokens are rejected. Multi-row atomicity is obtained by keeping the high-water in the same Postgres transaction as the protected rows; if the write spans multiple databases the same generation token must be checked in every participant or the operation must be made single-row. Fence state therefore lives either in the resource row itself or in a one-row-per-shard generations table that is updated under the same lock that protects the data.
This hybrid is stronger than either pure approach: broker epochs give us true sink rejection for the path that already failed in production, while high-water keeps the DB path simple and free of Kafka-specific machinery.
By Thursday end-of-day the following must be demonstrated on a production-like cluster, not a laptop:
If any of those five fail, the fallback that still ships is an exclusive produce gateway that itself owns the transactional.id and performs the high-water check before forwarding; that gateway becomes the sole acceptance point and can be load-tested independently. Without one of those two (native epoch or gateway) the Kafka path remains open.
I concede the clock-skew objection once and for all: when every CAS, renew, and expiry comparison uses the database’s clock_timestamp() (or now()) and expires_at is written from that same clock, client and lock-service node wall-clock skew become irrelevant to ownership decisions. The only remaining skew risk is an asynchronously replicated Postgres standby that is promoted after externally visible tokens or fenced writes have already been observed by Kafka; that is a durability configuration problem, not an acquire-CAS problem, and must be addressed by synchronous commit or an incarnation epoch in the token itself.
On the outbox: Friday’s cut ships the revocation stream fire-and-forget after the Postgres commit, with its best-effort status written into the runbook and the sidecar code. A transactional outbox (same transaction as the lease row, relayed by a background process) is the correct long-term shape and removes the permanent “commit-then-crash” gap, but it is not required for safety and adds migration and relay operational surface that we should not force into this window. Document the gap; do not pretend the stream is reliable.
BIGINT NO CYCLE CACHE 1, nextval performed inside the single atomic reclaim/insert statement that also stamps owner and expires_at from database time, renew predicated on name + owner + token + unexpired in one statement. Concurrent supersede tests must show strictly increasing tokens.(name, old_owner_id, old_token), independent cursors (no shared consumer group), and an explicit “this is UX only” flag that can disable them without touching safety.I will not sign off on Friday if any of the following are true:
CACHE 1 or token allocation can still issue a lower value after a higher value has been successfully used for the same name.@GPT’s insistence that the protected resource or an exclusive gateway must enforce, and @Muse’s observation that a GC-frozen process cannot act on a Redis message, both survived round 1 untouched because they are the actual failure mode that already occurred; everything else is implementation detail around them. The stream and sidecar remain useful for healthy processes; they must never be allowed to become the story operators tell themselves when the next regional blip freezes a worker past TTL.
For Friday, I ship sink-side high-water mark now. Broker epoch is stronger long term for Kafka but not shippable by EOD Thursday.
Enforcement must live where writes are accepted, not where they are sent. Middleware check-then-produce has TOCTOU. The frozen process that caused Tuesday passes the live check, freezes, then flushes after lease loss.
For the 3 non-compliant classes:
1. Fence state location: shard_fences(shard_id PK, cur_token BIGINT NOT NULL, cur_owner_id, updated_at) + applied_writes(shard_id, token, dedup_key, PK) Separate from leases table so fence survives lease deletion and is auditable. Storing fence in same transaction as lease row creates coupling between lease lifecycle and data-plane safety - keep it independent.
2. Atomicity story:
BEGIN;
SELECT cur_token FROM shard_fences WHERE shard_id=$ FOR UPDATE;
-- reject if $incoming_token < cur_token
-- if $incoming_token == cur_token, check applied_writes dedup
-- else if $incoming_token > cur_token: do business writes
INSERT INTO shard_fences ... ON CONFLICT DO UPDATE SET cur_token=$token WHERE cur_token < $token
INSERT INTO applied_writes...
COMMIT;
Fence advance and side-effect commit atomically. If fence update is not in same txn as data write, crash between them leaves gap where old token can re-win.
incoming_token < cur_token. That means duplicate offsets still exist in log (disk cost) but downstream state is safe. Document that. If you need log itself clean, you need gateway or transactional.id later.This turns lock service into async control plane, not sync per-write dependency. Live per-write PG check makes lock service availability = data plane availability, which turns correctness bug into outage.
3. Equal-token / retry objection:
Rule is not <= reject. It's:
< cur_token => hard reject stale generation== cur_token => allow only if dedup_key not in applied_writes (idempotent retry from still-valid owner)> cur_token => takeover, advance water markWithout dedup table, a legitimate retry from current owner after crash looks identical to stale replay. You need producer to send stable dedup_key - e.g., input offset batch id - in header.
Tradeoffs vs broker-native transactional.id per lease:
transactional.id. With 1000s of shards, transaction coordinator becomes bottleneck, txn timeout tuning vs lease TTL is non-trivial, and zombie with open txn blocks new owner until transaction.timeout.ms expires.I pick high-water now as Friday safety. Transactional.id as follow-up hardening if Thursday validation passes.
If you want to swap to broker fencing, demonstrate by Thursday EOD:
transactional.id = shard:{id}, sustained produce while injecting takeover every 5s: p99 fencing latency < lease TTL, no InvalidProducerEpoch leak.If any fails, fallback is high-water consumer gate above. Do not ship transactional.id in audit-only mode - that's false safety.
Concede the rebuttal is correct if you enforce DB time authoritatively. Using WHERE expires_at < clock_timestamp() with PG clock_timestamp() or now() inside txn eliminates cross-service-node skew. If any node uses app System.currentTimeMillis(), ahead clock can prematurely reclaim. Fix: all expiry comparisons use DB time, and fencing token allocation uses nextval('lease_token_seq' CACHE 1) inside the same atomic UPDATE ... WHERE expires_at < now() RETURNING. Remaining risk is PG host clock jump backward - mitigate with NTP monitoring, not logic.
Sequence config is go/no-go: CACHE 1, NO CYCLE, BIGINT. With CACHE>1, session holds lower values while another commits higher, allowing later acquisition with lower token than already-accepted max. Gaps from failed CAS are fine.
Fire-and-forget for Friday. Revocation stream is UX-only, never safety - frozen GC victim that caused 11-min dual-writer will never process it anyway. Commit-then-publish to Redis has crash window where revoked lease has no revoke event. That is acceptable because safety is in sink. Document stream as best-effort with on failure -> lag/trim/cursor lost = silence. Outbox relay in same PG txn as takeover comes next cycle, not blocking safety.
UPDATE ... WHERE expires_at < now() AND (owner+token match for renew) with DB time, CACHE 1 sequence, fail-closed renew requires owner_id=$ AND fencing_token=$ AND expires_at > now().shard_fences + applied_writes, ship sink enforcement in shadow/log mode, alert on violation.{token, owner, dedup_key} and write fence transactionally.POST /v1/leases for old binaries on protected shard:* names (require capability header).lease.revoked fire-and-forget sidecar for fast abort.token < cur_token under concurrent reclaim test.expires_at <= now() lease to renew on token match alone.CACHE !=1 or ability to setval.Position note: my stance on UX vs safety is unchanged from round 1 - revocation cannot be in safety proof.
GO for Friday only with authenticated generation barriers, atomic high-water-plus-dedup at both readers, and every Thursday gate passing; otherwise NO-GO and ship only lease acquire/renew hardening plus best-effort revocations.
The materializer and analytics job must be the only readers, and both must process the same partition-ordered barrier, fence, dedup, and business transaction; analytics must read the accepted ledger or use the identical fenced reader, never final-max filtering. ACLs must deny new groups, mirrors, replay tools, and bypass paths; a future consumer flips the decision and costs either the same fencing integration or broker-epoch migration.
The new owner must emit an authenticated barrier bound to the exact lease, shard, owner, and token, acknowledge it, and only then emit business records. Lower tokens are rejected; equal tokens require stable operation-level deduplication; only the validated barrier may advance the fence, preventing fabricated high tokens from poisoning the shard.
Partition ownership, stable dedup keys, checkpoint-after-transaction ordering, and load testing are mandatory: fence-row contention can create backlog and more lease expiry, while replicas, replay, and batch checkpoints can recreate split-brain behavior. If those throughput and recovery tests fail, do not ship the sink change.
Equal-token rejection was withdrawn because retries are legitimate and dedup is separate; the barrier and ordered analytics model changed the design on correctness grounds. Fire-and-forget revocation and Friday high-water instead of broker epochs were clock-driven concessions; without the deadline, broker-native fencing is the concession I would take back.
I move to @Muse’s materializer-side high-water mark plus deduplication for Friday, under the stipulated topology. My earlier fatal-distinction objection is closed only because we can redefine the raw shard topic as an untrusted ingress log, not the authoritative side effect, and force every correctness-relevant reader through one acceptance decision. If another consumer later interprets the raw topic directly, the original objection immediately returns.
I would not ship @Muse’s sketch unchanged, however. It needs an explicit generation barrier and authenticated token-to-shard binding.
The manifest is:
GENERATION_BARRIER(403) on that shard’s partition, and it waits for broker acknowledgement before emitting business records.shard_fences on a valid higher-token barrier; rejects later records with lower tokens; and handles equal-token business records through operation-level deduplication.Without the barrier, token 402 can append and be materialized after Postgres grants 403 but before the first token-403 business record advances the high-water mark. The barrier makes the Kafka partition position—not the Postgres lease CAS—the Kafka handover boundary. Records from 402 ordered before the barrier are treated as belonging to the old era; records from 402 ordered after it are stale and dropped. That is a narrower guarantee than “Kafka rejects 402 immediately when Postgres grants 403,” but it serializes accepted materialized effects.
Rejecting equality would break legitimate multiple writes and retries within one lease generation. The rule is therefore:
incoming_token < current_token: reject.incoming_token == current_token: accept only if the operation identity has not already been applied.incoming_token > current_token: accept only a valid generation barrier, not an arbitrary business record.A stable dedup key must identify the business operation independently of the lease generation—preferably source topic/partition/offset or a durable batch ID. Dedup retention must cover the entire possible replay horizon; deleting dedup entries while corresponding Kafka records remain replayable re-enables duplicate application.
The topology stipulation must be converted from institutional knowledge into an enforceable property:
ACL proof. Export Kafka ACLs and demonstrate that only the materializer principal can consume the raw shard topics. Remove the nightly analytics group’s raw-topic permission before enforcement activates. Configure deny-by-default so a newly created group does not inherit access.
Analytics proof. Run the nightly job against the accepted-event ledger and inject this sequence into staging: 402/data A, 403/barrier, 403/data B, 402/data C, duplicate 403/data B. Both the online materializer and analytics must include A and B exactly once and exclude C.
Partition/order proof. Demonstrate that all records for a shard—including its barrier—use the same partition key. Freeze token 402, acquire 403, append and acknowledge the barrier, resume 402, and prove no post-barrier token-402 record mutates materialized state.
Credential-binding proof. A valid token issued for shard:8842 must not advance shard:8841; a fabricated large token must not advance anything. A signed grant may establish this binding, but signature validation alone is not fencing—the high-water comparison remains essential.
Atomicity proof. Crash the materializer after each database statement boundary. There must be no state in which the fence advanced without the accepted effect or the effect committed without its fence/dedup record.
Replay proof. Rebuild materialized state from offset zero and obtain the same result. Test dedup-key collisions, dedup cleanup, partition reassignment, poison records, malformed barriers, and database deadlocks.
Bypass proof. Search is insufficient. Runtime telemetry must account for every read principal and consumer group on the raw topics, with an alert for any identity other than the approved materializer.
The migration price for a new consumer next quarter is explicit: it cannot read raw shard records as authoritative. It must consume the accepted ledger/curated output or implement and certify the identical barrier, fencing, and dedup state machine. If neither is acceptable, migrate to broker-native transactional-ID fencing before onboarding it.
The most dangerous unnamed failure is high-water poisoning. A consumer that advances cur_token merely because it sees a numerically larger header allows a buggy producer to attach a token from another lease—or BIGINT_MAX—and permanently fence the legitimate shard owner. A global sequence proves ordering only after the system also proves that the token was actually granted for the named shard. This is why a higher business record must not implicitly take over: only an authenticated, correctly bound generation barrier may advance the fence.
There are other sharp edges. A per-shard fence row becomes a serialization point for a hot shard. Dedup storage can grow without bound, while premature cleanup makes replays unsafe. Poison barriers can stall a partition if the error policy repeatedly retries instead of quarantining them. And database commit followed by Kafka offset commit has the usual replay window; atomic deduplication makes that replay harmless only if the operation ID is stable.
Most importantly, this design does not clean the Kafka log. Stale records will still be stored and visible to principals allowed to read the raw topic. ACL enforcement and the analytics migration are therefore part of the safety mechanism, not rollout hygiene.
I moved off broker-native fencing for Friday for two reasons.
First, the argument genuinely moved me: the stipulated topology means Kafka-log cleanliness is not itself the protected effect. If all authoritative interpretation is centralized in one materializer, sink-side generation filtering can provide the required state safety despite stale records remaining in the raw log.
Second, the Friday clock affects which adequate design I prefer. Stable per-lease transactional producers require more operational validation—coordinator capacity, transactional-path bypass removal, consumer isolation, and takeover semantics—than this team is likely to complete credibly by Thursday EOD. If the deadline vanished, I would take back the choice of high-water as the sole Kafka mechanism and add broker-native fencing in front of it. The materializer gate would remain as defense in depth.
I also moved from treating the missing revocation outbox as a notable deployment defect to accepting fire-and-forget for Friday. That move is deadline-driven, not argumentative: because revocation is explicitly outside the safety proof, an outbox improves operational delivery rather than correctness. I would add the transactional outbox in the next cycle.
The database-clock rebuttal moved me on the merits. When expiry is created and compared exclusively by authoritative Postgres expressions in the atomic CAS, client and service-node wall clocks do not decide reclamation. I would not take that concession back.
GO for Friday, conditional on all Thursday checks above passing.
I refuse to ship if:
The unresolved dissent I want preserved is that broker-native fencing remains the stronger end state. Materializer fencing protects the authoritative state only so long as topology and ACL assumptions remain true; broker rejection reduces the number of places where stale data exists and the number of future readers that can accidentally misuse it. Friday’s high-water design is acceptable because those assumptions can presently be enforced—not because consumer-side fencing is generally equivalent to broker fencing.
@GPT’s fatal-distinction argument is the cleanest architectural claim still standing: a consumer high-water mark leaves the Kafka log itself dirty, so any system that treats offsets or raw records as the externally visible side-effect remains exposed. Under the moderator’s stipulation, however, that distinction is no longer fatal for Friday. Only two consumer groups exist, both under our control and both able to run the identical fence transaction; topic ACLs can be tightened so nothing else can read. That closes the objection for the present topology, provided Thursday verification actually forces both readers through the fence and proves the analytics job cannot interpret a stale generation. If a third consumer group appears next quarter the call flips, and the migration cost is either forcing every new group through the same shard_fences + applied_writes path or completing the broker-epoch cut-over; I name that dependency explicitly rather than assume it away.
@Muse’s ship-now high-water design is therefore the only mechanism that can realistically clear the Thursday gate. I adopt it as the single Friday mechanism for the three classes: fence state lives in a dedicated shard_fences table (and companion applied_writes for equal-token deduplication), independent of the leases table; every materializing write and the analytics batch job execute the same single Postgres transaction that locks the fence row, rejects strictly lower tokens, accepts equal tokens only when the dedup key is new, and advances the high-water atomically with the business mutation. Kafka records may still land; the log is allowed to be dirty. Safety is defined solely at materialization and at the nightly analytics run.
Before I finalize that convergence I have to attack it. The design we just agreed on still fails in ways none of us has named. First, the materializer instances themselves form a second distributed system: if two materializer replicas process the same partition without their own mutual exclusion or if one freezes after reading a record but before the fence transaction commits, redelivery can re-apply under a token that has already been superseded elsewhere, and the dedup table only helps if the dedup key is present and stable across every producer path. Second, the analytics job is a batch that can lag by a full day; when it finally runs it must scan a log that already contains both generations, so a naïve “max token seen so far” implemented only inside the batch can still apply a late-arriving lower token if the fence row was not consulted for every record, or if the batch’s own checkpoint is not itself fenced. Third, any operational tool, replay script, or on-call human who reads the raw topic will observe dual-writer offsets and can re-inject them, defeating the materializer fence unless every such path is also forced through the same transaction. Fourth, the fence table becomes a hot row under high-frequency takeover; lock contention on FOR UPDATE can stall materialization long enough to create backlog that itself triggers more lease expiries. Finally, because the Kafka log remains dirty, any future compaction, tiered storage, or cross-region mirror that does not carry the fence state will re-surface the stale records as if they were authoritative. Those are not edge cases; they are the ordinary operational surface of a dirty-log design.
The concrete Thursday verification that closes @GPT’s objection under the stipulated topology is therefore:
cur_token; a synthetic dual-writer injection after a reclaim must produce zero materialization of the stale generation and zero analytics aggregation of it.If any of those four fail, high-water does not ship and Friday carries only the acquire/renew hardening (CACHE 1 sequence, atomic CAS on database time, fail-closed renew) plus fire-and-forget revocation labeled best-effort; the three classes remain on their present unfenced paths until a gateway or broker-epoch design clears a later gate.
I moved three positions this session. First, I entered treating broker-native epochs as the primary Friday Kafka path; the combination of @GPT’s gap analysis (no atomic span between Postgres CAS and InitProducerId) and @Muse’s realistic Thursday validation surface moved me to high-water for the ship-now cut. The Friday clock was the decisive pressure—without it I would still be arguing for epochs as the only clean log solution. Second, I conceded the original clock-skew objection once database-authoritative time was required in the CAS; that was argument, not deadline. Third, I accepted fire-and-forget for the revocation stream rather than insisting on an outbox in the same bundle; again the clock, because safety never depended on the stream. The one concession I would take back if the deadline vanished is the demotion of broker epochs: the dirty-log residual and the future-consumer migration cost make epochs the superior long-term mechanism, and I would hold Friday for them if the validation surface could be met.
Conditional GO for Friday on sink-side high-water mark enforcement at the two stipulated consumers, with the acquire/renew fixes and fire-and-forget revocation as supporting pieces.
Refuse-to-ship conditions (any one blocks):
CACHE 1 NO CYCLE, or any path can allocate a lower token after a higher one has been accepted for the same name.Unresolved dissent I want on the record: even under the stipulated topology the Kafka log remains a dirty audit trail; that is an accepted residual for Friday but it is not a permanent architecture. Broker epochs (or an exclusive produce gateway that owns the transactional.id) remain the only mechanism that makes the log itself free of stale generations, and the next consumer group will force us to pay that migration cost. The high-water design is correct for materialization safety today; it is incomplete for any system that ever treats the Kafka offsets as authoritative history.
For Friday I ship sink-side high-water at acceptance, not broker epochs.
Under your topology stipulation, @GPT's fatal distinction is closable but not free. If the Kafka log is side effect, a consumer fence does not clean the log. If the only readers are our materializer and one batch job, both owned by us, log pollution is contained to disk + ops cost, not correctness, provided both are forced through the same fence. That containment depends on ACLs and library enforcement. Break either and the bug returns.
Fence state: shard_fence(shard_id PK, cur_token, cur_owner, fenced_at DB time) + fence_history for audit. Not in leases table - must survive lease deletes.
Enforcement = where write is accepted:
SELECT ... FOR UPDATE, reject if incoming < cur, dedup if ==, else advance. Commit offset for Kafka consumer only after that txn commits. If you commit offset before fence row, crash -> fence regression.{token, owner, dedup_key, event_ts} in header. Broker does not reject. Materializer consuming in offset order maintains running max_seen. Pseudocode:running_max = load cur_token at start of segment
for rec in partition order:
if rec.token < running_max: drop stale late arrival
else if rec.token == running_max: dedup by dedup_key
else: // rec.token > running_max -> takeover point
begin tx; update shard_fence if rec.token > cur_token; do business; commit
running_max = rec.token
This is the fix for the <= bug moderator flagged. Fencing predicate is <, not <=. Equal-token is idempotency, not fencing.
Tradeoff vs broker epoch:
transactional.id = shard:{id}, TxnCoordinator becomes bottleneck at 1000s ids, txn init adds latency, zombie txn blocks takeover until timeout, and Postgres token 403 does not advance Kafka epoch by itself - you have two linearization points. Not demonstrable by Thursday EOD per your own gate 1-9.I back broker fencing as hardening next quarter, not Friday.
We converged on high-water + dedup at materializer, but none of us named retroactive invalidation:
SELECT cur_token at job start = 403 and then WHERE token < 403 DROP, it drops valid 402 records written before takeover. Valid history is 1..100 @402, 101..200 @403. Correct rule is not "drop all < cur_token", it's "drop < running_max after first higher token appears in offset order". Batch must replay in order, not filter by final max.Second, none of us named fence-then-offset ordering for materializer scale-out: two materializer replicas consuming different partitions of same shard can race on shard_fence update. That race is safe only with FOR UPDATE, but then materializer throughput becomes PG-contended. Also applied_writes grows unbounded - needs retention tied to dedup window.
Third, if materializer falls behind, nightly batch may see a takeover before materializer advances PG fence, leading to divergence between two readers about ownership at T. Solution: single fence source of truth is not PG table at start of job, it's Kafka offset order itself. PG fence is checkpoint for recovery.
fenced-kafka-consumer lib produces counts that exclude stale 50 but include early 100.fenced-kafka-consumer wrapper that loads fence logic + read_committed. CI bans raw KafkaConsumer/KafkaProducer for those topics. Topic ACLs: only materializer service account + analytics batch account can READ. New consumer group creation requires fencing review - flagged in infra repo.If analytics cannot adopt wrapper this week, fallback is to forbid direct Kafka read and make batch read materialized deduped snapshot at midnight. Document that.
If a new consumer group appears Q2, high-water requires that group adopt fence lib or move to gateway. Cost = re-plumbing or migration to broker fencing which we gated earlier at ~2-3 weeks + load tests. Assuming ACLs catch it is fragile - humans add consumers. Therefore manifest must say: "Friday's safety depends on only two enforced readers; new reader = new fence or broker epoch before it ships."
Settled: revocation stream UX-only, fire-and-forget, documented best-effort. Outbox not in Friday cut - commit-then-publish gap is acceptable because safety is in sink. Clock skew conceded: DB time authoritative in CAS.
nextval inside atomic UPDATE ... WHERE expires_at < now() / INSERT ... ON CONFLICT DO NOTHING, renew WHERE name=$ AND owner=$ AND token=$ AND expires_at > now() using DB time, sequence BIGINT NO CYCLE CACHE 1.shard_fence + dedup, ship sink enforcement shadow/log.shard:* leases, drain old fleet.incoming <= max hard reject after @GPT/@Grok pointed out equal-token retries are legitimate. Argument moved me - fencing is <, dedup is separate.max_token_for_shard in PG being sufficient to sequential-scan fence for analytics. Argument from moderator's batch concern moved me, not clock.Verdict: CONDITIONAL GO for Friday if and only if: sink rejects < cur, equal-token dedup exists, fence+offset commit order proven, both materializer and analytics batch forced through wrapper verified by test 1-3 above, CACHE 1 verified, and old client acquisition blocked before any reclaim.
Refuse if: any path can write without fence txn, analytics reads raw log without fence lib, sequence not CACHE 1, renew allows expired resurrect, offset committed before fence, or new consumer can appear via ACL bypass.
Dissent preserved: broker epoch is the correct long-term sink rejection for Kafka - high-water containment is brittle to future readers, and we are explicitly accepting log pollution as tech debt.