Auth · Session Revocation · Edge

auth-gateway

Google sign in for many sites on Cloudflare Workers, with sessions in D1 and no eventually consistent cache in the read path, so signing out takes effect on the very next request. One Durable Object per user pushes that to every open tab, so a new site integrates in two lines.

Released

View on GitHub ↗

In one line. auth-gateway holds sessions in a strongly consistent store with no cache in front of them, so a revoked session stops working on the very next request, and it pushes that revocation to every tab the user has open without any client writing sync code.

Signing out is a promise. A person clicks the button because they are on a shared laptop, or they think someone has their session. If it keeps working for another minute somewhere else, the system lied, and the person who trusted it is the one exposed. This is about finding that lie in my own code, why the cause is a default a popular Cloudflare integration still recommends, and what the corrected design looks like.

BEFORE Browser auth-gateway Worker Workers KV up to 60s stale authenticated revoked session read the database was consulted only when KV missed, so the stale answer always won AFTER Browser auth-gateway Worker D1 primary, no cache rejected on the next read read one authority, nothing cached in front of it Figure 1. Left: the read hits an eventually consistent cache first, so a colo that has not seen the delete keeps authenticating a revoked session. Right: one strongly consistent store, no cache in the read path, same request rejected.


Why this matters

  • A revocation that is not immediate is not a revocation. NIST SP 800-63B-4, finalized 31 July 2025, requires a session binding be terminated when the subscriber logs out. A store that takes up to 60 seconds to agree schedules a termination rather than performing one.
  • The defect ships as a default. The most visible Cloudflare integration for this auth library wires a cache its own documentation only warns about for rate limiting, while the same property silently delays revocation. Documented below with source, and filed upstream.
  • Multi tab agreement gets copied into every site that integrates. Behind the auth service instead, a new site adds two lines and inherits it.

🧭 If you only read this far: putting an eventually consistent cache in front of your session store means a signed out session keeps working for up to a minute somewhere else in the world, and the popular way to wire auth on Cloudflare does exactly that by default.


Terms in 30 seconds

Domain engineers: skip ahead, this is orientation for everyone else.

  • Eventual consistency: a write is not immediately visible everywhere, so readers in different places may see the old value for a while. Fine for a product catalogue, dangerous for a revocation. Strong consistency is the opposite guarantee, which Cloudflare D1 provides by sending queries to one primary instance.
  • Colo (colocation facility): one of the physical locations Cloudflare runs servers in. Your request hits whichever is nearest, which is why "the delete happened" and "the delete is visible here" are different statements.
  • Workers KV: Cloudflare's globally distributed key value store. Cloudflare's own storage guidance positions it for read heavy data that tolerates eventual consistency.
  • Durable Object: a Cloudflare primitive giving you exactly one addressable instance of an object, globally, that requests can be routed to. Used here as the single place a revocation and a tab's open connection can both reach.

The problem, for engineers who don't live in this space

Think about cancelling a hotel key card. The front desk deactivates it, and you would like to believe the card is dead. If the door locks only sync with the desk every so often, the card still opens some doors for a while. Nobody is lying; the deactivation is real, it is just not yet true everywhere, and "everywhere" is where the doors are.

The concrete version, from this repository. auth-gateway is a Cloudflare Worker handling Google sign in for my sites. Until this week it stored sessions in Workers KV: getSession asked KV first and consulted the backup store only when KV returned nothing, and deleteSession deleted the KV key.

That ordering is the whole bug. A user signs out in Singapore and the key is deleted. A request arrives in Frankfurt, where the value is still cached, so KV returns it and the gateway authenticates a revoked session. Cloudflare's own words are that changes "may take up to 60 seconds or more to be visible in other global network locations as their cached versions of the data time out", and that visibility takes longer still in locations which recently read a previous version of the key. Two independent reasons for the same stale answer.

Colo A Singapore Colo B Frankfurt Storage Workers KV t0 t1 user signs out DELETE key delete succeeds copy still cached request authenticated up to 60 seconds: the revoked session still works Figure 2. The revocation window. The delete is immediate in one colo; a request routed to any colo that has not caught up still authenticates, and the fresher store is never consulted because the stale one answered.

I found this while pricing the thing, not while auditing it. I was comparing KV against D1 for the session store, partly because KV's free tier caps writes at 1,000 a day and that ceiling had already cost me a fix. The line in Cloudflare's storage guidance that stopped the comparison was that KV suits read heavy data which tolerates eventual consistency. A session delete is the exact opposite of that.

Reading my own code with that in mind made it worse than the docs implied: the backup store's delete was wrapped in a try/catch that swallowed failures, so the bad case was not bounded at 60 seconds. If that delete failed, then once the KV key expired the read fell through to a backup row never removed, and the session came back. deleteUserSessions, the function you call on a password change, logged a warning and returned true without deleting anything.

Why this is genuinely hard. The instinct that produced the bug is a good one. Session reads happen on every authenticated request, so you put them in the fastest store and keep a durable one behind it. That is correct reasoning for a cache, and wrong here for a reason easy to miss: a cache in front of an authority can only make an answer staler, and for revocation, staler means insecure. The fix is not a better cache or a shorter TTL, it is recognising that this read must not be cached at all.


What exists today (and the gap)

ApproachSelf hostedRuns on Workers edgeNo eventually consistent store in the session read pathLive session sync provided by the service
Auth0, Clerk (hosted)❌⚠️ (SDK integration)✅✅ (via their SDK)
Better Auth, unconfigured✅✅ (D1 supported since 1.5)✅❌
Better Auth via better-auth-cloudflare✅✅❌❌
Rolling your own on Workers✅✅⚠️ (whatever you pick)❌
auth-gateway✅✅✅✅

The third column is the narrow claim, and it is empty for the integration most people would reach for. better-auth-cloudflare is the most visible package for wiring this library to Cloudflare, and its documented configuration passes a KV namespace as Better Auth's secondaryStorage. Better Auth's findSession consults secondary storage before the database and returns on a hit:

// better-auth/dist/db/internal-adapter.mjs
if (secondaryStorage) {
  const sessionStringified = await secondaryStorage.get(token);
  if (!sessionStringified && (!options.session?.storeSessionInDatabase ||
      ctx.options.session?.preserveSessionInDatabase)) {
    return null;
  }
  if (sessionStringified) {
    // returns the session; the database is never consulted
  }
}

The eventually consistent store short circuits the strongly consistent one. The package's README warns that KV's 60 second minimum TTL breaks rate limiting, which is true; the same property applied to session reads has the more serious consequence, and I could not find it documented anywhere. That is the gap this fills: an auth gateway on Cloudflare Workers with no eventually consistent store in the session read path, with the two settings that would reintroduce one pinned by regression tests.

The fourth column is a real differentiator but a softer one, deliberately not the novelty claim: hosted providers do give consumers session sync through their SDKs. What differs here is where it lives.


How it works

Moving the store made most of the old code redundant: the hand written OAuth handling, JWT minting and session management came out, roughly 10,800 lines, and runtime dependencies went from eight to four. What is left is one Cloudflare Worker that owns four things and deliberately owns nothing else.

  • Better Auth as a library, not a service. It runs inside the Worker: the code exchange happens in your account with your client secret, sessions are rows in your database, cookies are signed with your secret. The only third party in a sign in is the identity provider.
  • D1 as the sole session authority. It sends every query to the primary instance unless the Sessions API is explicitly used, so a delete is visible to the next read.
  • A Durable Object per user for live sync, holding tab connections with the WebSocket hibernation API.
  • A browser client served by the gateway, at /client.js, so consuming sites install nothing.

your Cloudflare account auth-gateway Worker D1: sessions sole authority, no cache in front Durable Object one per user, hibernating KV: rate limit counters only client sites Google /client.js code exch. every read publish session.changed (no session state on the wire) Figure 3. One authority, no cache in front of it. KV survives only for rate limit counters, where staleness is harmless. The Durable Object carries notifications, never session state.

Sessions in D1

One table, read by an indexed lookup on a unique token; a validation reads two rows, the session and its user, measured rather than assumed. There is no fallback store, which is a decision rather than an omission: a second store that can disagree with the first can hand back a session the first has revoked. If the D1 write fails at sign in, the sign in fails, because a session id no read can validate is worse than an error.

The two settings that would undo it

Both are configuration, both are one line, and both are covered by tests that fail if anyone adds them:

  • secondaryStorage, for the reason dissected above.
  • session.cookieCache, which keeps the session in a signed cookie on the client. A server cannot delete a cookie on someone else's device, so a revoked session survives until it expires. Better Auth's documentation is candid about this trade off; it is a legitimate feature, and not compatible with this project's central claim.

The Durable Object, and why the event carries no state

The request that revokes a session and the request holding a tab's open connection run in different isolates, usually in different colos, with no way to reach each other. A Durable Object addressed by idFromName(userId) is the one place both can find.

That addressing creates a subtlety, and getting it wrong would be a correctness bug. The hub is per user, so a push reaches every device, but signing out ends the session on one device only; the others are still legitimately signed in. Broadcasting "you are signed out" would log out devices that were not. So the event carries no state, only {type, reason, at}, and every tab asks get-session again and reaches its own conclusion. A test asserts no session identifier goes on the wire.

The client the gateway serves

Three mechanisms, because none covers every case: BroadcastChannel for same origin tabs, which works while signed out and so carries a sign in across tabs; the WebSocket to the hub for cross device revocation, which needs a session to address the hub and so cannot carry a sign in; and a refetch on visibilitychange as a backstop.


The decisions that actually mattered

D1 rather than keeping KV as a read through cache

The fork. The cheap fix is to move sessions to D1 and keep KV in front as a read through cache, which preserves the edge read latency that made KV attractive in the first place.

Options. (a) D1 with a KV cache, preserving edge read latency. (b) D1 alone. (c) Keep KV primary and add a revocation list.

Chosen. (b), which means giving up the edge read altogether. A read through cache reintroduces the precise defect: KV's minimum cacheTtl is 60 seconds, so a cached hit can vouch for a session D1 has deleted, and no cache configuration fixes it because the floor is the problem. Option (c) means maintaining a second consistency mechanism to compensate for the first being wrong.

Trade off accepted. Every authenticated request now queries the primary across regions instead of reading at the edge; mine is in APAC, so a European request pays that trip. Correctness on revocation is the product, and the old design already called a separate backend on every validation, so this replaces a hop rather than adding one. If it bites, read replication is the answer, and session reads must stay pinned to the primary or the window returns.

Verifying on a deployed worker, not only in tests

The fork. How much to trust a local suite before shipping auth to a live domain.

Options. (a) Ship on green tests. (b) Stand up a separate staging worker with its own database.

Chosen. (b), partly by force: Cloudflare refuses wrangler versions upload, the usual zero traffic preview, for any Worker carrying a Durable Object migration.

Trade off accepted. More infrastructure for a project with one user, and it paid for itself immediately by catching two bugs the local harness structurally cannot see. Sign out sent no content-type header and came back 415, which would have broken sign out for every consuming site. The session endpoint answered 426 rather than 401 to unauthenticated callers, because the Workers runtime drops a non conforming Upgrade header, so it answered about its protocol before its authentication. The local harness returns 200 to the exact request the real runtime returns 415 to: the two disagree about whether a bodyless POST has a body to type check. For this class of bug, a green suite was evidence of nothing.


Why this is new, and why it's significant

What's novel. An auth gateway on Cloudflare Workers with no eventually consistent store in the session read path, with the two settings that would reintroduce one pinned by regression tests. The narrow part is the last clause: choosing a strongly consistent database for sessions is ordinary practice, but treating "no cache in front of the session read" as an invariant with executable guards is not, in an ecosystem where the most visible integration package wires that cache on by default.

Why it's significant to the field. Two named standards make session termination a requirement rather than a preference. NIST SP 800-63B-4 (finalized 31 July 2025) requires the session binding be terminated on subscriber logout; OWASP ASVS 5.0.0 (released 30 May 2025) requires sessions be invalidated when no longer required. Neither says "eventually." Meanwhile the edge platform that most cheaply satisfies the rest of an auth stack ships a storage primitive whose documented consistency model is incompatible with that requirement, and the community integration layer wires it in without flagging the conflict. That gap between what the standards require and what the default tooling produces is worth naming precisely.

Standards and ecosystem alignment.

  • OAuth 2.0, RFC 6749: authorization code flow with the state parameter.
  • RFC 9700, Best Current Practice for OAuth 2.0 Security (BCP 240, January 2025): updates RFC 6749, 6750 and 6819; this deployment uses PKCE, which the BCP makes a baseline expectation rather than an option.
  • NIST SP 800-63B-4 (31 July 2025): session termination on logout.
  • OWASP ASVS 5.0.0 (30 May 2025): session management verification requirements.

Upstream contribution from this work (status as of September 2026):

  • zpg6/better-auth-cloudflare#61, the KV as secondaryStorage ordering analysed above. Fixed in #62, merged. The maintainer then took it further than the report: Better Auth also keeps each user's active session list in secondary storage with separate reads and writes, so concurrent changes can lose a token reference and a bulk revocation can miss a cached token until its original expiry. A strongly consistent store removes the propagation lag but does not make that update atomic. He also shipped #63, narrowing the KV adapter's type rather than faking the atomic getAndDelete and increment that Better Auth 1.7 requires and Workers KV cannot provide.

📌 Honest scope. One provider (Google). Single account Cloudflare deployments, not multi tenant SaaS. The revocation property is verified by automated tests and a live two connection fan out against a deployed worker; it is not load tested. The WebSocket upgrade is not exercised in continuous integration at all, because the local runtime forbids constructing the 101 response the real one requires, so hibernation and reconnect are verified only by deployment. Expired rows are rejected on read but not yet swept on a schedule. Two people have used this system, and one of them is me.


See it run

bun test in the repository. The revocation property is the first test, written to leave no room for a propagation window:

✓ a signed out session is rejected on the very next read
✓ a session revoked directly in the database is rejected on the very next read
✓ revoking one session leaves the user other sessions
✓ no secondary storage is configured          # ← guard, not behaviour
✓ cookie cache is not enabled                 # ← guard, not behaviour

The last two assert configuration, not behaviour, which is the point: they stop a future change reopening the window quietly.

Fan out, against a deployed worker with two real connections:

opening two connections...
both open (101 upgrade succeeded)
signing out from tab-A...
  sign-out status: 200
  tab-B received: {"type":"session.changed","reason":"signed-out",...}   # ← the other tab
  tab-A received: {"type":"session.changed","reason":"signed-out",...}
session after sign-out: null

🔬 Go deeper. The two guarded settings and the reasoning are in src/auth.ts; the invalidation design is in src/durable/session-hub.ts. Open auth.in8.sh/demo in two tabs and sign out of one.

Tab A Tab B Gateway Hub (DO) POST /sign-out resolve user id first the session is about to be gone publish {type, reason, at} no session id on the wire session.changed session.changed get-session get-session Tab A -> null (same device, signed out) Tab B -> still valid (another device) Figure 4. Sign out in one tab. The gateway resolves the user before the handler runs, because afterwards no session remains to address the hub with; the hub notifies every connection; each tab re-asks and reaches its own answer.


What's next

The nearest step is running the test suite inside the real Workers runtime rather than beside it, with @cloudflare/vitest-pool-workers. Both bugs that reached a deployment were invisible locally, and one made a passing test actively misleading. After that, a scheduled sweep for expired rows, since the database has no expiry of its own.

The larger arc: this is the piece the other projects sign in through. Making the auth service own live session state, rather than exporting that problem to every site that integrates, is what lets the next site be two lines instead of a subscription implementation.


Appendix / references

</invoke>