Skip to main content
Kodelyth ECC
engineering

The Service Worker That Poisoned Its Own Dev Server

A hydration error named an innocent component. The real cause was a cache-first service worker serving dev chunks that Turbopack reuses without content hashes — and the stale bundle re-registered the worker, closing the loop.

· 7 min read

A loop: the worker serves a stale chunk, that chunk holds the old registration code, which re-registers the worker. Hard reload, deleting the build directory and restarting the dev server all have no effect.

Every page of our dev server threw this:

Hydration failed because the server rendered HTML didn't match the client.
  <SupportMenu>
    <div
+     className={"fixed bottom-5 right-5 z-50 flex flex-col items-end …"}
-     className="support-fab"

Clear enough: SupportMenu renders one thing on the server and another on the client. Except SupportMenu renders support-fab. That is what its source says. The Tailwind string in the diff appears nowhere in the working tree — we grepped the whole repo.

The error named a component that was innocent.

Three wrong answers first

Worth recording, because each looked right.

The custom Cache-Control header. Next.js warns on every dev boot: "Custom Cache-Control headers detected … can break Next.js development behavior." We had max-age=31536000, immutable on /_next/static/. That is correct in production — those filenames carry a content hash — and wrong in dev. We gated it to production. The error persisted.

A stale .next. Deleted it, restarted. The error persisted.

"The stale code is in a served chunk." We grepped the served chunks and found the Tailwind string, which seemed conclusive. It was circular: the only file containing it was the dev error log, echoed back by the devtools overlay. We had found our own error message and called it evidence.

The actual cause

sw.js serves /_next/static/ cache-first, with a comment explaining that this is safe because Next.js content-hashes build output.

In production, true. Under Turbopack's dev server, chunk names are reused across compiles without a content hash. So the worker pinned whichever build it saw first and kept serving that copy after the source moved on.

And the stale bundle contained the old SwRegister, which re-registered the worker on every load. A closed loop that healed only if you knew to open DevTools → Application → Service Workers and unregister by hand.

const regs = await navigator.serviceWorker.getRegistrations();
const keys  = await caches.keys();
// ecc-v1 held 26 entries under /_next/static/ —
// one of them the chunk the server was serving correctly.

That is what finally proved it: the dev server was serving the right code (support-fab, 9 occurrences, curl'd straight from it), while the browser was rendering the wrong code.

Why every recovery failed

This is the part that makes it genuinely hard to diagnose.

What you'd tryWhy it doesn't work
Hard reloadThe worker intercepts before the HTTP cache
rm -rf .nextThe stale copy is in Cache Storage, not on disk
Restart the dev serverSame
fetch(url, { cache: 'reload' })The worker intercepts that too

cache: 'reload' is the one that fools you. It bypasses the HTTP cache, so it feels authoritative — but a service worker's fetch handler runs first, and ours answered from its own cache without ever consulting the network.

The fix needs two halves

if (process.env.NODE_ENV !== "production") {
  navigator.serviceWorker.getRegistrations()
    .then((regs) => Promise.all(regs.map((r) => r.unregister())));
  caches.keys().then((keys) => Promise.all(keys.map((k) => caches.delete(k))));
  return;
}
navigator.serviceWorker.register("/sw.js");

Gating registration is the obvious half. Unregistering is not optional, and this is the part that is easy to skip: gating alone fixes new machines and leaves every developer who already opened the site permanently broken, with no indication why. Their browser keeps serving the old bundle, which keeps re-registering.

Verified on a clean origin — a different port, so no cached entries — support-fab renders, zero registrations, zero caches, zero console errors.

What it cost, and what to take away

Three wrong hypotheses and a lot of time, because the error pointed at the wrong file and every instinctive fix appeared to do nothing.

Production was never affected. Worth saying plainly, since we initially reported otherwise: we checked five live pages before changing anything, and server and client both render support-fab with a clean console. Deployed filenames really are content-hashed, which is the assumption the worker was built on. The assumption was just never true in dev.

Three things generalise:

A service worker is a second cache you didn't think about, sitting in front of the one you did. When client and server disagree and the source says otherwise, check Cache Storage before you touch the source.

Content-hashing assumptions do not survive the dev server. Anything that reasons about immutable URLs — a worker, a CDN rule, a long max-age — needs a production guard.

When a diagnosis predicts something testable, test it before acting. Each of our three wrong answers was plausible and none was checked against the thing it predicted. The one that worked came from asking "what is actually in the cache?" instead of "what might be wrong?"