Frontend Platform Engineering — A Working Reference

A reference for the areas a frontend platform engineer is expected to reason about: web protocols and security, core web platform concepts, performance and bundling, and testing. Every topic follows the same shape — What it is · Example · Trade-offs · Follow-ups — so it can be used for lookup and revision rather than read front to back.

1. How to use this reference

Each topic is short on purpose. The definitions are written to survive a follow-up question, and the trade-off line is the part worth memorising — knowing that a technique exists is cheap, knowing what it costs is the part that transfers.

Three passes. First pass: read the definitions and examples without memorising. Second pass: explain each topic aloud without looking. Third pass: answer the follow-up questions — they are where the real understanding shows.

1.1 The areas this covers

Area What it spans Sections
Web protocol and security Browser interactions, security layers, data flow, resource delivery and persistence 6, 7, 8
Core web platform concepts Rendering strategies, component models, modern JS, build tools, styling 2, 3, 4, 5, 11
Performance and bundling DevTools fluency, root-cause diagnosis, systemic fixes at scale 9, 10
Testing Designing strategies and frameworks, raising quality across teams 12, 13

1.2 How each topic is laid out

2. JavaScript core

2.1 Execution context and the call stack

The stack grows with every call and unwinds on every return Each frame holds that call's arguments, locals and where to return to function third() { return 3; } function second() { return third(); } function first() { return second(); } first(); global 1 · script runs global first() 2 · first() global first() second() 3 · second() global first() second() third() 4 · third() — deepest global first() second() 5 · third returned global 6 · unwound push ↑ base A stack trace prints these frames, deepest first. Unbounded recursion overflows the stack → RangeError: Maximum call stack size exceeded. Every frame here is synchronous, which is why one slow function blocks rendering, input and everything else on the page.
The call stack across six moments — push on call, pop on return.

What it is. Before code runs the engine creates an execution context holding a variable environment, a lexical environment and a this binding. Contexts stack on the call stack; the engine runs one at a time, LIFO.

Example. A stack trace is the call stack printed. RangeError: Maximum call stack size exceeded is that stack overflowing — typically unbounded recursion.

Trade-offs. The stack is synchronous and single-threaded. Anything long-running on it blocks rendering and input, which is the root of every INP problem.

Follow-ups. What's on the stack when a promise resolves? Why does a stack trace lose frames across an await?

More examples.

// 1 — the stack explains why this catch never fires
try {
  setTimeout(() => { throw new Error('boom'); }, 0);
} catch (e) {
  // unreachable: the callback runs later, on a fresh, empty stack
}

// 2 — but this one does, because await resumes the same logical frame
try {
  await Promise.reject(new Error('boom'));
} catch (e) {
  // caught
}

// 3 — recursion depth is finite; depth is roughly 10k frames in V8
const depth = (n = 0) => depth(n + 1);
try { depth(); } catch (e) { e.name; } // 'RangeError'
Trace it — what order do the logs print?
function a() { console.log('a'); b(); console.log('a done'); }
function b() { console.log('b'); c(); console.log('b done'); }
function c() { console.log('c'); }
a();

Answer: a · b · c · b done · a done.

Each call pushes a frame and the caller is suspended, not finished. b done cannot print until c pops, and a done cannot print until b pops. The "done" lines print in reverse call order — that is the unwinding.

2.2 Hoisting and the temporal dead zone

All four are registered before line one runs — only two are usable A single scope, left to right in execution order · ◆ marks the declaration statement creation execution begins var a = 1 undefined 1 function foo(){} callable — hoisted whole, before its own line let b = 2 temporal dead zone — ReferenceError 2 const c = 3 temporal dead zone — ReferenceError 3 registered, value undefined registered, reading throws initialised and usable Why the TDZ exists: const must be guaranteed assigned exactly once, so silently reading undefined would make that unenforceable. typeof is not safe inside the TDZ either — it throws, which is the one case where typeof can fail.
When each kind of declaration becomes usable, and where the temporal dead zone sits.

What it is. Declarations are registered during the creation phase. var is initialised to undefined; function declarations are initialised fully; let, const and class are registered but uninitialised — reading them before the declaration throws. That window is the temporal dead zone.

Example.

console.log(a); // undefined
console.log(b); // ReferenceError — TDZ
var a = 1;
let b = 2;

foo();          // works — declaration hoisted whole
bar();          // TypeError: bar is not a function
function foo() {}
var bar = function () {};

Trade-offs. The TDZ turns a class of silent undefined bugs into loud errors, at the cost of ordering discipline.

Follow-ups. Why does the TDZ exist? Are function declarations hoisted inside blocks? What does typeof b do in the TDZ? (It throws — the one case where typeof is not safe.)

More examples.

// 1 — the loop-variable classic
for (var i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 3 3 3
for (let i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 0 1 2

// 2 — const protects the binding, not the value
const user = { name: 'Ada' };
user.name = 'changed';   // fine — the object is mutable
// user = {};          // TypeError: assignment to constant variable

// 3 — a function declaration inside a block is block-scoped in modules
if (true) { function f() {} }
// f is not reliably visible outside the block — use a const arrow instead
Predict the output
console.log(typeof x);
console.log(typeof y);
var x = 1;
let y = 2;

Answer: 'undefined', then a ReferenceError.

typeof is famously safe on undeclared identifiers — but not inside the temporal dead zone. y is declared, just not yet initialised, so the engine throws rather than reporting 'undefined'. This is the one case where typeof can fail, and it is a common interview trap.

2.3 Scope, closures and the scope chain

The outer call is gone — its scope is not A closure is a function plus a live reference to the scope it was created in function makeCounter() { let count = 0; return () => ++count; } const inc = makeCounter(); inc(); // 1 inc(); // 2 GLOBAL ENVIRONMENT inc makeCounter() ENVIRONMENT — the call already returned count: 2 not collectable () => ++count the returned function holds [[Environment]] The same mechanism is the most common leak in a long-lived SPA An event listener closing over a large DOM subtree keeps that subtree alive after unmount. Fix with an AbortController or an effect cleanup, not by avoiding closures.
A closure keeping its defining scope alive after the outer call returned.

What it is. A closure is a function plus a reference to the lexical environment it was created in. Lookup walks the scope chain outward to the global scope.

Example.

for (var i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 3 3 3 — one binding
for (let i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 0 1 2 — per-iteration binding

Trade-offs. Closures are the most common memory leak in long-lived SPAs: a listener closing over a DOM subtree keeps it alive after unmount. Guard with cleanup functions and AbortController.

Follow-ups. How would you implement once, memoise, or a private counter with a closure? How do closures cause detached-DOM leaks?

More examples.

// 1 — private state, no class needed
function counter() {
  let n = 0;                       // unreachable from outside
  return { inc: () => ++n, value: () => n };
}

// 2 — run-once, a closure over a boolean
const once = fn => {
  let done = false, result;
  return (...args) => done ? result : (done = true, result = fn(...args));
};

// 3 — memoise: the cache lives in the closure, not in global scope
const memo = fn => {
  const cache = new Map();
  return x => cache.has(x) ? cache.get(x) : (cache.set(x, fn(x)), cache.get(x));
};
Why does this leak, and what is the one-line fix?
function attach(node) {
  const bigData = new Array(1e6).fill(node.textContent);
  node.addEventListener('click', () => console.log(bigData.length));
}

Answer: the listener closes over bigData and over node. Even after the node is removed from the DOM, the listener keeps both alive — a detached DOM node plus a million-element array, retained for the life of the page.

The fix is to tie the subscription to a lifetime:

const ac = new AbortController();
node.addEventListener('click', handler, { signal: ac.signal });
// later: ac.abort();  — removes the listener and releases the closure

Note the cause is not "closures are leaky" — it is that nothing ever removed the subscription. The closure is just what made the retention large.

2.4 The event loop

Microtasks drain completely before the browser is allowed to paint One turn of the event loop 1 · Run ONE macrotask to completion 2 · Drain ALL microtasks including newly queued ones 3 · rAF callbacks requestAnimationFrame 4 · Style → Layout → Paint the frame reaches the screen next turn · idle time left over goes to requestIdleCallback Macrotask queue — one taken per turn setTimeout · setInterval · setImmediate DOM events (click, scroll, input) I/O · postMessage · MessageChannel Nested setTimeout clamps to 4 ms after 5 levels Microtask queue — drained entirely promise .then / .catch / .finally everything after an await queueMicrotask · MutationObserver A microtask that queues a microtask never yields Starvation: Promise.resolve().then(spin) freezes the tab · setTimeout(spin, 0) does not — it yields at step 4 every turn.
One turn of the event loop. Microtasks drain completely before the browser may paint.

What it is. One turn: take one macrotask → run it to completion → drain the entire microtask queue → run requestAnimationFrame callbacks → style, layout, paint → requestIdleCallback if time remains.

Example.

console.log('1 sync');
setTimeout(() => console.log('5 macrotask'), 0);
Promise.resolve().then(() => console.log('3 microtask'));
queueMicrotask(() => console.log('4 microtask'));
console.log('2 sync');
// 1 sync, 2 sync, 3 microtask, 4 microtask, 5 macrotask

Trade-offs. A microtask that queues another microtask starves rendering forever; a setTimeout loop does not (nested timeouts clamp to 4ms after five levels). Microtasks give ordering guarantees at the cost of being unyielding.

Follow-ups. Predict the output of a mixed setTimeout/promise/async snippet. Why does Promise.resolve().then(spin) freeze the tab? Where does requestAnimationFrame run relative to microtasks?

More examples.

// 1 — await is microtask scheduling, so this prints 1 4 2 3
async function f() { console.log(2); await null; console.log(3); }
console.log(1); f(); console.log(4);
// 1 · 2 · 4 · 3  — the body runs synchronously up to the first await

// 2 — rAF runs after microtasks, before paint
Promise.resolve().then(() => console.log('micro'));
requestAnimationFrame(() => console.log('frame'));
// micro · frame

// 3 — yielding so input can be handled between chunks
async function process(items) {
  for (const [i, item] of items.entries()) {
    work(item);
    if (i % 100 === 0) await new Promise(r => setTimeout(r, 0));
  }
}
Order these eight logs
console.log('1');
setTimeout(() => console.log('2'), 0);
Promise.resolve().then(() => { console.log('3'); return Promise.resolve(); })
                 .then(() => console.log('4'));
queueMicrotask(() => console.log('5'));
(async () => { console.log('6'); await null; console.log('7'); })();
console.log('8');

Answer: 1 · 6 · 8 · 3 · 5 · 7 · 4 · 2

  • 1, 6, 8 are synchronous. The async IIFE body runs immediately up to its await.
  • The microtask queue then drains in the order things were queued: 3, then 5, then 7.
  • 4 is late because returning a promise from .then costs two extra microtask ticks to unwrap — this is the detail almost everyone misses.
  • 2 is a macrotask, so it runs on the next turn, last.

2.5 Promises and async/await

await in a loop serialises what could have run together Six requests, each one unit of latency · bars are time, stacked bars are concurrent 0 1 2 3 4 5 6 units for … await serial 6 units · 1 request in flight Promise.all all at once 1 unit · 6 sockets at once — fails fast if any one rejects mapLimit(3) bounded 2 units · never more than 3 in flight — the answer that scores all → fail fast · allSettled → never rejects, use for telemetry any → first fulfilled, AggregateError · race → first settled, the timeout idiom Unbounded Promise.all over a large list floods the server and exhausts the browser's connection pool. Bound it.
Serial await vs Promise.all vs bounded concurrency, on the same six requests.

What it is. await x splits the function: everything after it becomes a microtask continuation. async functions always return a promise.

Example. Serial vs parallel vs bounded concurrency:

for (const id of ids) out.push(await fetch(url(id)));        // N × latency
const out = await Promise.all(ids.map(id => fetch(url(id)))); // 1 × latency, N sockets

async function mapLimit(items, limit, fn) {                   // bounded
  const out = [], running = new Set();
  for (const [i, item] of items.entries()) {
    const p = Promise.resolve(fn(item, i)).then(r => { running.delete(p); return r; });
    out.push(p); running.add(p);
    if (running.size >= limit) await Promise.race(running);
  }
  return Promise.all(out);
}

Trade-offs. all fails fast (one rejection discards the rest); allSettled never rejects, right for dashboards and telemetry; any returns the first fulfilled and rejects with AggregateError; race returns the first settled and is the timeout idiom. Unbounded Promise.all over a large list floods the server and the browser's connection pool.

Follow-ups. How do you add a timeout to a fetch? How do you cancel in-flight work? What is an unhandled rejection and how do you report it?

More examples.

// 1 — timeout any promise, using race
const withTimeout = (p, ms) => Promise.race([
  p,
  new Promise((_, rej) => setTimeout(() => rej(new Error('timeout')), ms)),
]);

// 2 — retry with exponential backoff and jitter
async function retry(fn, tries = 3) {
  for (let i = 0; i < tries; i++) {
    try { return await fn(); }
    catch (e) {
      if (i === tries - 1) throw e;
      const wait = 2 ** i * 100 + Math.random() * 100;   // jitter avoids a thundering herd
      await new Promise(r => setTimeout(r, wait));
    }
  }
}

// 3 — genuinely sequential, when each step needs the previous result
const result = await steps.reduce(
  (acc, step) => acc.then(step),
  Promise.resolve(initial),
);
What does this log, and why is it a bug?
const results = [];
[1, 2, 3].forEach(async (n) => {
  results.push(await double(n));
});
console.log(results.length);

Answer: 0.

forEach ignores the returned promise, so all three callbacks are started and abandoned — console.log runs before any of them resume. await inside a callback does not make the outer function wait.

const results = await Promise.all([1, 2, 3].map(double));   // 3

The general rule: forEach is not async-aware. Use map + Promise.all, or a for…of loop if you genuinely need them sequential.

2.6 this, call/apply/bind, arrow functions

`this` is decided by the call, not by where the function was written Check these four in order — the first match wins How was it called? resolved at call time 1 · new Fn() the brand-new instance 2 · fn.call / .apply / .bind whatever you passed 3 · obj.fn() obj — the receiver left of the dot 4 · fn() undefined in modules / strict · globalThis otherwise Arrow functions skip all four They have no `this` binding of their own and close over the enclosing scope — so .call() and .bind() cannot change them. The classic break const f = obj.method; f() — the receiver is lost, so rule 3 no longer applies and you fall through to rule 4.
The four call-site rules that decide `this`, in priority order.

What it is. this is resolved at call time by four rules in priority order: new → explicit call/apply/bind → method receiver → default (undefined in strict mode and modules, globalThis otherwise). Arrow functions have no this binding and close over the enclosing scope.

Example. const f = obj.method; f() loses the receiver; obj.method.bind(obj) or () => obj.method() keeps it.

Trade-offs. Class-field arrows bind per instance — convenient but they cost memory per instance and are not on the prototype, so they cannot be overridden or spied on easily.

Follow-ups. What is this in a standalone function in a module? In a setTimeout callback? In an event handler? Implement bind yourself.

More examples.

// 1 — the receiver is lost the moment you detach the method
const obj = { n: 1, get() { return this.n; } };
const f = obj.get;
obj.get(); // 1      — rule 3, receiver is obj
f();       // throws — rule 4, this is undefined in a module

// 2 — array callbacks take an explicit thisArg
[1, 2].map(function () { return this.k; }, { k: 9 }); // [9, 9]

// 3 — prototype method vs class field, and what each costs
class A {
  onClick() {}            // on A.prototype — one copy, needs binding
  onTap = () => {};       // per instance — pre-bound, costs memory per object
}
Which of these four log 10?
const o = {
  n: 10,
  regular() { return this.n; },
  arrow: () => this?.n,
  nested() { return [1].map(function () { return this?.n; })[0]; },
  nestedArrow() { return [1].map(() => this.n)[0]; },
};

Answer: regular() and nestedArrow() return 10.

  • regular — rule 3, the receiver is o.
  • arrow — defined at module top level, so it closed over module scope, where this is undefined. The object literal creates no scope.
  • nested — the inner function is called by map with no receiver, so rule 4 applies and this is undefined.
  • nestedArrow — the arrow closes over nestedArrow's this, which is o.

The takeaway: an arrow is only useful for this when it is nested inside a function that already has the this you want.

2.7 Prototypes and classes

A property miss walks up the chain until it hits null class Dog extends Animal · const d = new Dog() d name: 'Rex' own properties only Dog.prototype bark() constructor Animal.prototype speak() eat() Object.prototype hasOwnProperty() toString() ∅ [[Prototype]] [[Prototype]] [[Prototype]] null d.speak() — not own, not on Dog.prototype, found on Animal.prototype What lives where Methods → the prototype, shared by every instance (cheap) Class fields and constructor assignments → the instance (per-object cost) Class-field arrows → per instance, and not overridable or easily spied on Questions this answers `in` walks the chain · hasOwnProperty does not instanceof checks whether a .prototype appears anywhere on the chain Mutating a built-in prototype breaks every library sharing the realm
Property lookup walking the prototype chain until it reaches null.

What it is. Every object has an internal prototype link (Object.getPrototypeOf(obj) === Ctor.prototype). Property lookup walks that chain. class is syntax over it: methods land on the prototype, class fields on the instance.

Example. Array.prototype.map is found by walking from the array instance to Array.prototype. obj.hasOwnProperty('x') checks only the own object; 'x' in obj walks the chain.

Trade-offs. Prototype sharing saves memory; prototype mutation at runtime deoptimises V8's hidden classes and should be avoided in hot paths. Monkey-patching built-in prototypes breaks other libraries.

Follow-ups. Difference between __proto__ and prototype? How does instanceof work? How would you implement inheritance without class?

More examples.

// 1 — instanceof is a walk, not a type tag
class A {} class B extends A {}
const b = new B();
b instanceof A;                              // true — A.prototype is on the chain
Object.getPrototypeOf(B.prototype) === A.prototype; // true

// 2 — own vs inherited
const o = Object.create({ inherited: 1 });
o.own = 2;
'inherited' in o;                  // true  — `in` walks the chain
o.hasOwnProperty('inherited');     // false — own properties only
Object.keys(o);                    // ['own'] — own and enumerable

// 3 — a null-prototype object has no inherited methods at all
const dict = Object.create(null);
dict.toString;                     // undefined — safe as a plain string map
Why does this break, and what is the safer check?
const data = JSON.parse('{"hasOwnProperty": 1}');
data.hasOwnProperty('x');   // TypeError: not a function

Answer: the parsed object shadows the inherited method with a number, so calling it fails. Any object built from untrusted input can do this.

Safe forms, in order of preference:

Object.hasOwn(data, 'x');                          // modern, clearest
Object.prototype.hasOwnProperty.call(data, 'x');   // the classic borrow

The same reasoning is why Object.create(null) is the right choice for a lookup map — there is nothing on the chain to shadow or to collide with.

2.8 Memory and garbage collection

Nothing is freed because it is "unused" — only because it is unreachable Mark and sweep walks out from the roots; whatever it never reaches is collected GC roots globals, stack REACHABLE — retained app state router rendered component tree UNREACHABLE — swept on the next cycle old route's data components that unmounted cleanly × The leak: one surviving reference keeps a whole subtree reachable GC roots → a module-level array of handlers → a closure → the detached DOM node it captured → its entire subtree. Tools that break the path WeakMap / WeakSet keys · WeakRef + FinalizationRegistry · AbortController How you find it Snapshot → interact → snapshot → compare → sort by retained size → retainer path
Reachability from the GC roots — and the single reference that causes a leak.

What it is. V8 uses generational mark-and-sweep: a scavenged young space plus a mark-compact old space. Objects surviving two scavenges are promoted. Anything reachable from a root is retained.

Example. Leak taxonomy: detached DOM nodes, forgotten timers and listeners, unbounded module-level caches, closures over large objects. WeakMap/WeakSet keys do not retain; WeakRef + FinalizationRegistry for caches; AbortController for listener and fetch lifecycle.

Trade-offs. Weak collections prevent leaks but make eviction non-deterministic — you cannot rely on a finalizer running.

Follow-ups. How do you find a leak in DevTools? (Heap snapshot → interact → second snapshot → "objects allocated between 1 and 2" → sort by retained size → follow the retainer path.) What does a sawtooth memory graph that never returns to baseline mean?

More examples.

// 1 — a WeakMap lets the key be collected; a Map does not
const meta = new WeakMap();
meta.set(node, { seen: true });   // when `node` dies, the entry dies with it

// 2 — the three subscriptions that outlive a component
const id = setInterval(tick, 1000);          // must clearInterval
window.addEventListener('resize', onResize); // must removeEventListener
const sub = store.subscribe(onChange);       // must unsubscribe
// one controller can cover the listeners:
const ac = new AbortController();
window.addEventListener('resize', onResize, { signal: ac.signal });

// 3 — an unbounded module-level cache is a leak with good intentions
const cache = new Map();                    // grows forever
const bounded = new Map();                  // evict by size or time instead
Two snapshots, same page, 40 MB apart. What do you look at first?

Answer: the Comparison view, filtered to objects allocated between snapshot 1 and snapshot 2, sorted by retained size — not shallow size.

Shallow size is the object itself; retained size is everything that would be freed if it went away. A leak is usually a small object retaining a large graph, so sorting by shallow size hides it.

Then follow the retainer path from the suspect up to a GC root. That path is the answer: it names the exact reference that must be released. A path ending in a module-level Map, an array of handlers, or a detached DOM node is the overwhelmingly common shape.

Quick sanity check before any of that: in the Performance panel, record an interaction loop and look for a sawtooth that never returns to its starting baseline. If memory does return to baseline, you have churn, not a leak.

2.9 The patterns asked by name

Pattern What it is Typical use
Debounce Run after a quiet period Search-as-you-type
Throttle At most once per interval Scroll, resize, mousemove
Event delegation One listener on a container, dispatch via event.target 1,000-row tables
Capture vs bubble Root→target, then target→root; {capture: true} Intercepting before a child handles it
passive: true Promises not to preventDefault Lets scroll run on the compositor
structuredClone Deep clone with cycles, Map, Date Replaces JSON.parse(JSON.stringify())
Generators Lazy sequences via Symbol.iterator Pagination, async streams
Proxy/Reflect Intercept property access Vue 3 reactivity, MobX

Follow-ups. Implement debounce with a leading edge and cancel(). Why does JSON.parse(JSON.stringify(x)) lose undefined, Date and cycles? When is stopImmediatePropagation needed over stopPropagation?

3. TypeScript

More examples.

// 1 — debounce with a cancel, the version worth memorising
function debounce(fn, ms) {
  let t;
  const wrapped = (...a) => { clearTimeout(t); t = setTimeout(() => fn(...a), ms); };
  wrapped.cancel = () => clearTimeout(t);
  return wrapped;
}

// 2 — throttle: trailing edge preserved
function throttle(fn, ms) {
  let last = 0, timer;
  return (...a) => {
    const now = Date.now(), wait = ms - (now - last);
    if (wait <= 0) { last = now; fn(...a); }
    else { clearTimeout(timer); timer = setTimeout(() => { last = Date.now(); fn(...a); }, wait); }
  };
}

// 3 — event delegation: one listener for any number of rows
table.addEventListener('click', (e) => {
  const row = e.target.closest('[data-hotel-id]');
  if (row && table.contains(row)) select(row.dataset.hotelId);
});
Search-as-you-type: debounce or throttle, and what else?

Answer: debounce — you want one request after the user stops typing, not a steady stream while they type. Throttle is for continuous streams you must sample (scroll, resize, pointermove).

But debounce alone is not enough. Three more things belong in a real implementation:

  1. Cancel the in-flight request when a newer keystroke supersedes it — AbortController — otherwise a slow early response can overwrite a fast later one.
  2. Guard against out-of-order responses even so: tag each request and drop any response that is not the latest.
  3. Mark the update as non-urgent with useDeferredValue or startTransition, so re-filtering a large list never blocks the keystroke itself.

The race condition in point 2 is the one candidates usually miss, and it is a real bug class: the user sees results for a query they already edited away.

3.1 Structural typing and branded types

What it is. TypeScript compares types by shape, not by name. Two unrelated types with the same members are interchangeable.

Example. To get nominal behaviour for identifiers:

type HotelId = string & { readonly __brand: 'HotelId' };
type CityId  = string & { readonly __brand: 'CityId' };
const asHotelId = (s: string) => s as HotelId;
// passing a CityId where a HotelId is expected is now a compile error

Trade-offs. Branding costs a cast at the boundary and slightly noisier types, and buys you a whole class of ID-mixup bugs caught at compile time.

Follow-ups. Why does TypeScript use structural typing? Where does structural typing surprise people? (Excess property checks apply to object literals only.)

3.2 any vs unknown vs never

What it is. any disables checking and propagates silently. unknown is the safe top type — assignable from anything, assignable to nothing without narrowing. never is the bottom type: no value inhabits it.

Example. never gives exhaustiveness for free:

function render(s: FetchState) {
  switch (s.status) {
    case 'loading': return spinner();
    case 'error':   return alert(s.error.message);
    case 'success': return list(s.data);
    default: { const _never: never = s; return _never; } // errors if a variant is added
  }
}

Trade-offs. Banning any outright creates friction at untyped third-party boundaries; the usual policy is unknown plus a parse, with any allowed only behind an explicit lint suppression and a comment.

Follow-ups. Why is unknown safer than any? What does never mean as a return type? How do you type a function that always throws?

3.3 Discriminated unions

What it is. A union of object types sharing a literal-typed discriminant field, which lets the compiler narrow by checking that field.

Example.

type FetchState<T> =
  | { status: 'idle' }
  | { status: 'loading' }
  | { status: 'error'; error: Error }
  | { status: 'success'; data: T };

Trade-offs. More verbose than { loading: boolean; error?: Error; data?: T }, but it makes impossible states unrepresentable — you can never have loading: true with data present.

Follow-ups. Model a booking flow's states. How does narrowing work with in, typeof and custom type guards (x is Foo)?

More examples.

// 1 — the discriminant does not have to be called "status"
type Shape =
  | { kind: 'circle'; r: number }
  | { kind: 'rect'; w: number; h: number };
const area = (s: Shape) => s.kind === 'circle' ? Math.PI * s.r ** 2 : s.w * s.h;

// 2 — a custom type guard narrows across function boundaries
function isError<T>(s: FetchState<T>): s is Extract<FetchState<T>, { status: 'error' }> {
  return s.status === 'error';
}

// 3 — a booking funnel where impossible states cannot be constructed
type Booking =
  | { step: 'dates' }
  | { step: 'rooms'; dates: DateRange }
  | { step: 'pay';   dates: DateRange; room: Room }
  | { step: 'done';  confirmation: string };
// you cannot reach 'pay' without a room — the type enforces the order
Why is the boolean version worse?
interface State<T> { loading: boolean; error?: Error; data?: T }

Answer: it permits combinations that cannot happen, and forces every reader to defend against them. { loading: true, error: e, data: d } type-checks perfectly, so each consumer invents its own precedence rule — and they disagree.

A discriminated union makes illegal combinations unrepresentable, and gives you exhaustiveness for free: add a 'refetching' variant and every switch that forgot it fails to compile. With the boolean shape, adding a state means hunting down every if (loading) by hand and hoping you found them all.

3.4 Generics, conditional and mapped types

What it is. Generics parameterise types; conditional types branch on assignability with infer; mapped types transform every key of a type, optionally remapping keys with as.

Example.

type Unwrap<T> = T extends Promise<infer U> ? Unwrap<U> : T;
type Getters<T> = { [K in keyof T as `get${Capitalize<string & K>}`]: () => T[K] };
// Getters<{ name: string }> === { getName: () => string }

type PolymorphicProps<E extends React.ElementType> =
  { as?: E } & Omit<React.ComponentPropsWithoutRef<E>, 'as'>;

Trade-offs. Deep recursive conditional types are expressive but slow the compiler and produce unreadable errors. A platform rule of thumb: if the type needs a comment to explain it, prefer an explicit overload.

Follow-ups. Implement DeepPartial. What are Partial, Required, Pick, Omit, Record, ReturnType built from? Why do large unions blow up compile time?

3.5 Variance and where TypeScript is unsound

What it is. Function parameters are contravariant and returns covariant. TypeScript deliberately accepts some unsound patterns for ergonomics.

Example. Known unsound spots: arrays are covariant (Dog[] assignable to Animal[], then you can push a Cat); method parameters are bivariant unless strictFunctionTypes applies; index access without noUncheckedIndexedAccess claims T where T | undefined is true; as assertions bypass everything.

Trade-offs. Full soundness would reject large amounts of idiomatic JavaScript. Enabling noUncheckedIndexedAccess is correct but noisy on existing code.

Follow-ups. Where is TypeScript deliberately unsound, and why? What does strictFunctionTypes change?

3.6 Runtime validation at the boundary

What it is. Types are erased at compile time, so anything crossing a boundary — API responses, URL params, localStorage, postMessage, feature flags — must be parsed, not asserted.

Example.

const Hotel = z.object({ id: z.string(), price: z.number(), currency: z.string().length(3) });
type Hotel = z.infer<typeof Hotel>;          // one source of truth
const hotel = Hotel.parse(await res.json()); // throws on drift, at the edge

Trade-offs. Zod adds bundle weight (Valibot is lighter, ArkType faster) and runtime cost per parse. Validate at boundaries only, not on every internal call.

Follow-ups. Where exactly do you validate? How do you keep frontend types in sync with the backend? (Generate from OpenAPI/GraphQL/protobuf in CI so a breaking change fails the build.)

More examples.

// 1 — one schema, one type, parsed once at the edge
const Hotel = z.object({ id: z.string(), price: z.number() });
type Hotel = z.infer<typeof Hotel>;
const hotel = Hotel.parse(await res.json());

// 2 — safeParse when a bad payload should degrade, not crash
const r = Hotel.safeParse(raw);
if (!r.success) { report(r.error); return fallback; }

// 3 — the boundaries people forget
const flags  = Flags.parse(JSON.parse(localStorage.getItem('flags') ?? '{}'));
const params = Search.parse(Object.fromEntries(new URLSearchParams(location.search)));
window.addEventListener('message', (e) => Msg.parse(e.data));  // postMessage is a boundary too
The backend renames a field. Where should that break?

Answer: at build time in CI, not at runtime in a user's browser.

Three layers, and a mature setup has all three:

  1. Generated types — derive frontend types from the backend's OpenAPI or GraphQL schema in CI. A rename now fails the frontend build, before merge.
  2. Runtime parsing at the boundary — because generated types still assume the deployed server matches the schema you generated from. It may not, during a rollout.
  3. Contract tests — the consumer publishes what it expects; the provider's CI replays it. This catches the change in the backend's pipeline, which is where it is cheapest to fix.

Types alone are not enough: they vanish at runtime and describe what you were promised, not what arrived. Parsing alone is not enough either: it finds the problem in production. You want the schema to be the shared artifact.

3.7 Declaration merging and module augmentation

What it is. Interfaces with the same name in the same scope merge; declare module extends a third-party library's types.

Example.

declare module 'styled-components' {
  export interface DefaultTheme extends AppTheme {}
}

Trade-offs. Powerful for design-system theming, but global augmentation is invisible action-at-a-distance and can conflict between packages.

Follow-ups. How do you type a library that ships no types? What goes in exports conditions (import, require, types, browser) when publishing a package?

3.8 TypeScript at monorepo scale

Problem Lever
Slow tsc Project references + composite + incremental; build only changed projects
Editor lag skipLibCheck, shallower conditional types, smaller unions
Slow CI Type-check per package in parallel; transpile with esbuild/SWC separately from tsc --noEmit
Published types drift Generate .d.ts from source; run publint and arethetypeswrong in CI
Breaking a shared type @deprecated JSDoc, a codemod, an overload kept for one minor

Strictness ratchet. You cannot flip strict: true on a legacy codebase in one PR. Enable one flag at a time (noImplicitAny, then strictNullChecks), write existing failures to a baseline file, fail CI only on new violations, and burn the baseline down.

Follow-ups. How would you migrate a 500k-line JS codebase to TypeScript? What do you do about the long tail of @ts-expect-error?

4. Browser internals

4.1 The critical rendering path

Only transform and opacity skip the main thread Critical rendering path · and where each kind of change re-enters it MAIN THREAD COMPOSITOR + GPU HTML → DOM streaming parser CSS → CSSOM render-blocking Render tree minus display:none Layout geometry Paint into layers Composite GPU width · height · top · margin font-size · display color shadow · radius transform · opacity filter on a layer re-enters at LAYOUT — most expensive PAINT only COMPOSITE only — 60 fps safe
The critical rendering path. Only composite-stage changes stay off the main thread.

What it is. The sequence from bytes to pixels: network → HTML parse to DOM → CSS parse to CSSOM → render tree → layout → paint → composite.

Example. Script loading changes the path: a classic <script> blocks the parser; defer runs after parsing, before DOMContentLoaded, in document order; async runs whenever it arrives, out of order. CSS is render-blocking by default — the browser will not paint without the CSSOM.

Trade-offs. Inlining critical CSS removes a round trip but bloats the HTML and loses caching. Inlining a script can hurt, because the preload scanner can no longer discover and fetch subresources ahead of the blocked parser.

Follow-ups. What blocks first paint? Where should <script> go and why? What is the preload scanner? What's the difference between DOMContentLoaded and load?

4.2 Layout, paint, composite — what each change costs

What it is. A style change re-enters the pipeline at a stage determined by the property changed.

Change Triggers Properties
Layout layout + paint + composite width, height, top, margin, font-size, display
Paint paint + composite color, background-color, box-shadow, border-radius
Composite composite only transform, opacity, filter (on a promoted layer)

Example. Animate transform: translateX() rather than left, and opacity rather than visibility. Promote with will-change: transform immediately before the animation and remove it after.

Trade-offs. Each promoted layer costs GPU memory; leaving will-change on many elements causes layer explosion and can make things slower than not promoting at all.

Follow-ups. Which properties are cheap to animate and why? What does promoting a layer actually do? How would you debug a janky animation in DevTools? (Rendering panel → paint flashing, layer borders, FPS meter.)

More examples.

/* 1 — same visual motion, three very different costs */
.a { left: 100px; }                  /* layout → paint → composite */
.b { background-color: red; }        /* paint → composite */
.c { transform: translateX(100px); } /* composite only */

/* 2 — promote only for the duration of the animation, then release */
.card.animating { will-change: transform; }

/* 3 — stop a widget's reflows escaping into the host page */
.widget { contain: layout paint; content-visibility: auto;
          contain-intrinsic-size: 0 420px; }
It animates smoothly in isolation but janks in the real page. Why?

Answer: three candidates, in the order worth checking.

  1. Layer explosion — will-change left permanently on many elements. Each layer costs GPU memory, and past a threshold the compositor spends more time managing layers than it saves. Check Rendering → Layer borders.
  2. It is not actually composited — something in the real page forces it back to the main thread: an ancestor filter, a box-shadow animating alongside, or a property you assumed was cheap.
  3. Main-thread contention — the animation is fine, but long tasks elsewhere starve the frame budget.

The test that separates them: deliberately block the main thread with a long task while the animation runs. If it keeps going, it is genuinely composited and your problem is elsewhere. If it freezes, it never was.

4.3 Reflow and layout thrashing

Interleaving reads and writes forces one synchronous layout per iteration Same work, same result — the only difference is the order Interleaved — n forced reflows boxes.forEach(b => { b.style.width = b.offsetWidth + 10 + 'px'; }); read LAYOUT write read LAYOUT write read LAYOUT write … × n each write dirties style, so the next read must recompute geometry Batched — one reflow const w = boxes.map(b => b.offsetWidth); // read phase boxes.forEach((b,i) => b.style.width = w[i] + 10 + 'px'); // write phase read · read · read · read LAYOUT write · write · write · write style stays clean through the read phase Properties that force layout when style is dirty offsetTop/Left/Width/Height · scrollTop/Height · clientWidth/Height getComputedStyle() · getBoundingClientRect() · focus() · scrollIntoView() Better than batching by hand ResizeObserver and IntersectionObserver deliver measurements off the critical path entirely.
Layout thrashing: n forced reflows, versus one when reads and writes are batched.

What it is. Reading a geometry property while style is dirty forces a synchronous layout. Interleaving reads and writes in a loop produces O(n) forced reflows.

Example.

// BAD — read/write/read/write, n forced synchronous layouts
boxes.forEach(b => { b.style.width = b.offsetWidth + 10 + 'px'; });

// GOOD — batch reads, then batch writes
const widths = boxes.map(b => b.offsetWidth);
boxes.forEach((b, i) => { b.style.width = widths[i] + 10 + 'px'; });

Forcing properties: offsetTop/Left/Width/Height, scrollTop/Height, clientWidth/Height, getComputedStyle(), getBoundingClientRect(), focus(), scrollIntoView().

Trade-offs. A read/write scheduler (FastDOM-style) fixes it generically but adds a frame of latency and indirection. ResizeObserver/IntersectionObserver deliver measurements off the critical path and are usually the better answer.

Follow-ups. Why is getBoundingClientRect() expensive? How does DevTools surface a forced reflow? Where would you put the fix — the component or the framework?

4.4 Main thread vs compositor thread

A tap that lands during a long task waits for it to finish The main thread is single-threaded: input cannot be handled until the current task returns One 300 ms task filter 10,000 results — uninterruptible handler paint user taps here input delay ≈ 330 ms — this is what INP measures Chunked, yielding between pieces handler paint remaining chunks resume after the user is served same tap input delay ≈ 30 ms — one chunk, not the whole job How to yield await scheduler.yield() · scheduler.postTask with a priority · setTimeout(…, 0) Or get off the thread entirely Web Worker for parsing and indexing · useTransition to mark work interruptible
Why a tap during a long task feels slow, and what yielding changes.

What it is. The main thread runs JavaScript, style, layout and paint. The compositor thread assembles layers, often on the GPU, and can scroll and animate independently.

Example. Escape hatches from the main thread:

Tool Use
Web Worker Parsing large JSON, search indexing, image processing. No DOM access; postMessage structured clone, or zero-copy via Transferable/SharedArrayBuffer
requestIdleCallback Non-urgent work in spare frame time
scheduler.postTask Explicit priorities: user-blocking, user-visible, background
Yielding Chunk work and await scheduler.yield() between chunks so input is handled
content-visibility: auto Skip layout and paint for offscreen subtrees — large win on long lists
contain: layout paint size Bound reflow scope inside a widget

Trade-offs. Workers buy parallelism at the cost of serialisation and a more complex API (Comlink helps). content-visibility can break find-in-page and anchor scrolling if contain-intrinsic-size is wrong.

Follow-ups. What is a long task and why does 50ms matter? What can't a worker do? How do you move a 300ms JSON parse off the main thread?

4.5 Observers and platform APIs

API Use it for
IntersectionObserver Lazy loading, infinite scroll, impression tracking — no scroll listeners
ResizeObserver Element-level resize, chart sizing — no window.resize
MutationObserver Watching third-party DOM; delivered as a microtask
PerformanceObserver LCP, CLS, INP, long tasks, resource timing — the basis of RUM
BroadcastChannel Cross-tab sync: logout everywhere, cart updates
AbortController One cancellation token for fetch, listeners and subscriptions
View Transitions / Navigation API Native SPA transitions and routing
OffscreenCanvas Canvas rendering inside a worker

Trade-offs. Observers are asynchronous and batched — better for performance, but you cannot read a value synchronously right after a change.

Follow-ups. Implement infinite scroll without a scroll listener. How does AbortController clean up both a fetch and its listeners? What are the rootMargin and threshold options for?

More examples.

// 1 — infinite scroll with no scroll listener at all
const io = new IntersectionObserver(([e]) => e.isIntersecting && loadMore(),
                                    { rootMargin: '400px' });   // prefetch before it is visible
io.observe(sentinel);

// 2 — impression tracking: fire once, when half the card has been seen
new IntersectionObserver((es) => es.forEach((e) => {
  if (e.isIntersecting) { track(e.target.dataset.id); io2.unobserve(e.target); }
}), { threshold: 0.5 });

// 3 — one controller cancels a fetch and its listeners together
const ac = new AbortController();
fetch(url, { signal: ac.signal });
el.addEventListener('click', onClick, { signal: ac.signal });
ac.abort();   // both gone
Why is IntersectionObserver faster than a scroll handler?

Answer: a scroll handler runs on the main thread, on every scroll event, and almost always calls getBoundingClientRect() — which forces a synchronous layout. Scrolling is exactly when you can least afford that.

IntersectionObserver computes intersections off the main thread, in the compositor, and only calls you when a threshold is actually crossed. Your callback runs a handful of times rather than hundreds, and it receives already-computed geometry, so there is no forced reflow.

The same argument applies to ResizeObserver versus a window.resize handler, and it is why { passive: true } exists for the scroll listeners you genuinely cannot avoid: it promises not to preventDefault, so the compositor need not wait for your handler before scrolling.

4.6 The browser as a security boundary

What it is. Site isolation puts each site in its own renderer process, so a compromised renderer cannot read another site's memory.

Example. This is why SharedArrayBuffer requires COOP and COEP headers, and why window.opener access from target="_blank" links is restricted (rel="noopener").

Trade-offs. Process-per-site costs memory — a visible issue on low-end devices, which matters for emerging-market traffic.

Follow-ups. Why do Spectre mitigations require COOP/COEP? What is the same-origin policy and what does "origin" mean exactly? (scheme + host + port.)

5. Rendering architectures and React

5.1 CSR, SSR, SSG, ISR, streaming, islands, edge

Where the work happens decides when the page is usable Schematic — relative order of events, not measured timings server / build work HTML transfer JS transfer hydration first paint interactive CSRempty shell nothing visible until JS runs SSRper request visible early, but not interactive — the uncanny valley SSG / ISRprebuilt best first paint — but cannot personalise StreamingSSR + Suspense shell first, slow data later, hydration in chunks Islandspartial hydration only the interactive bits ship JS — lowest INP risk, most framework lock-in time →
Schematic order of events per rendering strategy — relative sequence, not measured timings.

What it is. Where the HTML comes from, and when.

Strategy HTML from Best for LCP INP risk Cost
CSR Empty shell + JS Logged-in dashboards, internal tools Poor Medium Cheapest infra, worst first load, SEO workarounds
SSR per request Server, per request Personalised, price-sensitive, SEO pages TTFB-dependent High (hydration) Server CPU per request
SSG Build time Marketing, city landing pages, docs Excellent Low Build time grows with page count
ISR / on-demand revalidate Build + background refresh Large slow-changing catalogues Excellent Low Needs a revalidation and purge story
Streaming SSR + Suspense Server, in chunks Fast shell, slow data Best perceived Medium Streaming-capable runtime
Islands / partial hydration Server HTML + selective JS Content pages with a few interactive bits Excellent Lowest Framework lock-in (Astro, Qwik)
Edge SSR CDN edge node Geo-personalised, low TTFB Excellent Medium Limited runtime, cold starts, DB distance

Example. A hotel search page is best segmented by volatility: static marketing pages SSG; hotel detail ISR; search shell streamed from the edge; prices client-fetched because they must never be cached.

Trade-offs. One global strategy is almost always wrong for a marketplace. SSR improves LCP and SEO but shifts cost to servers and adds hydration cost; SSG is fastest but cannot personalise.

Follow-ups. How would you render a hotel search page and why? What breaks if you cache an SSR page at the CDN? When is CSR the right answer?

5.2 Hydration

What it is. The client re-runs the component tree over server-rendered HTML to attach event listeners and rebuild state.

Example. The ladder of fixes, cheapest first:

  1. Ship less — server-render static parts and don't hydrate them.
  2. Selective/progressive hydration — hydrate above-the-fold and interactive islands first, the rest on IntersectionObserver or on interaction.
  3. Streaming + Suspense boundaries, so hydration happens in chunks rather than one long task.
  4. Resumability (Qwik) — serialise listener state into HTML, no hydration pass at all.
  5. Server Components — whole subtrees never reach the client.

Trade-offs. Hydration is why a server-rendered page can look ready and not respond — the "uncanny valley" that wrecks INP. Avoiding it entirely usually means framework lock-in.

Follow-ups. What causes a hydration mismatch? (Date.now(), locale formatting, window access during render, random ids, UA-dependent markup.) How do you fix a genuinely client-only value? What does React do when markup doesn't match?

More examples.

// 1 — the mismatch nobody expects: locale and clock differ server vs client
<span>{new Date().toLocaleDateString()}</span>

// 2 — the two-pass escape hatch for genuinely client-only values
const [mounted, setMounted] = useState(false);
useEffect(() => setMounted(true), []);
return <span>{mounted ? localTime() : null}</span>;

// 3 — hydrate a heavy island only once it is visible
const Map = lazy(() => import('./Map'));
<Suspense fallback={<MapSkeleton />}>{visible && <Map />}</Suspense>
LCP is 1.8s but INP is 500ms on a server-rendered page. Diagnose it.

Answer: the hydration uncanny valley. The HTML painted fast — hence the good LCP — but the page is not yet interactive, so taps queue behind hydration.

Confirm it: in the field, check whether bad INP clusters in the first seconds after load. In the lab, look for one long task immediately after FCP.

Fixes, cheapest first:

  1. Ship less — server-render static regions and never hydrate them.
  2. Break the single long task — Suspense boundaries hydrate in chunks instead of one uninterruptible pass.
  3. Defer non-critical islands to IntersectionObserver or first interaction.
  4. Move subtrees to Server Components so their code never reaches the client.

What will not help: optimising LCP further. The failing metric is caused by JavaScript execution, not resource loading — saying that distinction out loud is most of the answer.

5.3 Fiber and reconciliation

Render can be thrown away. Commit cannot. That is why render must be pure, and why StrictMode double-invokes it in dev CURRENT TREE — what is on screen App Filters Results Card WORK-IN-PROGRESS TREE App Filters Results Card ×24 reused as-is rebuilt RENDER PHASE — interruptible call components, diff, mark effects can be paused, aborted or restarted keys decide what is reused — index keys attach state to position, not identity then swap current ↔ WIP COMMIT PHASE — one synchronous, uninterruptible pass 1 · mutate the DOM 2 · useLayoutEffect (blocks paint) 3 · browser paints 4 · useEffect (after paint)
React's two trees. The render phase is interruptible; commit is one synchronous pass.

What it is. A fiber is a plain object describing a unit of work with child, sibling and return pointers. React keeps two trees — current and work-in-progress. The render phase is interruptible; the commit phase is one synchronous pass.

Example. Diffing heuristics: different element type → unmount the whole subtree; same type → update props in place; list children matched by key. So index-as-key breaks on reorder — a sorted list of checkboxes keeps the wrong checked state because state attaches to position, not identity.

Trade-offs. O(n) heuristic diffing instead of an optimal O(n³) tree diff: fast, but it relies on you giving stable keys and not changing component types.

Follow-ups. Why must render be pure? Why does StrictMode double-invoke in dev? What happens in commit vs render? When is useLayoutEffect correct over useEffect?

More examples.

// 1 — index keys attach state to a slot, not to an item
{items.map((it, i) => <Row key={i} />)}     // sort → checkbox state stays behind
{items.map((it)    => <Row key={it.id} />)} // sort → state travels with the row

// 2 — changing the element type unmounts the whole subtree
{editing ? <input value={v} /> : <Input value={v} />}   // remounts, focus lost

// 3 — a key change is the deliberate way to reset state
<Form key={userId} />   // switching user clears every field inside
A sorted table keeps the wrong rows checked. Walk through it.

Answer: the rows are keyed by array index.

Before the sort React holds fibers keyed 0,1,2, and fiber 1 carries checked: true. After sorting, the data at index 1 is a different hotel — but the key is still 1, so React matches the old fiber to the new data and reuses its state. The box stays checked, now for the wrong hotel.

With key={hotel.id} React matches by identity: it reorders the existing fibers rather than reusing them positionally, and the checked state travels with its row.

The rule generalises: index keys are safe only when a list is append-only and never reordered, filtered or sorted. Since that is a claim about future code as well as current code, most teams simply ban them.

5.4 Hooks

Hooks are matched by call order, not by name That single fact is the whole reason for the rules of hooks function Search() { const [q, setQ] = useState(''); const [p, setP] = useState(1); const r = useMemo(…); useEffect(…); } fiber.memoizedState → 0 · state '' 1 · state 1 2 · memo 3 · effect Now wrap one in a condition if (q) useMemo(…) ← skipped on the next render 0 · state '' 0 · state '' 1 · state 1 1 · state 1 2 · memo 2 · effect ← wrong slot 3 · effect 3 · missing render 1 render 2 Why the list exists The fiber survives re-renders; the function body does not. The list is where state lives between calls. What actually breaks State from one hook is handed to another. React detects the count change and throws — but the cause is slot drift.
Hooks as an ordered list on the fiber, and what a conditional hook shifts.

What it is. Hooks are a linked list on the fiber, resolved by call order — which is why they cannot be called conditionally.

Hook The nuance
useState Updates are queued and batched; setX(x => x+1) avoids stale closures
useEffect A synchronisation primitive, not a lifecycle. Cleanup runs before every re-run and on unmount
useLayoutEffect Runs before paint — only to measure and correct; warns in SSR
useMemo / useCallback A hint, not a guarantee; has its own cost. Largely unnecessary with React Compiler
useRef Mutable box that doesn't trigger renders; the latest-value escape hatch
useReducer State machines, and when next state depends on prior state
useSyncExternalStore Keeps external stores tear-free under concurrent rendering
useTransition / useDeferredValue Mark updates interruptible so typing stays responsive
useId SSR-safe ids — required for accessible design-system components

Trade-offs. Most useEffects that derive state should be plain computation during render; effects that sync to external systems are the legitimate use.

Follow-ups. Why can't hooks be conditional? When does useEffect run relative to paint? Write a useDebounce. What is a stale closure and how do you avoid it?

5.5 Concurrent rendering

What it is. React 18+ assigns lanes/priorities: discrete input is urgent, transitions are not. Automatic batching now covers promises and timeouts, not just event handlers.

Example. startTransition(() => setFilters(next)) keeps the input responsive while a large result list re-filters.

Trade-offs. A component can render without committing, so side effects in render became actively dangerous; and tearing becomes possible with external mutable stores, which is why useSyncExternalStore exists.

Follow-ups. What is tearing? What's the difference between useTransition and useDeferredValue? What does Suspense actually suspend on?

More examples.

// 1 — keep typing responsive while a large list re-filters
const [q, setQ] = useState('');
const deferred = useDeferredValue(q);         // list renders from the stale value
const rows = useMemo(() => filter(all, deferred), [all, deferred]);

// 2 — mark a navigation as interruptible and show pending state
const [isPending, startTransition] = useTransition();
startTransition(() => setRoute(next));

// 3 — an external store that cannot tear under concurrent rendering
const width = useSyncExternalStore(
  (cb) => { window.addEventListener('resize', cb); return () => window.removeEventListener('resize', cb); },
  () => window.innerWidth,
  () => 1024,            // server snapshot
);
What is tearing, and why can't a plain module variable be the fix?

Answer: tearing is one commit showing two different values of the same external state — the top half of the screen rendered before a change, the bottom half after.

It became possible with concurrent rendering because React can now pause mid-render, let other work run, and resume. If a module variable or an external store mutates during that pause, components rendered before and after the pause read different values, and React commits the inconsistent mix.

A plain module variable has no way to tell React it changed mid-render, so React cannot detect the inconsistency or restart. useSyncExternalStore exists exactly for this: it gives React a getSnapshot to re-read and compare, so React can notice the store moved during the render pass and redo the work consistently.

Component state (useState) is immune, because React controls it and keeps a consistent snapshot per render pass.

5.6 Server Components

What it is. Components that run only on the server, never ship their code to the client, can read data directly, and serialise their output into an RSC payload (a stream, not HTML). Client Components are marked 'use client' and form the boundary.

Example. A hotel detail page can fetch and render description, amenities and policies as RSC, with only the date picker and booking widget as Client Components.

Trade-offs. Smaller client bundles and no client-side data waterfall, versus a new mental model, a server runtime requirement, harder incremental adoption, and payload size for deeply nested trees.

Follow-ups. What can't a Server Component do? How do props cross the boundary? (They must be serialisable.) How does this differ from SSR?

5.7 State management

Most "global state" is server cache wearing a disguise Ask these in order and the answer is usually the first yes Does it come from the server? TanStack Query / SWR / RTK Query caching, dedup, retries, cancellation Should a link to it reproduce the view? The URL — filters, dates, occupancy, sort, page Does only one subtree care? useState / useReducer, colocated as low as it will go Genuinely cross-cutting? Zustand / Jotai — or Context, split by concern yes no Context is not a state manager — every consumer re-renders on any change. Give teams one blessed default, or twelve squads ship twelve patterns.
Where a piece of state belongs — the first yes wins.
Need Tool Why
Server data TanStack Query / SWR / RTK Query Caching, dedup, stale-while-revalidate, retries, cancellation. Most "global state" is this
URL-shaped state The URL Filters, dates, occupancy must be shareable and back-button safe
Small client state Zustand / Jotai / Context Context re-renders all consumers — split by concern or use selectors
Complex workflows Redux Toolkit / XState Middleware, devtools, explicit state machines for a booking funnel
Form state React Hook Form Uncontrolled inputs avoid a render per keystroke

Trade-offs. Keep server cache separate from client state; pick the narrowest tool; give teams one blessed default so twelve squads don't ship twelve patterns.

Follow-ups. Why is Context not a state manager? Where does search-filter state belong and why? How do you prevent a context change re-rendering the whole tree?

5.8 React vs signals

What it is. React re-renders subtrees and diffs a virtual DOM. Vue 3, Svelte 5 and Solid use fine-grained reactivity (signals) that updates only the bound DOM nodes.

Example. In Solid, a component function runs once; the signal read inside a JSX expression creates a direct subscription to that text node.

Trade-offs. Signals do less work at runtime but need a compiler or proxy layer and have a smaller ecosystem. React Compiler attacks the same problem from the other direction by auto-memoising. Web Components remain the interop layer when you must ship a widget into partner sites you don't control.

Follow-ups. When would you not choose React? What does React Compiler change about how you write components? Why is the virtual DOM not inherently fast?

6. Web protocols

6.1 What happens when you type a URL

Every round trip before the first byte is latency you can remove One block = one round trip · the win grows with distance and mobile latency 0 1 RTT 2 3 4 Cold · TCP + TLS 1.3 DNS TCP handshake TLS 1.3 request → HTML first byte after 4 round trips HTTP/3 QUIC over UDP DNS QUIC + TLS together request → HTML transport and crypto handshake collapse into one Resumed 0-RTT request → HTML 0-RTT data is replayable — idempotent requests only preconnect performs DNS + TCP + TLS early, so the first request starts at block 1 instead of block 4.
Round trips before the first byte of HTML, cold vs QUIC vs resumed.

What it is. URL parse → HSTS check → DNS resolution → TCP handshake (or QUIC's single round trip) → TLS handshake, ALPN negotiating h2/h3 → HTTP request → CDN or origin response → parse and render.

Example. Counting round trips is the point: a cold cross-origin image host costs DNS + TCP + TLS before the first byte, which is exactly what preconnect removes.

Trade-offs. Every extra origin costs a connection setup; that's why domain sharding stopped being an optimisation and became a cost.

Follow-ups. Where does the browser cache fit in this sequence? What does HSTS preload change? Why does preconnect help more on mobile?

6.2 DNS

What it is. Hierarchical name resolution: browser cache → OS cache → resolver → root → TLD → authoritative, each hop honouring a TTL.

Example. dns-prefetch resolves a third-party hostname early; preconnect goes further and completes TCP + TLS too.

Trade-offs. A low TTL buys fast failover and costs more lookups. A CNAME chain to your CDN adds a lookup on cold connections.

Follow-ups. What's the difference between dns-prefetch and preconnect? How many preconnect hints should a page have? (A handful — they compete for bandwidth.)

6.3 TCP vs QUIC

What it is. TCP is a reliable, ordered, connection-oriented byte stream with a three-way handshake. QUIC runs over UDP, builds in TLS 1.3, and multiplexes independent streams.

Example. On a lossy mobile network, one lost TCP segment stalls every HTTP/2 stream on that connection; with QUIC only the affected stream stalls.

Trade-offs. QUIC moves congestion control into userspace (more CPU, faster iteration) and some corporate networks block or throttle UDP.

Follow-ups. What is head-of-line blocking, at which layer? Why does connection migration matter when a phone switches from Wi-Fi to cellular?

6.4 TLS

What it is. The handshake negotiates cipher and protocol version, ALPN selects h2/h3, the server presents a certificate chain validated against a trust store, and session keys are derived.

Example. TLS 1.3 completes in one round trip; session resumption and 0-RTT reduce it further.

Trade-offs. 0-RTT data is replayable, so it must only carry idempotent requests. Certificate pinning improves security but causes outages when rotation is mishandled.

Follow-ups. What does Strict-Transport-Security do and why does preload matter? What breaks with mixed content? What is ALPN for?

6.5 HTTP versions

Head-of-line blocking moves down a layer with each HTTP version Six requests on one origin · ✕ marks a lost packet HTTP/1.1 6 TCP connections, one request each at a time Requests 7+ queue behind these. Sharding and sprite sheets made sense here — and only here. HTTP/2 One connection, streams interleaved as frames one TCP connection ↑ ✕ A lost TCP segment stalls EVERY stream — TCP must deliver bytes in order, so all six wait. HOL blocking removed at the application layer, still present at the transport layer. HTTP/3 QUIC over UDP, streams independent ✕ Only the affected stream waits. The other five keep delivering — the mobile win. Also: connection migration survives a Wi-Fi → cellular switch. Consequence: under HTTP/2 and /3, domain sharding is an anti-pattern — many small cacheable chunks are fine.
How each HTTP version handles concurrency, and where a lost packet stalls things.
HTTP/1.1 HTTP/2 HTTP/3
Transport TCP TCP QUIC over UDP
Concurrency \~6 connections per origin Multiplexed streams on one connection Multiplexed, independent streams
Head-of-line blocking At the application layer Removed at app layer, remains at TCP layer Removed
Headers Plaintext, repeated HPACK compression QPACK compression
Handshake TCP + TLS, 2–3 RTT TCP + TLS 1 RTT, 0-RTT on resume

Example. Under HTTP/1.1, sprite sheets, concatenated bundles and domain sharding were correct. Under HTTP/2 sharding is an anti-pattern and many small cacheable chunks are fine.

Trade-offs. H2 Server Push is dead (removed from Chrome) — the replacement is 103 Early Hints plus preload.

Follow-ups. Does HTTP/2 remove the need for bundling? (No — compression ratio and request overhead still favour reasonable chunk sizes.) Why did Server Push fail?

6.6 HTTP methods and idempotency

What it is. Safe means no state change; idempotent means repeating has the same end effect as doing it once.

Method Safe Idempotent Notes
GET, HEAD, OPTIONS Yes Yes Cacheable; never use GET for a mutation — it gets prefetched and logged
PUT No Yes Full replacement
DELETE No Yes Second call returns 404/204, end state matches
POST No No Create or trigger — the retry hazard
PATCH No Not necessarily A merge patch can be; "increment by 1" is not

Example. A booking POST carries a client-generated Idempotency-Key; the server stores key → response for a window, so a retry after a timeout returns the original result instead of double-booking.

Trade-offs. Idempotency is an API contract, not a client property. Retries without a circuit breaker and jittered backoff turn a blip into a retry storm.

Follow-ups. How do you make a POST retry-safe? Where do you put retry policy — the component, the data layer, or the gateway? What does Retry-After do?

More examples.

# 1 — the same booking submitted twice, deduplicated by the server
POST /bookings
Idempotency-Key: 6f3a1c2e-9b7d-4e11-ae02-1d9c4f8b2a77
# a retry with the same key returns the original response, no second booking

# 2 — honour the server's own backoff instruction
HTTP/1.1 429 Too Many Requests
Retry-After: 30
// 3 — retry reads freely, mutations only with a key
const retryable = (method, hasKey) =>
  ['GET', 'HEAD', 'PUT', 'DELETE'].includes(method) || hasKey;
Payment request times out. The user sees a spinner. What now?

Answer: a timeout tells you the response was lost — not that the request was. The charge may well have succeeded. So the one thing you must not do is blindly retry a bare POST.

The correct sequence:

  1. Retry the same request with the same idempotency key. The server either replays the original response (it did complete) or processes it for the first time (it did not). Either way you end up with exactly one charge.
  2. If retries are exhausted, reconcile rather than guess — poll GET /bookings?idempotencyKey=… to discover the true state.
  3. Never show "failed" on a timeout. Show "we're confirming this" and resolve it from the server's answer, because a false failure causes the user to try again manually, which is the worst outcome.

The general principle worth stating: idempotency is a property the server provides; the client can only supply the key that makes it possible.

6.7 Realtime transports

Long polling SSE WebSocket
Direction Request/response Server → client Bidirectional
Protocol HTTP HTTP (text/event-stream) Upgrade from HTTP, then its own framing
Reconnect Manual Built in, with Last-Event-ID Manual
Proxy/CDN friendliness Best Good Often needs special config
Use for Fallback Price pushes, notifications, progress Chat, collaborative editing, trading

Example. Subscribe per visible item rather than per result set — drive subscriptions from IntersectionObserver so offscreen cards unsubscribe; coalesce bursts into one flush per animation frame.

Trade-offs. WebSockets are stateful, which complicates load balancing, scaling and deploys; SSE rides normal HTTP infrastructure but is one-way and limited to text.

Follow-ups. How do you reconnect without hammering the server? (Exponential backoff with jitter, a sequence number to detect gaps, full refetch after a gap.) How do you scale a WebSocket fleet? How do you handle backpressure?

6.8 API shapes: REST, GraphQL, BFF

What it is. REST exposes resources; GraphQL exposes one endpoint with a client-specified query; a BFF is a server owned by the frontend team that aggregates and trims payloads for one surface.

Example. For a travel marketplace, a BFF per surface (web, iOS, Android) keeps the mobile payload small without forcing every backend service to change.

Trade-offs. GraphQL removes over-fetching and client-side waterfalls, and costs you CDN cacheability (POST by default), N+1 risk on the server, and the need for query-cost limits and persisted queries. REST stays trivially cacheable.

Follow-ups. How do you cache a GraphQL response at the edge? Cursor vs offset pagination, and why does offset break on a live inventory? How do you version an API the frontend depends on?

7. Caching, storage and offline

7.1 HTTP cache headers

Three questions pick the cache header Get the second one wrong and a shared cache serves one user's page to another Is the filename content-hashed? Cache-Control: public, max-age=31536000, immutable never revalidates — a new build means a new filename Is it specific to one user? Cache-Control: private, no-store bookings, prices, anything behind auth — never in a shared cache Can it be a few seconds stale? Cache-Control: s-maxage=60, stale-while-revalidate=300 hotel descriptions, reviews, static copy — instant, refreshed behind the scenes Otherwise — the HTML document Cache-Control: no-cache + ETag may be stored, must revalidate — a 304 costs headers, not the body yes no no-cache ≠ no-store: no-cache may be stored and revalidated, no-store is never written to disk at all.
Choosing a Cache-Control policy in three questions.

What it is. Cache-Control sets the policy; ETag/Last-Modified enable conditional revalidation; Vary defines the cache key's extra dimensions.

Directive Meaning Use for
max-age=31536000, immutable Never revalidate for a year Content-hashed static assets
no-cache May store, must revalidate before use HTML documents
no-store Never written to disk Authenticated or PII responses
private / public Whether shared caches may store it private for anything user-specific
stale-while-revalidate=60 Serve stale instantly, refresh in background Content APIs that tolerate seconds of staleness
s-maxage CDN-only lifetime, overrides max-age Short browser TTL, long edge TTL

Example.

GET /app.3f9a1c.js  →  Cache-Control: public, max-age=31536000, immutable
GET /                →  Cache-Control: no-cache
                        ETag: "v842"
   next request:        If-None-Match: "v842"  →  304 Not Modified

Trade-offs. private vs public is a correctness decision, not a performance one — getting it wrong serves one user's booking page to another. Over-varying (e.g. on User-Agent) destroys hit rate.

Follow-ups. What's the difference between no-cache and no-store? What exactly does a 304 save? Strong vs weak ETags? Why must Vary: Accept-Encoding be set?

More examples.

# 1 — hashed asset: cache hard, forever, never revalidate
GET /app.3f9a1c.js
Cache-Control: public, max-age=31536000, immutable

# 2 — the document: always revalidate, cheap when unchanged
GET /
Cache-Control: no-cache
ETag: "v842"
→ If-None-Match: "v842" → 304 Not Modified   (headers only, no body)

# 3 — short browser TTL, long edge TTL, instant-but-fresh
Cache-Control: max-age=0, s-maxage=600, stale-while-revalidate=3600

# 4 — correctness with compression and locales
Vary: Accept-Encoding, Accept-Language
Users report seeing another user's name in the header. One header is wrong.

Answer: a personalised response was served with Cache-Control: public (or simply no private), so a shared cache — the CDN, or a corporate proxy — stored one user's HTML and served it to the next.

The fix is Cache-Control: private, no-store on anything user-specific, but the more useful answer is the architecture that makes this impossible:

  • Keep the cached document anonymous. Cache the shell publicly and fill in personalisation client-side, or at the edge from the session cookie.
  • If you must vary at the edge, put the discriminator in the cache key explicitly rather than relying on Vary with a cookie, which is easy to get subtly wrong and destroys hit rate.
  • Treat public as something you opt into deliberately, never a default.

This is worth rehearsing as an incident story: it is high-severity, it is a one-line cause, and the prevention is a policy rather than a patch.

7.2 The cache-busting pattern

What it is. Immutable, long-lived, content-hashed assets plus a short-lived or revalidated HTML document that points at them.

Example. index.html is no-cache; it references app.3f9a1c.js which is immutable. A deploy changes the hash, so no purge is needed and no user ever gets a mismatched chunk.

Trade-offs. You still must handle a chunk 404 for users who loaded old HTML before a deploy — catch ChunkLoadError, retry with a cache-bust, then force a reload.

Follow-ups. What happens to a user who's been sitting on a tab for an hour when you deploy? How do you roll back without breaking in-flight sessions? (Keep the previous N builds' assets on the CDN.)

7.3 CDN and the cache layer stack

A request stops at the first layer that can answer it Each miss costs latency; each layer you add protects the one behind it Memory this tab only Disk cache Cache-Control Service Worker your code — offline CDN edge s-maxage, Vary CDN shield collapses misses Origin the thing you protect BROWSER NETWORK miss miss miss miss miss response flows back and is written into every layer it passed Cached far from the user Content-hashed JS/CSS/images · immutable, one year Anonymous HTML shell · s-maxage with a purge or hash swap Never cached in a shared layer Prices, availability, anything personalised · Cache-Control: private Get public/private wrong and one user sees another's booking page
Every layer a response can be served from, nearest to the user first.

What it is. A response passes through: browser memory cache → browser disk cache → Service Worker cache → preload cache → CDN edge → CDN shield/regional → origin.

Example. A shield tier collapses many edge misses into one origin request, which protects the origin during a cache purge or a traffic spike.

Trade-offs. Caching personalised HTML at the edge requires the personalisation to be in the cache key, which fragments the cache. The usual resolution is to cache an anonymous shell and personalise client-side or via an edge function.

Follow-ups. What should never be cached at the CDN? How do you invalidate — purge by URL, surrogate key, or just change the hash? What does a cache HIT/MISS/STALE header tell you when debugging?

7.4 Browser storage

Pick storage by who needs the data and when The first yes wins · two of these are synchronous and block the main thread Does the server need it on every request? Cookie — HttpOnly, Secure, SameSite. Costs bytes on every call. Is it a secret, like an access token? In memory only — gone on refresh, which is the point. Silently re-fetch. Large, structured, or needed offline? IndexedDB — async, quota-based. Draft bookings, search index, queues. Should it die with this tab? sessionStorage — per-tab funnel state. Synchronous. Otherwise — small and durable localStorage — prefs, flag cache. Synchronous: never in a hot path.
Choosing a storage mechanism by who needs the data and when.
Mechanism Size Sync? Sent to server Use for
Cookie \~4KB n/a Yes, every request Session/auth (HttpOnly, Secure, SameSite)
localStorage \~5MB Synchronous — blocks the main thread No Small prefs, flag cache. Never tokens
sessionStorage \~5MB Synchronous No Per-tab state, multi-step funnel
IndexedDB Quota-based Async No Offline data, search index, request queue
Cache Storage Quota-based Async No Service Worker asset and response cache
In-memory n/a n/a No Access tokens, anything secret-adjacent

Example. Reading localStorage inside a scroll handler is a genuine performance bug — it is synchronous disk I/O on the main thread.

Trade-offs. Cookies travel on every request, adding bytes to every call; localStorage is convenient but XSS-readable and synchronous; IndexedDB is the only option at size but has an awkward API (use idb).

Follow-ups. Where would you store an auth token and why? What is storage quota and what happens when it's exceeded? How do you sync state between tabs? (BroadcastChannel or a storage event.)

What it is. HttpOnly blocks JavaScript access; Secure restricts to HTTPS; SameSite controls cross-site sending; Domain/Path set scope; Max-Age/Expires set lifetime.

Example. Set-Cookie: sid=...; HttpOnly; Secure; SameSite=Lax; Path=/; Max-Age=1209600

Trade-offs. SameSite=Strict is the safest and breaks inbound links from email and partner sites (the user arrives logged out). Lax is the modern default and allows top-level GET navigation. None requires Secure and is needed for genuine third-party contexts.

Follow-ups. Why does SameSite=Lax mostly solve CSRF? What is a cookie-prefix (__Host-)? How do cookies behave across subdomains, and what breaks when several apps share an apex domain?

7.6 Service Workers and offline

The new worker waits until every old tab is gone Which is why a broken Service Worker is the stickiest bug a frontend team can ship installing waiting new code, not in control activating activated — in control redundant skipWaiting() — takes over immediately unblocks only when all tabs using the old worker close clients.claim() to control pages opened before it The skipWaiting trap An open tab running old JS suddenly gets a new worker serving new, incompatible assets. Prompt the user to reload instead of forcing it. Always ship the kill switch A remote flag that makes the worker call registration.unregister() and Clear-Site-Data as the escape hatch of last resort. Strategy by resource — the worker sits in front of the HTTP cache Hashed assets → cache-first · HTML → network-first with a cached fallback Descriptions and reviews → stale-while-revalidate · prices, availability, payments → network-only
Service Worker lifecycle, and why a bad one is so hard to dislodge.

What it is. A proxy worker sitting between the page and the network, with its own cache, lifecycle and no DOM access.

Example. Lifecycle is install → waiting → activate; skipWaiting and clients.claim take over immediately. Strategy by resource type:

Resource Strategy
Hashed static assets Cache-first
HTML documents Network-first with a cached fallback
Content APIs (descriptions, reviews) Stale-while-revalidate
Prices, availability, payments Network-only

Trade-offs. A Service Worker is the most dangerous thing a frontend team deploys — a bad one is sticky and can serve broken assets indefinitely. Always ship a kill switch (remote flag that makes it unregister) and know Clear-Site-Data. skipWaiting risks mixing an old page with a new worker.

Follow-ups. How do you update a Service Worker safely? What is Background Sync for? How would you make a booking flow survive a tunnel? (Draft in IndexedDB, queued submit with an idempotency key, explicit offline UI — never an optimistic success.)

8. Security

8.1 Same-origin policy and CORS

CORS relaxes the same-origin policy — it does not protect your API A non-simple request asks permission before the real one is sent Browser origin: www.example.com api.example.com different origin OPTIONS /book — the preflight Origin · Access-Control-Request-Method: POST · -Request-Headers: authorization 204 — permission granted Access-Control-Allow-Origin: www.example.com · -Allow-Credentials: true · -Max-Age: 600 POST /book — the real request 200 — and only now may JS read the body Allow-Credentials: true cannot be paired with Origin: * — echo a validated origin, never reflect it blindly. Simple requests (GET, or form-encoded POST) skip the preflight.
A CORS preflight. The browser asks permission before sending the real request.

What it is. An origin is scheme + host + port. The same-origin policy blocks a document from reading cross-origin responses. CORS is the server's mechanism to relax that — it is not a defence.

Example. A non-simple request triggers a preflight:

OPTIONS /api/book          Origin: https://www.example.com
                           Access-Control-Request-Method: POST
                           Access-Control-Request-Headers: content-type, authorization
→ Access-Control-Allow-Origin: https://www.example.com
  Access-Control-Allow-Credentials: true
  Access-Control-Max-Age: 600

Trade-offs. Allow-Credentials: true cannot be combined with Origin: * — you must echo a specific, validated origin. Blindly reflecting Origin is a common vulnerability. A high Max-Age cuts preflights but delays policy changes.

Follow-ups. What makes a request "simple" (no preflight)? Does CORS protect your API? (No — curl ignores it; authorization is the server's job.) Why can a <img> or <form> go cross-origin but fetch cannot read the response?

8.2 XSS

What it is. Executing attacker-controlled script in your origin.

Type How Defence
Stored Malicious input saved then rendered (a hotel review) Encode on output, sanitise on input
Reflected Payload in the URL echoed into the page Context-aware escaping
DOM-based Client-side sink: innerHTML, document.write, eval, location, dangerouslySetInnerHTML Avoid sinks; DOMPurify; Trusted Types

Example. React escapes interpolated values, so the real React vectors are exactly four: dangerouslySetInnerHTML, href={userInput} with a javascript: URL, spreading unvalidated props onto an element, and server-rendered JSON injection (an unescaped </script> inside a serialised payload).

Trade-offs. Sanitising HTML is a losing arms race compared with not rendering HTML at all; Trusted Types converts a code-review problem into a browser-enforced one, at the cost of a migration.

Follow-ups. How would you safely render user-authored rich text? Why is encodeURIComponent not enough in an HTML attribute? What does Trusted Types enforce?

More examples.

// 1 — the four real React vectors
<div dangerouslySetInnerHTML={{ __html: userHtml }} />   // sink
<a href={userUrl}>link</a>                               // javascript: URL
<Component {...untrustedProps} />                        // can inject onError etc.
<script>{`window.__DATA__ = ${JSON.stringify(data)}`}</script>  // </script> breakout

// 2 — safe versions
<a href={/^https?:\/\//.test(userUrl) ? userUrl : '#'}>link</a>
const safe = JSON.stringify(data).replace(/</g, '\\u003c');   // escape the breakout

// 3 — when you genuinely must render HTML
import DOMPurify from 'dompurify';
<div dangerouslySetInnerHTML={{ __html: DOMPurify.sanitize(userHtml) }} />
A hotel review renders fine but fires an alert. Where did it get in?

Answer: it is stored XSS, and the review text reached a sink somewhere that bypasses React's escaping. The three places to look, in order:

  1. dangerouslySetInnerHTML — because the product wanted bold text in reviews.
  2. Server-rendered JSON — the review was interpolated into a <script> tag and contained </script>, which closes the tag early and starts markup.
  3. A third-party widget given the raw string and writing it with innerHTML.

Note what is not the cause: {review.text} in JSX. React escapes that.

The systemic fix is not "sanitise harder" — sanitisers are an arms race. It is Trusted Types (require-trusted-types-for 'script'), which makes the browser throw on any assignment to a dangerous sink unless the value passed through a registered policy. That converts a code-review problem into a runtime guarantee, which is the difference between a fix and a control.

8.3 Content Security Policy

What it is. A response header restricting which sources of script, style, images and frames may load and execute.

Example. A strict modern policy uses nonces plus strict-dynamic, not a host allowlist:

Content-Security-Policy:
  script-src 'nonce-{random}' 'strict-dynamic' https: 'unsafe-inline';
  object-src 'none'; base-uri 'none'; require-trusted-types-for 'script';
  report-uri /csp-report

'unsafe-inline' and https: here are fallbacks for old browsers that modern browsers ignore once a nonce is present.

Trade-offs. Host allowlists are routinely bypassable via JSONP endpoints on allowed CDNs. Nonces are incompatible with full-page CDN caching (each response needs a fresh nonce) — which pushes you to hashes or edge-injected nonces. Inline styles from runtime CSS-in-JS and tag managers are the usual blockers.

Follow-ups. How do you roll out CSP without breaking the site? (Content-Security-Policy-Report-Only, collect violations for two weeks, then enforce.) Why is strict-dynamic better than an allowlist? How does CSP interact with a CDN?

8.4 CSRF

CSRF needs no access to your page — only to the browser that holds your cookie The attacker never reads the response. They only need the write to happen. evil.example the victim's browser bank.example earlier: Set-Cookie: session=… (the user logged in normally) victim visits the attacker's page <form action="bank.example/transfer" method="POST"> auto-submitted POST /transfer — the browser attaches the cookie automatically the server sees a perfectly valid authenticated request SameSite=Lax (today's default) Cookies are not sent on cross-site POSTs, which kills this exact attack. Still allows top-level GET navigation from an email link. Defence in depth A synchroniser or double-submit token the attacker cannot read, plus checking Origin / Sec-Fetch-Site server-side. When it does not apply A bearer token in an Authorization header is not attached automatically — no CSRF, but XSS-readable instead.
A CSRF attack end to end, and the cookie attribute that stops it.

What it is. An attacker-controlled page causes the victim's browser to make an authenticated write using cookies it sends automatically.

Example. Defences, in order of preference: SameSite=Lax/Strict cookies, a synchroniser or double-submit token, and checking Origin/Sec-Fetch-Site server-side.

Trade-offs. CSRF defences are only needed when you authenticate with cookies — a bearer token in an Authorization header is not sent automatically, so it is not CSRF-exposed (but is XSS-exposed instead).

Follow-ups. CORS vs CSRF in one sentence each. Does SameSite=Lax fully solve CSRF? (Not for top-level GET state changes — another reason GET must stay safe.)

8.5 Security headers

Header Purpose
Strict-Transport-Security Force HTTPS; preload covers the first visit
X-Content-Type-Options: nosniff Stop MIME-confusion attacks
Referrer-Policy: strict-origin-when-cross-origin Stop leaking paths and query strings
Permissions-Policy Disable geolocation/camera for embedded third parties
Cross-Origin-Opener-Policy / -Embedder-Policy Process isolation; required for SharedArrayBuffer
Cross-Origin-Resource-Policy Block cross-origin reads of your resources
X-Frame-Options / CSP frame-ancestors Clickjacking

Follow-ups. Which of these would you set first on a new service, and why? What breaks when you turn on COEP?

8.6 Authentication vs authorization

What it is. Authentication is "who are you"; authorization is "what may you do."

Example. Hiding an admin button is a UX affordance. The server must independently enforce the permission on the endpoint.

Trade-offs. Fetching a capabilities object lets the UI render correctly without duplicating policy, at the cost of an extra round trip and a cache-invalidation question when roles change.

Follow-ups. Where must authorization be enforced? How do you handle a user whose permissions change mid-session?

8.7 OAuth, OIDC and JWT

Authorization Code + PKCE — the only correct flow for a browser app The verifier never leaves the client, so an intercepted code is useless SPA (public client) Authorization server Resource API 1 · generate random verifier challenge = S256(verifier) 2 · GET /authorize code_challenge · state · nonce · redirect_uri 3 · redirect back with code verify state matches — this is the CSRF check 4 · POST /token — code + code_verifier server re-hashes the verifier and compares 5 · access token + refresh token access token in memory · refresh in an HttpOnly cookie 6 · Authorization: Bearer … API validates signature, pinned alg, iss, aud, exp Never Implicit flow — leaked tokens into history and Referer Tokens in localStorage · trusting JWT claims client-side Refresh rotation with reuse detection Each refresh issues a new token and kills the old one An old one reappearing = theft → revoke the whole family
Authorization Code with PKCE — the only correct OAuth flow for a browser app.

What it is. OAuth 2.0 is a delegated authorization framework; OpenID Connect is the authentication layer on top of it (the id_token); JWT is merely a token format. An OAuth implementation need not use JWTs.

Example. For a browser app, use Authorization Code with PKCE. The implicit flow is deprecated because it leaked tokens into URL fragments, history and referrers. Validate state (CSRF on the callback) and nonce (id_token replay).

A JWT is header.payload.signature, base64url — encoded, not encrypted. Server-side validation must check the signature, pin the expected alg (reject none and algorithm confusion), and check iss, aud, exp, nbf.

Trade-offs. Stateless JWTs scale well but cannot be revoked before expiry — so either keep access-token lifetimes very short or maintain a denylist and give up some statelessness. Opaque tokens with introspection are the opposite trade.

Follow-ups. Access vs refresh tokens? How does rotation with reuse detection work? (Each refresh issues a new token and invalidates the old; if an old one reappears, revoke the whole family.) What stops ten parallel 401s triggering ten refreshes? (A single-flight lock.)

8.8 Token storage

What it is. The choice between localStorage, cookies and memory for auth material.

Example. The defensible architecture: short-lived access token in memory only, refresh token in an HttpOnly; Secure; SameSite=Strict cookie scoped to the refresh path, silent refresh on 401.

Trade-offs. localStorage is XSS-readable; cookies are CSRF-exposed but can be HttpOnly. The honest caveat: under XSS you lose either way, because the attacker can simply call your API from the victim's session. That is why CSP and Trusted Types matter more than storage choice.

Follow-ups. Why is "just use HttpOnly cookies" not a complete answer? What survives a page refresh with in-memory tokens? (Nothing — you silently refresh on load.)

More examples.

// 1 — access token in memory only, never persisted
let accessToken = null;                  // dies on refresh, by design

// 2 — refresh via an HttpOnly cookie, single-flight so N 401s cause one refresh
let inflight = null;
const refresh = () => (inflight ??= fetch('/auth/refresh', { credentials: 'include' })
  .then(r => r.json())
  .finally(() => { inflight = null; }));

// 3 — the cookie that carries it
// Set-Cookie: rt=…; HttpOnly; Secure; SameSite=Strict; Path=/auth/refresh; Max-Age=1209600
“Just use HttpOnly cookies, then XSS can’t steal the token.” Rebut it.

Answer: HttpOnly stops the attacker reading the token. It does not stop them using it.

With script execution in your origin, the attacker simply calls your API from the victim's own session. The browser attaches the cookie automatically, same-origin checks pass, and every request looks legitimate. They do not need the token's value — they have something better: the ability to act as the user, from the user's browser, for as long as the page is open.

So HttpOnly is worth having (it prevents exfiltration, which enables offline and later abuse), but it is a mitigation, not a boundary.

The honest framing: under XSS you have lost, whatever the storage choice. Storage decisions trade one exposure for another — localStorage is XSS-readable, cookies are CSRF-exposed. Neither is a defence against script injection. That is why CSP and Trusted Types matter more than the storage debate, and saying so is what separates a considered answer from a memorised one.

8.9 Supply chain and third parties

What it is. Risk arriving through dependencies and third-party scripts rather than your own code.

Example. Controls: committed lockfile, npm ci only, Renovate/Dependabot with grouped patch automerge, npm audit/Socket in CI, ignore-scripts plus a private registry proxy, and provenance/sigstore where available. For third-party tags: Subresource Integrity hashes, crossorigin, sandboxed iframes, and a review gate.

Trade-offs. SRI breaks when the vendor updates the file — which is the point, but it means vendor-hosted "always latest" scripts cannot use it. Sandboxing a tag often breaks the feature the business wanted.

Follow-ups. What is dependency confusion and how do scoped private packages prevent it? A marketing tag on your checkout page — what's the blast radius? (Full compromise.) How would you audit what third-party scripts are actually on the page?

8.10 Privacy and compliance on the frontend

What it is. Consent, data minimisation and scope reduction as frontend responsibilities.

Example. Consent management must gate analytics before it loads; PII never goes in URLs (they land in logs and Referer); session-replay tools need field masking; and a payment iframe or hosted fields keeps card data out of your DOM, which removes you from most PCI scope.

Trade-offs. Hosted payment fields reduce compliance scope and cost you styling control and a slightly worse checkout UX.

Follow-ups. How do you keep analytics working under a strict CSP? What would you mask in session replay on a booking form?

9. Performance

9.1 Core Web Vitals

Each vital measures a different moment — and a different failure Measured at the 75th percentile of real users, not in the lab time → NETWORK server + network HTML CSS + JS + hero image TTFB server + network only · diagnostic, not a vital FCP first pixel of anything · blocked by CSS and sync JS LCP — good ≤ 2.5 s largest in-viewport element painted · usually the hero image or H1 MAIN THREAD long tasks ≥ 50 ms sum = TBT (the lab proxy) INP — good ≤ 200 ms a tap landing on a long task: input delay + processing + next paint CLS — good ≤ 0.1 measured across the whole session: images without dimensions, late fonts, injected banners
Where each Core Web Vital is measured during a page's life.

What it is. Three field metrics, measured at the 75th percentile of real users.

Metric Good Needs work Measures Top causes
LCP ≤ 2.5s ≤ 4.0s Render time of the largest in-viewport element Slow TTFB, render-blocking CSS/JS, lazy-loaded hero, unoptimised images
INP ≤ 200ms ≤ 500ms Worst interaction latency: input delay + processing + presentation Long tasks, heavy handlers, large re-render trees, hydration
CLS ≤ 0.1 ≤ 0.25 Sum of unexpected layout shift scores Images without dimensions, injected banners, late fonts, dynamic content above the fold

Example. Supporting diagnostics: TTFB (server + network), FCP (first paint of anything), TBT (sum of long-task time over 50ms), TTI. TBT is the lab proxy; INP is the field metric.

Trade-offs. p75 hides the tail — always look at p95 too, and segment by device class and country, because a global average flatters you while low-end Android users suffer.

Follow-ups. Why did INP replace FID? What's the difference between lab and field data? How do you attribute an INP regression to a specific interaction? (event timing entries and the Long Animation Frames API.)

More examples.

// 1 — attribute a regression to a deploy, not just a route
onLCP(m => beacon({ v: m.value, build: __BUILD_ID__, route, device, country }));

// 2 — find WHICH element is the LCP, in the field
new PerformanceObserver((l) => {
  const e = l.getEntries().at(-1);
  beacon({ lcp: e.startTime, el: e.element?.tagName, url: e.url });
}).observe({ type: 'largest-contentful-paint', buffered: true });

// 3 — attribute INP to the actual interaction
new PerformanceObserver((l) => l.getEntries()
  .filter(e => e.duration > 200)
  .forEach(e => beacon({ inp: e.duration, type: e.name, target: e.target?.id })),
).observe({ type: 'event', durationThreshold: 200, buffered: true });
p75 LCP is 2.4s — green. Users still complain. What did you miss?

Answer: p75 globally is an average of populations, and it hides the segments that are actually suffering.

Slice the same data four ways before believing it:

  1. Device class — a mid-range Android may sit at p75 of 5s while desktop drags the aggregate down. Most travel traffic in emerging markets is low-end Android.
  2. Country / network — a distant origin or a 3G tail behaves nothing like the office.
  3. Route — the home page may be fine while search, where the money is, is not.
  4. p95, not just p75 — the complaining users are, by definition, in the tail.

Then check you are measuring the right thing at all: if complaints are about responsiveness rather than loading, LCP is simply the wrong metric and INP is where to look.

The general lesson: a green aggregate is a hypothesis, not a conclusion. Always ask "green for whom?"

9.2 Measurement

What it is. Lab tools give deterministic, repeatable numbers; field tools (RUM) give the truth about real users.

Example.

import { onLCP, onINP, onCLS } from 'web-vitals';
const send = m => navigator.sendBeacon('/rum', JSON.stringify({
  name: m.name, value: m.value, id: m.id,
  route: currentRoute(), buildId: __BUILD_ID__,   // attribute regressions to a deploy
  conn: navigator.connection?.effectiveType,
}));
onLCP(send); onINP(send); onCLS(send);

Trade-offs. Lighthouse CI is good for gating because it is deterministic, and blind to real device and network diversity. RUM is true but noisy and needs volume before a signal appears — hence synthetic monitoring as the early-warning layer.

Follow-ups. How do you read a Performance panel trace? (Long tasks, forced reflow warnings, the main-thread flame chart, Recalculate Style and Layout blocks.) What is CrUX? Why beacon on visibilitychange rather than unload?

9.3 Loading optimisation

LCP is late because the hero image is discovered late Same bytes, same server — only the discovery order changed Before HTML CSS — render-blocking app.js — parser-blocking JS runs, renders the card hero image — only now requested LCP After HTML critical CSS inlined, rest deferred hero image — preload + fetchpriority="high" LCP app.js, deferred — no longer on the critical path preconnect to the image origin happens here, before the HTML even finishes Order of payoff: inline critical CSS → preconnect to the image origin → preload the LCP image → fetchpriority="high" on the hero. Defer everything else. Never put loading="lazy" on the LCP image — that is the classic self-inflicted regression.
A loading waterfall before and after fixing the critical path.

What it is. Shortening the critical path from request to largest paint.

Example. In rough order of payoff:

  1. Critical CSS inline, rest deferred; per-route CSS extraction.
  2. Resource hints: preconnect to the image/API origin, preload the LCP image and the critical font, fetchpriority="high" on the hero and low below the fold, prefetch the next likely route.
  3. Images: AVIF with WebP fallback, responsive srcset/sizes, explicit width/height or aspect-ratio, loading="lazy" only below the fold, decoding="async", CDN resizing on the fly.
  4. Fonts: font-display: swap (or optional), preload one critical woff2, subset glyph ranges, size-adjust metric overrides so the swap doesn't shift layout.
  5. JavaScript: route-level splitting, then component-level for heavy widgets, then shrink what remains.
  6. Third parties: facade pattern (a fake chat button that loads the real widget on click), iframe sandboxing, or Partytown.

Trade-offs. Preload competes with itself — more than about three hints and you delay the thing you cared about. loading="lazy" on the LCP image is a classic self-inflicted regression.

Follow-ups. What's the difference between preload, prefetch and preconnect? Why can inlining a script hurt? How would you optimise a hotel listing page with forty images?

9.4 Compression

What it is. Brotli (br) beats gzip by roughly 15–20% on text assets; gzip is the universal fallback. Negotiated via Accept-Encoding, which makes Vary: Accept-Encoding mandatory.

Example. Pre-compress static assets at build time with Brotli quality 11 and serve the .br file; use quality 4–6 for dynamic responses so you don't trade TTFB for bytes.

Trade-offs. Never re-compress already-compressed payloads (JPEG, AVIF, WebP, woff2, video) — you burn CPU and can grow the file. Compressing a response containing a secret alongside reflected user input is the BREACH class.

Follow-ups. Why state bundle budgets in compressed bytes? How does compression interact with CDN caching? Why not Brotli images?

9.5 Runtime performance

What it is. Keeping interactions under the INP threshold once the page is loaded.

Example. The standard levers: virtualise long lists (react-window, TanStack Virtual); useDeferredValue/startTransition for filtering while typing; debounce input and throttle scroll; move big JSON parses to a worker; animate on the compositor via CSS or the Web Animations API; memoise deliberately rather than everywhere.

Trade-offs. Memoisation has its own cost (comparison plus retained references) and useMemo everywhere makes code harder to change for no measured gain — prefer state colocation and context splitting first. React Compiler changes this calculus.

Follow-ups. A 10,000-row price calendar is janky — what do you do? Why doesn't React.memo help when you pass an inline object or arrow? What is a long task and how do you break one up?

9.6 Budgets and enforcement

What it is. A per-route limit on transferred bytes and lab metrics, enforced in CI.

Example. Search route: 170KB gzip JS, 50KB CSS, LCP image ≤ 150KB, Lighthouse LCP ≤ 2.5s. Enforce with size-limit or bundlesize; fail the PR and post the delta plus a treemap link as a bot comment.

Trade-offs. Hard budgets block legitimate features; the usual resolution is a documented exception process with an owner and an expiry, not a silently raised threshold.

Follow-ups. Where do you set the budget — total repo or per route? Who owns a budget breach? How do you stop budgets drifting upward over a year?

9.7 A performance debugging workflow

What it is. A repeatable sequence from symptom to systemic fix.

Example.

  1. Quantify in RUM: which metric, which p, which segment, since when.
  2. Correlate with build ids and releases to find a candidate cause.
  3. Reproduce in the lab with matching throttling and device class.
  4. Trace in the Performance panel; find the long task or the blocking resource.
  5. Attribute to a module — bundle analyser treemap, or the LoAF script attribution.
  6. Fix, then verify in the lab, then confirm in RUM at p75 and p95.
  7. Guardrail: a budget, a lint rule, or a CI check so the class of regression cannot return.

Trade-offs. Steps 1–2 are often skipped in favour of jumping to the lab — which finds a problem, rarely the problem.

Follow-ups. Walk me through a real regression you diagnosed. How did you know the fix worked? What stops it recurring?

10. Bundling and build systems

10.1 Module formats

What it is. CommonJS uses dynamic, synchronous require resolved at runtime. ESM uses static import/export, hoisted, with live bindings, resolved before execution.

Example. That static structure is what makes tree shaking, top-level await and import maps possible. CJS cannot be reliably tree-shaken because require('./x')[name] is only knowable at runtime.

Trade-offs. Interop is the pain: "type": "module", dual exports maps, no __dirname in ESM, and packages that ship both. A platform team's job is to publish correct exports conditions (import, require, types, browser, default) and run publint/arethetypeswrong in CI.

Follow-ups. Why can't CJS be tree-shaken? What is the dual-package hazard? What does an import map solve?

10.2 What a bundler does

What it is. Resolve imports from the entry points → load and transform each module → build the module graph → optimise (tree shake, scope hoist, minify, split chunks) → emit content-hashed chunks plus a manifest.

Example. The manifest is what the server or HTML uses to map a route to its chunks — and what makes "build once, promote many" possible.

Trade-offs. More aggressive optimisation means slower builds; most teams run full optimisation only in production builds and transpile-only in dev.

Follow-ups. What is scope hoisting and why does it help? What's in a source map and why hidden-source-map in production?

10.3 Tree shaking

The import that looks identical and ships 200× the code Tree shaking can only drop what the bundler can prove has no side effects Through the barrel import { Button } from '@ds'; @ds/index.ts export * from './button'; export * from './chart'; // + d3 export * from './editor'; // + prosemirror … 200 more all 200 evaluated Side effects cannot be ruled out, so every module is kept — and every one must be parsed first. ~900 KB Direct path import { Button } from '@ds/button'; @ds/button one module Nothing else is in the graph, so nothing else is parsed, bundled or shipped. ~4 KB Other silent breakers: a CJS dependency · transpiling to CJS before bundling · class-property mutation at module scope · a missing "sideEffects": false Fixes: ban deep barrels in lint · publish per-component entry points · declare "sideEffects" honestly · check the treemap in CI, not after launch
Why a barrel import ships 200x the code of a direct one.

What it is. Dead-code elimination over the ESM module graph.

Example. It works only when modules are ESM, the bundler can prove no side effects, and the package declares "sideEffects": false (or lists the files that do have them).

Trade-offs. Silent breakers: re-exporting a whole barrel (export * from './everything'), CJS dependencies, class-property mutation at module scope, and transpiling to CJS before bundling. Barrel files are the number-one cause of bloated frontend bundles — the fix is to ban deep barrels, import from paths, or generate per-component entry points.

Follow-ups. Why did importing one icon pull in 400KB? What does sideEffects actually tell the bundler? How do you prove a dependency is tree-shakeable?

10.4 Code splitting and chunking

Split so the chunks that rarely change stay cached A chunk's hash changing invalidates it for every user — so isolate what churns index.html no-cache framework.a91f.js react, react-dom hash changes ~never commons.5c20.js used by ≥ 3 routes search.3f9a.js hotel.7b1c.js checkout.e2d4.js map.js — on interaction datepicker.js — on focus payment.js — on step 3 LOADED ON EVERY ROUTE ONE PER ROUTE LAZY — React.lazy + Suspense prefetch these on hover or pointer intent, so splitting never costs a visible delay What silently breaks this export * from './everything' — one import pulls the barrel A CJS dependency, or transpiling to CJS before bundling The failure you must handle After a deploy, a stale tab requests a chunk that no longer exists Catch ChunkLoadError → retry cache-busted → force reload
A chunk graph. Stable chunks stay cached; route chunks change independently.

What it is. Breaking the graph into chunks loaded on demand.

Example.

const SearchPage = React.lazy(() => import('./SearchPage'));          // route level
const Map = React.lazy(() => import('./Map'));                        // heavy widget
onMouseEnter={() => import('./CheckoutPage')}                          // prefetch on intent

Chunk strategy: a framework chunk (react, react-dom) that changes rarely so it stays cached, a commons chunk for modules used by ≥N routes, and per-route chunks.

Trade-offs. Over-splitting costs request overhead and deepens the waterfall; under-splitting ships code nobody runs. Splitting also introduces the chunk-404-after-deploy failure, which needs a retry-then-reload handler.

Follow-ups. Where do you split first? How do you avoid a loading spinner on every navigation? Why does a shared chunk's hash changing invalidate everything downstream?

More examples.

// 1 — route split with prefetch on intent, so splitting costs nothing visible
const Checkout = lazy(() => import('./Checkout'));
<Link onMouseEnter={() => import('./Checkout')} onFocus={() => import('./Checkout')} />

// 2 — survive a deploy that removed the chunk you are asking for
window.addEventListener('vite:preloadError', () => location.reload());
// webpack equivalent: catch ChunkLoadError, retry with a cache-bust, then reload

// 3 — keep the framework chunk stable so it stays cached across deploys
// splitChunks: { cacheGroups: { framework: { test: /[\\/]node_modules[\\/](react|react-dom)/ } } }
You split aggressively and it got slower. How?

Answer: several ways, and they compound.

  1. Waterfall depth. Chunk A imports B imports C. Each level costs a round trip, and the browser cannot discover C until B has arrived. Splitting trades bytes for round trips, and on high-latency mobile that trade can lose.
  2. Lost compression ratio. Many tiny chunks compress worse than one larger one — shared dictionary, per-response overhead. Below roughly 20KB a chunk often is not worth its own request.
  3. Duplicated shared code. Without a sensible commons group, the same module gets inlined into several route chunks.
  4. Spinner-per-navigation. Technically faster first load, worse felt performance, because now every click waits.

The fix is not "split less" but "split along the right seams": route level first, then genuinely heavy conditional widgets, with the framework in its own long-lived chunk — and prefetch on intent so the user never waits for a split you chose.

10.5 Tool comparison

Tool Engine Strength Choose when
Webpack 5 JS Largest plugin ecosystem, Module Federation, filesystem cache Large legacy monorepos, federation-heavy setups
Vite esbuild (dev) + Rollup (prod) Native-ESM dev server, near-instant HMR Default for new apps
Rspack / Turbopack Rust Webpack-compatible (Rspack), 5–10× faster Migrating a big Webpack build without a rewrite
esbuild Go Extremely fast transform and bundle Libraries, tooling, dev transforms
Rollup / tsup JS Cleanest library output, multiple formats Publishing design-system packages
SWC / Babel Rust / JS Transform only SWC for speed, Babel when you need AST plugins or codemods

Follow-ups. Why is Vite's dev server fast? (Serves source as native ESM, transforms on demand, pre-bundles deps with esbuild.) Why does Vite still use Rollup for production?

10.6 HMR

What it is. Swapping a module in the running app and propagating the update up the import graph to the nearest boundary that accepts it.

Example. If a change to a context provider full-reloads the page, it is because the boundary is effectively the root — usually a module with side effects or a root-level export change.

Trade-offs. HMR preserves state, which speeds iteration but can mask bugs that only appear on a fresh mount.

Follow-ups. Why does our HMR full-reload every time? What state survives an HMR update and what doesn't?

10.7 Build speed at scale

Lever Effect
Persistent filesystem cache (Webpack 5, Vite) 5–10× on warm local builds
Remote build cache (Turborepo, Nx Cloud, Bazel) CI reuses teammates' and previous runs' outputs
Affected-only builds from the dependency graph PR CI proportional to the change, not the repo
Transpile-only in dev, type-check in a parallel job Removes tsc from the hot loop
hidden-source-map in prod Large emit-time saving, still uploadable to Sentry
Babel → SWC, Webpack → Rspack Usually the single biggest step change

Trade-offs. Remote caching requires deterministic builds — timestamps, absolute paths and non-pinned tool versions all poison the cache.

Follow-ups. Our CI is 35 minutes; get it to 10. What makes a build non-deterministic? How would you measure build time as an SLO?

11. Platform architecture

11.1 Styling approaches

Approach Runtime cost Pros Cons
Plain CSS + BEM None Simple, cacheable Discipline-dependent, no scoping guarantee
CSS Modules None Real scoping, zero runtime Needs CSS variables for dynamic theming
Tailwind / utility None Tiny shipped CSS, no naming debates Verbose markup, needs token discipline
Runtime CSS-in-JS (styled-components, Emotion) Per-render style computation Colocated, dynamic props Runtime overhead, breaks with RSC, hurts INP at scale
Zero-runtime CSS-in-JS (vanilla-extract, Linaria, Panda) None Type-safe tokens, extracted at build Build complexity, less dynamic

Trade-offs. The defensible modern recommendation is CSS variables for tokens plus a zero-runtime or utility layer: it survives Server Components, costs nothing at runtime, and makes theming a CSS concern rather than a React concern.

Follow-ups. Why does runtime CSS-in-JS hurt INP? How would you theme without re-rendering? What breaks when you use Emotion inside a Server Component?

11.2 Design tokens

Three tiers, so a theme switch is one attribute and zero re-renders Components only ever reference tiers 2 and 3 — never a raw value TIER 1 · PRIMITIVE --blue-500: #1a73e8 A raw value with no meaning. Never referenced by a component. TIER 2 · SEMANTIC --color-action-primary What it means, not what it is. This is the layer a theme swaps. TIER 3 · COMPONENT --button-bg-primary The documented override point for a one-off that must differ. A theme is a tier 1 → tier 2 remapping [data-theme="light"] --color-action-primary: var(--blue-500) [data-theme="dark"] --color-action-primary: var(--blue-300) One attribute on <html>. CSS recalculates; React never re-renders. One source, many outputs Tokens live in JSON (or Figma Tokens) and Style Dictionary emits CSS custom properties, TypeScript types, iOS and Android. Lint bans raw hex, so the correct path is also the easy one.
Three token tiers, and why a theme switch costs zero re-renders.

What it is. Three tiers: primitives (--blue-500: #1a73e8), semantic (--color-action-primary: var(--blue-500)), component (--button-bg-primary: var(--color-action-primary)).

Example. Components consume only tiers 2 and 3. Themes swap the tier-1→tier-2 mapping via a data-theme attribute, so a theme switch is one DOM attribute and zero React re-renders. Tokens live in one source (JSON / Figma Tokens) and generate CSS, TS types, iOS and Android outputs via Style Dictionary.

Trade-offs. Three tiers is more indirection than a small team needs; it pays off the moment you have a second theme, a white-label partner, or a native platform.

Follow-ups. How do you add dark mode without touching every component? How do you stop teams using raw hex values? (Lint rule plus a Stylelint config.)

11.3 Component API design

What it is. The rules that make a shared library usable by hundreds of engineers without tickets.

Example.

Trade-offs. Headless libraries save you months of a11y work and add a dependency you must track. Escape hatches improve adoption and make future refactors harder — which is why they should be explicit and documented, not accidental.

Follow-ups. Design the API for a date-range picker. How do you handle a team that needs a one-off variant? How do you forward a ref through a polymorphic component?

11.4 Publishing shared packages

What it is. Internal versioned packages as the distribution mechanism for platform code.

Example. The breaking-change playbook:

  1. Additive change first — new prop, old behaviour unchanged by default.
  2. Deprecate with @deprecated JSDoc, a dev-only console warning, and a lint rule.
  3. Ship a codemod (jscodeshift/ts-morph) and run it across the monorepo yourself.
  4. Publish the major with a migration guide and a compatibility layer for one minor cycle.
  5. Track "packages on latest major" on a dashboard and chase the tail personally.

Trade-offs. Semver discipline slows you down and is the only thing that makes a shared library trustworthy. Steps 3 and 5 are the ones most teams skip, and they are where adoption actually happens.

Follow-ups. How do you prevent two versions of React in a consumer's bundle? (Peer dependencies plus a duplicate check in CI.) How do you version design tokens?

11.5 Monorepo vs polyrepo

Monorepo Polyrepo
Cross-cutting change One atomic PR N coordinated PRs and a release dance
Dependency versions Single version policy, enforced Drift is the default
CI cost Needs affected-graph tooling or it explodes Naturally bounded
Release independence Requires tooling (changesets) Free
Ownership CODEOWNERS paths Repo boundaries
Tooling investment High, centralised, pays back at scale Low per repo, duplicated N times

Trade-offs. Choose by how often changes cross boundaries. If a typical feature touches the design system and two apps, monorepo — and you must fund the build tooling. If teams genuinely ship independently against stable contracts, polyrepo plus published packages is cheaper.

Follow-ups. How do you keep monorepo CI fast? (pnpm, affected-only, remote cache.) What's the failure mode of each? (Unowned shared code — identical in both.)

11.6 Micro-frontends

Independent deploys, one page — and one copy of React Module Federation resolves remotes at runtime, not at build time Host shell routing · auth · layout · error boundaries Search team exposes: ./SearchPage remoteEntry.js · own deploy Account team exposes: ./AccountPage remoteEntry.js · own deploy Payments team exposes: ./Checkout remoteEntry.js · own deploy loaded at runtime shared: { react: { singleton: true, requiredVersion } } Two copies of React is the canonical production incident — hooks throw, context is empty, nothing is debuggable. What you must build before shipping this Error boundary + fallback per remote · version-pinned manifest for rollback · contract tests between host and every remote Centralised regardless: design tokens and components, auth, analytics, routing contract, error reporting, performance budget.
Module Federation at runtime. React must be a shared singleton.

What it is. Independently deployable frontend applications composed into one user-facing surface.

Approach Mechanism Pros Cons
Build-time packages npm versions in one host Simple, optimal bundles No independent deploy
Module Federation Runtime remoteEntry.js + shared scope True independent deploy, shared singletons Version negotiation, runtime failures, hard debugging
Import maps + native ESM Browser-level resolution Standards-based Caching care needed, immature tooling
iframes Hard isolation Perfect CSS/JS isolation, security boundary Routing, sizing, a11y, duplicated runtime
Web Components Custom elements per team Framework-agnostic, works in partner sites Styling/slotting friction, weak SSR story
Server-side composition (ESI, Tailor, Podium) Edge stitches HTML fragments Great first load, SEO-safe Needs edge infra; interactivity still needs a plan

Example. With Module Federation the host declares remotes, the remote declares exposes, and both declare shared with singleton: true and requiredVersion for React — two React copies is the canonical production incident.

Trade-offs. Micro-frontends are justified only when teams need independent deploy cadence and the org boundary is real. The price: duplicated dependencies, version skew, cross-app inconsistency, harder E2E testing, worse aggregate performance, and a much heavier platform team.

Follow-ups. When would you not use micro-frontends? What must stay centralised regardless? (Tokens and components, auth, analytics and experiments, routing contract, error reporting, perf budget.) How do you roll back one remote?

11.7 Many apps, one domain

What it is. Several independently deployed apps served under one hostname via path-based routing at the CDN, reverse proxy or edge function.

Example. /search/* → origin A, /account/* → origin B, via CloudFront behaviours or an edge function. The details that matter:

Trade-offs. Edge routing isolates apps and adds infrastructure and routing-rule complexity that someone must own.

Follow-ups. How do you deploy one app without touching the others? What happens to a user navigating between two apps — full page load or client-side? What breaks with relative asset paths?

11.8 The frontend data layer

What it is. One place that owns request behaviour: auth headers, retries, caching, serialisation, error normalisation and cancellation.

Example. Instead of every component calling fetch differently, a query layer standardises cache keys, loading and error states, dedup, and AbortController wiring for stale searches.

Trade-offs. Not every request should retry — reads yes, non-idempotent mutations only with an idempotency key. Blanket retry policy is how a blip becomes an outage.

Follow-ups. Where should retry policy live? How do you cancel a stale search? How do you normalise errors from three different backends?

11.9 Developer experience

What it is. Treating the local loop and CI as a product with metrics.

Example. The four numbers worth quoting: cold install, dev server start, HMR round-trip, CI wall-clock p95. The five-step adoption playbook for any platform change: prove it on one team with before/after numbers → make it opt-in and trivial → provide the codemod → warn, then enforce (dev warning → CI warning → CI error, with announced dates) → dashboard the tail and close it by pairing.

Trade-offs. Enforcement too early creates resentment; too late and you have permanent drift. The dates are the contract.

Follow-ups. How do you measure DX? (PR cycle time, first-run CI pass rate, flake rate, time-to-first-PR, cache hit rate, plus a survey.) How do you get 12 teams to adopt something without a mandate?

12. Testing

12.1 The shape: trophy, not pyramid

Most frontend bugs are integration bugs — so that is where the mass goes Width ≈ how many tests belong at that layer End to end · 20–50 tests Playwright · revenue-critical journeys only slow, flaky, highest confidence Integration / component · thousands Testing Library + MSW · a component with its children, real DOM, mocked network getByRole queries survive refactors and assert accessibility as a side effect Unit · hundreds Vitest · pure functions, reducers, hooks, formatters Static · TypeScript + ESLint free on every keystroke · catches a whole class of bugs before a test runs cost & confidence per test ↑ cheap, fast, narrow Alongside, not inside, the trophy: contract tests vs the API schema · visual regression per story · axe on every component · Lighthouse CI per route.
The testing trophy — width is roughly how many tests belong at each layer.

What it is. For frontend the useful distribution is a thin base of static analysis, a few unit tests, the bulk in integration/component tests, and a small high-value E2E layer.

Layer Tool Scope Count Runs in
Static TypeScript, ESLint Types, patterns n/a Pre-commit + CI
Unit Vitest / Jest Pure functions, reducers, hooks Hundreds Seconds, on save
Component / integration Vitest + Testing Library, Playwright CT A component with children, real DOM, mocked network Thousands 1–5 min in CI
Contract Pact, or schema diff vs OpenAPI/GraphQL Frontend↔backend payload shape Per endpoint CI, both repos
Visual regression Chromatic / Percy / Playwright screenshots Rendered pixels per state Per story CI on PR
E2E Playwright, Cypress Critical journeys on a real build 20–50, not 500 CI on merge + scheduled
Performance Lighthouse CI, size-limit Budgets per route Per route CI on PR
A11y axe-core in component tests WCAG violations Every component CI on PR

Trade-offs. The classic pyramid came from backend services where unit tests catch most defects. Most frontend bugs are integration bugs — wrong props, wrong state transition, wrong fetch shape — so the mass belongs one layer up.

Follow-ups. Why not more E2E? (Slow, flaky, expensive to maintain, and they fail for reasons unrelated to the change.) What would you test at each layer for a booking form?

12.2 Testing behaviour, not implementation

What it is. Query the DOM the way a user perceives it — by role and accessible name — rather than by class or internal structure.

Example.

// good — survives refactors, and asserts accessibility as a side effect
await userEvent.click(screen.getByRole('button', { name: /search/i }));
expect(await screen.findByRole('list', { name: /results/i })).toBeInTheDocument();

// bad — couples the test to markup
wrapper.find('.btn-primary').simulate('click');

Trade-offs. Role-based queries are slower to write and occasionally force you to fix the markup first — which is the point. Test ids are a legitimate fallback for things with no accessible identity, used sparingly.

Follow-ups. Why is getByRole preferred? When is a test id acceptable? Why are large DOM snapshots worse than no test? (They get rubber-stamped on failure.)

More examples.

// 1 — query the way a user perceives the UI
await userEvent.click(screen.getByRole('button', { name: /search/i }));
expect(await screen.findByRole('list', { name: /results/i })).toBeInTheDocument();

// 2 — assert the absence of a loading state, not an implementation detail
await waitForElementToBeRemoved(() => screen.queryByRole('status'));

// 3 — the escape hatch, used sparingly and deliberately
screen.getByTestId('price-cell');   // only when there is no accessible identity
The test passes but the feature is broken in production. Classic causes?

Answer: the test asserted the implementation rather than the behaviour.

The usual shapes:

  1. Mocked at the module boundary. jest.mock('./useSearch') means the real data layer — the part that broke — never ran. Mock at the network boundary with MSW so the component exercises its real hooks, cache and error handling.
  2. Queried by test id. The element exists but is aria-hidden, covered, or disabled. A real user could not click it; getByRole would have failed.
  3. No assertion on the async settled state. The test asserted the spinner, which always appears, then finished before the failure surfaced.
  4. The bug is in integration. Both units pass in isolation; the contract between them changed. This is precisely why the mass of frontend tests belongs at the integration layer rather than the unit layer.

The diagnostic question worth asking of any test: if the feature broke, would this test fail? A surprising number of green tests cannot answer yes.

12.3 Mocking and MSW

What it is. Intercepting at the network boundary rather than the module boundary, so the component under test runs its real data layer.

Example.

const server = setupServer(
  http.get('/api/search', () => HttpResponse.json({ hotels: [...] })),
  http.get('/api/prices', () => new HttpResponse(null, { status: 500 })), // error path
);

Trade-offs. Mocking the hook instead tests nothing but your mock. MSW handlers are another artefact that can drift from the real API — which is what contract tests are for.

Follow-ups. What's the difference between a stub, a mock and a fake? How do you test a loading state deterministically? How do you keep mocks honest?

12.4 Contract testing

What it is. Verifying that the shape the frontend expects matches what the backend actually produces, without running both end to end.

Example. Consumer-driven contracts (Pact) publish the frontend's expectations; the provider's CI replays them. Alternatively, generate types from OpenAPI/GraphQL in CI so a breaking schema change fails the frontend build.

Trade-offs. Pact adds a broker and process overhead; schema-generated types are cheaper but only catch shape changes, not semantic ones.

Follow-ups. Who owns a broken contract? How do you ship a backward-incompatible API without breaking the web client?

12.5 Visual regression and accessibility testing

What it is. Screenshot diffing per component state, and automated WCAG rule checks.

Example. Every design-system component ships a Storybook story per state, which doubles as both the visual-regression fixture and the axe test target — across themes and both text directions.

Trade-offs. Visual diffs catch what assertions can't and generate noise from font rendering and animation; you need deterministic rendering (frozen clock, disabled animations, pinned browser). axe catches roughly a third of a11y issues — claiming automation covers accessibility is a red flag.

Follow-ups. How do you stop visual tests being flaky? What does axe not catch? (Focus order, meaningful alt text, correct heading structure, keyboard traps in practice.)

12.6 Flaky tests

What it is. Tests that pass and fail on identical code, destroying trust in the suite.

Example. A programme that actually works:

  1. Quarantine, don't ignore — a detected flaky test moves to a suite that still runs but doesn't block, with an auto-filed ticket and an owner from CODEOWNERS.
  2. Measure flake rate per test and per suite, and publish it. A test over the threshold is deleted, not retried forever.
  3. Ban blanket retries at suite level; allow one retry, recorded as a flake signal.
  4. Attack root causes: auto-waiting locators instead of fixed timeouts, per-test data seeded via API not UI, frozen time and network, containerised environment identical to CI.
  5. Gate the gate — if main's suite pass rate drops below \~99%, the pipeline is the incident.

Trade-offs. Deleting a flaky test loses coverage and is often still correct — a test nobody trusts provides zero coverage already while costing wall-clock and attention.

Follow-ups. What causes flakiness in E2E specifically? How do you detect a flaky test automatically? (Re-run on main on a schedule and diff outcomes.)

12.7 Coverage

What it is. The proportion of code executed by the test suite.

Example. Enforce coverage on changed lines in a PR rather than a global percentage. Use mutation testing (Stryker) on critical modules — pricing, currency, date handling — to check whether the tests actually assert anything.

Trade-offs. A global target produces coverage theatre: tests that execute code without asserting. Changed-lines coverage avoids both that and the untested-new-code problem.

Follow-ups. Is 100% coverage a good goal? What does mutation testing measure that line coverage doesn't?

12.8 A testing strategy for a large frontend

What it is. The platform-level answer: own the infrastructure, not the individual tests.

Example. Publish the test utilities as a package — a renderWithProviders wrapper carrying theme, i18n, router and query client; MSW handlers generated from the API schema; a Playwright fixture that logs in via API and seeds data. Every team then starts from the right setup, and improvements propagate with a version bump.

Trade-offs. A shared harness becomes a bottleneck if the platform team is the only one who can change it — hence a contribution model and clear extension points.

Follow-ups. How do you raise testing standards across twelve teams? What do you do about a team that writes no tests? How do you decide what gets an E2E test?

13. Operations

13.1 Pipeline shape

Build the artifact once, then only ever promote it Rebuilding per environment breaks the guarantee that what you tested is what you shipped ON EVERY PULL REQUEST — run in parallel lint + types unit + component a11y (axe) bundle-size diff Lighthouse CI visual diff Preview URL per PR designers and PMs can click the change before it merges Build ONCE immutable artifact hashed chunks + manifest staging canary 1% production same artifact watch RUM by build id staged by geo / device Rollback = flip a flag, or point the CDN at the previous artifact The frontend-specific catch Users mid-session still hold the old HTML and will request old chunks. Keep the previous N builds on the CDN, and handle ChunkLoadError with a cache-busted retry then a reload.
A frontend pipeline. One artifact is built, then promoted; rollback is a pointer flip.

What it is. The sequence of checks on a PR and the promotion path after merge.

Example. On every PR: cached install → lint, type-check and unit/component tests in parallel → build affected packages → bundle-size diff comment → Lighthouse CI on key routes → visual regression → a11y checks → a preview URL per PR. On merge: build once into an immutable artifact and promote that same artifact through environments.

Trade-offs. "Build once, promote many" requires runtime configuration rather than build-time inlining, which is slightly more work and is the only way to guarantee that what you tested is what you shipped.

Follow-ups. Why not rebuild per environment? What goes in a PR check vs a nightly job? How do preview deploys change code review?

13.2 Progressive delivery

Mechanism Gives you Watch out for
Feature flags Deploy decoupled from release; kill switch without a rollback Flag debt — require an expiry date and a cleanup ticket
Canary by percentage Regression caught on 1% of traffic Needs per-cohort metrics or there's no signal
Blue/green at the CDN Instant switch and rollback Both versions' chunks must stay fetchable
Server-side experiment assignment No flicker, no blocking script Must be in the cache key or you serve the wrong variant
Staged rollout by geo/device Limited blast radius Observability must segment the same way

Example. An experiment platform that doesn't regress performance: assign at the edge or in the BFF, pass the assignment into SSR, fire the exposure event only when the component actually renders, add sample-ratio-mismatch alerting, auto-expire experiments, and run a per-experiment performance budget check.

Trade-offs. Client-side assignment is far easier to ship and works with static caching, at the cost of flicker, CLS and INP. Server-side is correct and couples experiments to the render path and the cache key.

Follow-ups. How do you A/B test without a flash of the wrong variant? How do you stop flags accumulating? How would you roll back a frontend change in 60 seconds?

13.3 Frontend rollback

What it is. Reverting the served version without breaking users who are mid-session.

Example. Users holding old HTML will request old chunks. Keep the previous N builds' assets on the CDN (immutable hashed assets make this free) and handle ChunkLoadError with a cache-busted retry then a reload.

Trade-offs. Flipping a feature flag is faster and safer than a deploy rollback, but only covers code that was behind a flag — which is an argument for flagging more than feels necessary.

Follow-ups. What's different about rolling back a frontend vs a backend service? What breaks if you purge the CDN on every deploy?

13.4 Observability

What it is. Knowing what real users experience, and being able to attribute it.

Example.

Trade-offs. RUM is truthful and lagging; synthetic is fast and artificial. You need both, and the cost is two systems to maintain and reconcile.

Follow-ups. How do you attribute an LCP regression to a specific release? What do you do about errors from browser extensions? How much RUM do you sample and why?

13.5 SLOs and error budgets

What it is. A target for a user-facing metric over a window, with a budget for failing it.

Example. "p75 LCP on search ≤ 2.5s over 28 days"; "JS error rate ≤ 0.5% of sessions"; "booking funnel success ≥ 99.5%". Attach a written policy: burn the budget and feature work pauses for reliability work.

Trade-offs. SLOs only work if someone will actually honour the policy; an SLO with no consequence is a dashboard. Setting them too tight produces alert fatigue.

Follow-ups. What would you set as the three SLOs for a frontend platform team? What's the difference between alerting on a threshold and on burn rate?

13.6 Incident handling

What it is. Detect → mitigate → diagnose → review.

Example. Alert on SLO burn rate rather than raw spikes; mitigate before diagnosing (flip the flag, roll back the CDN pointer); diagnose afterwards; run a blameless postmortem whose action item removes the class of failure, not just the instance.

Trade-offs. Mitigating first sometimes destroys the evidence — so capture the trace, the build id and a HAR before you roll back.

Follow-ups. Walk me through a frontend incident you owned. What was the prevention item? How do you know it worked?

14. Accessibility and internationalization

14.1 WCAG essentials

What it is. POUR — perceivable, operable, understandable, robust — with WCAG 2.2 AA as the usual bar.

Example. The concrete AA requirements you'll be asked about: 4.5:1 contrast for body text (3:1 for large text and UI components), a visible focus indicator, 24×24 CSS px minimum target size (44×44 is the stricter AAA/mobile guidance), no information conveyed by colour alone, and 200% zoom without loss of content.

Trade-offs. Meeting contrast ratios constrains brand palettes; the resolution is to bake compliant pairs into semantic tokens so product teams cannot pick a failing combination.

Follow-ups. What's the difference between AA and AAA? How do you handle a brand colour that fails contrast?

14.2 Semantic HTML and ARIA

What it is. Native elements carry role, keyboard behaviour and focus for free. ARIA only adds semantics — it adds no behaviour.

Example. <button> gives focusability, Enter/Space activation and the right role; <div onClick> gives a bug report. The first rule of ARIA is don't use ARIA.

Trade-offs. Custom components sometimes need ARIA because no native element exists (combobox, tabs, tree) — and then you own every keyboard interaction yourself.

Follow-ups. When is aria-label wrong? What does role="presentation" do? Why is aria-hidden on a focusable element a bug?

14.3 Keyboard and focus management

What it is. The hard part of accessibility: what has focus, where it goes, and what is announced.

Example. Focus trap inside a modal, return focus to the trigger on close, aria-live="polite" region for async results ("24 hotels found"), skip links, logical tab order, and no positive tabindex.

Trade-offs. aria-live announcements that fire on every keystroke are worse than none — debounce them and announce results, not progress.

Follow-ups. How do you make an autocomplete accessible? (role="combobox", aria-expanded, aria-controls, aria-activedescendant — the hardest standard pattern.) Where does focus go after deleting a row?

14.4 Testing accessibility

What it is. A layered approach, because no single method is sufficient.

Example. axe-core in component tests catches roughly a third of issues; add keyboard-only test runs, a screen-reader pass (VoiceOver/NVDA) on critical flows, and periodic audits with disabled users.

Trade-offs. Automation is cheap and shallow; manual testing is expensive and real. The platform play is to enforce axe on every design-system component so product teams inherit compliance.

Follow-ups. What can't axe detect? How would you stop accessibility regressing across twelve teams?

14.5 Internationalization architecture

Concern Approach
Message catalogues ICU MessageFormat for plurals, gender and select; keys namespaced by feature; never concatenate sentences
Loading Split translations per route and locale; load the active locale only — shipping 40 locales to every user is a common bundle bug
Dates, numbers, currency Native Intl.DateTimeFormat, Intl.NumberFormat, Intl.RelativeTimeFormat, Intl.PluralRules; Temporal as it lands
Wire format Send ISO/UTC; format on the client in the user's locale and timezone
Currency Never string-template a price — locale controls symbol position, grouping and decimals; conversion server-side with a rate timestamp
Timezones Hotel check-in dates are local calendar dates, not instants — a genuine travel-domain trap
Pseudo-localisation A build-time pseudo-locale ([!!! Ṡéárçh !!!]) to catch hardcoded strings and layout breakage in CI
Text expansion German and Finnish run 30–50% longer than English — never fix widths to English text
RTL CSS logical properties (margin-inline-start, inset-inline), dir="rtl" on <html>, mirrored icons, visual regression in both directions
SEO hreflang alternates per locale, localised URLs, correct lang attribute
Workflow Extract keys in CI, push to the TMS, fail the build on missing keys for launched locales, fall back to English with a logged warning rather than showing the key

Trade-offs. ICU is verbose and is the only sane way to handle languages with six plural forms. Server-side locale detection is accurate and fragments the CDN cache; client-side is cacheable and causes a flash.

Follow-ups. How do you format a price for 40 markets? Why can't you concatenate "Showing " + n + " results"? How would you support RTL in an existing design system? What breaks when a translator writes a string 60% longer?

15. Applying it

15.1 A system-design framework

Use the same eight steps every time, and announce them at the start.

  1. Clarify and scope (5 min). Users and devices, locales and markets, logged-in or anonymous, SEO needed, scale (QPS, catalogue size), and what is explicitly out of scope.
  2. Define success metrics (2 min). Two user metrics (p75 LCP, INP), one business metric (search→book conversion), one engineering metric (build time or deploy frequency).
  3. Sketch the architecture (8 min). Client, CDN/edge, BFF, backend services, third parties. Label arrows with protocol and cache policy.
  4. Pick the rendering strategy (5 min). Segment the page by volatility.
  5. Component hierarchy and state (8 min). Mark which state is URL, server cache, local, global. Name the data-fetching and Suspense boundaries.
  6. Data layer and contract (5 min). Endpoint shapes, pagination, error model, optimistic updates, cache keys and invalidation.
  7. Cross-cutting concerns (7 min). Budget, a11y, i18n, security headers, analytics and experiments, error/empty/loading states, offline.
  8. Trade-offs and next steps (5 min). Two rejected alternatives and why. End with the metric you'd watch after launch.

15.2 Worked scenarios

Hotel search results page. Streaming SSR at the edge for shell and first cards; filters, dates and occupancy in the URL; results in a server cache keyed by the normalised filter object; prices fetched separately and never cached; virtualised list with content-visibility: auto; AVIF with fixed aspect ratios; fetchpriority="high" on the first card image; map library loaded on interaction; filter changes in startTransition with AbortController cancellation; cursor pagination because offset breaks on live inventory. Rejected: client-side filtering (only works for small result sets) and a single global rendering strategy.

Image delivery pipeline. One high-res master; an image CDN derives variants from URL parameters (?w=800&fm=avif&q=70), cached immutably at the edge; a design-system <Image> component generates srcset/sizes and enforces width/height, loading, decoding and fetchpriority; LQIP inside a fixed aspect-ratio box so CLS stays 0. Guardrails: CI fails any raw <img>, a budget on LCP image bytes, RUM reports the LCP element type.

Design system rollout to 12 teams. Phase 0 audit — script the codebase to count distinct button implementations and colour values; that number creates the mandate. Phase 1 tokens only, adopted by codemod. Phase 2 the ten highest-traffic primitives on headless libraries with a11y and visual regression from day one. Phase 3 migration via codemods plus pairing, with a per-team adoption dashboard. Phase 4 enforcement via lint rules. Risks: the library becoming a bottleneck (contribution model, documented escape hatches) and version skew in a federated setup (singleton shared scope, compatibility matrix).

Offline-capable booking flow. Service Worker with precached shell, network-first HTML, network-only for prices; funnel state in IndexedDB keyed by a draft booking id; a request queue with Background Sync and a client-generated idempotency key so a retry cannot double-book; explicit offline UI rather than optimistic success; on reconnect, re-validate price and availability server-side and show a diff; a remote kill switch for the worker.

Real-time price updates. SSE for one-way pushes (simpler, auto-reconnect) or WebSocket if you need client→server too; subscribe per visible card via IntersectionObserver so offscreen items unsubscribe; coalesce bursts into one startTransition flush per animation frame; exponential backoff with jitter plus a sequence number to detect gaps and refetch; animate price changes without moving scroll position, and require confirmation if the price changed between selection and payment.

Experimentation platform. Covered in 13.2 — server-side assignment, variant in the cache key, typed variant names generated from the experiment registry, exposure on render, SRM alerting, auto-expiry, per-experiment performance budget.

15.3 Sources and further reading

MDN Web Docs · web.dev · Refactoring Guru · Awesome Scalability