Frontend Platform Engineering — A Working Reference
A reference for the areas a frontend platform engineer is expected to reason about: web protocols and security, core web platform concepts, performance and bundling, and testing. Every topic follows the same shape — What it is · Example · Trade-offs · Follow-ups — so it can be used for lookup and revision rather than read front to back.
1. How to use this reference
Each topic is short on purpose. The definitions are written to survive a follow-up question, and the trade-off line is the part worth memorising — knowing that a technique exists is cheap, knowing what it costs is the part that transfers.
Three passes. First pass: read the definitions and examples without memorising. Second pass: explain each topic aloud without looking. Third pass: answer the follow-up questions — they are where the real understanding shows.
1.1 The areas this covers
| Area | What it spans | Sections |
|---|---|---|
| Web protocol and security | Browser interactions, security layers, data flow, resource delivery and persistence | 6, 7, 8 |
| Core web platform concepts | Rendering strategies, component models, modern JS, build tools, styling | 2, 3, 4, 5, 11 |
| Performance and bundling | DevTools fluency, root-cause diagnosis, systemic fixes at scale | 9, 10 |
| Testing | Designing strategies and frameworks, raising quality across teams | 12, 13 |
1.2 How each topic is laid out
- What it is — the definition, stated precisely enough to survive a follow-up.
- Example — code, a header, or a config fragment. Concrete, not illustrative prose.
- Trade-offs — what you give up, and the condition under which you'd choose otherwise.
- Follow-ups — the questions that come next. If you can answer these, you know the topic.
2. JavaScript core
2.1 Execution context and the call stack
What it is. Before code runs the engine creates an execution context holding a variable environment, a lexical environment and a this binding. Contexts stack on the call stack; the engine runs one at a time, LIFO.
Example. A stack trace is the call stack printed. RangeError: Maximum call stack size exceeded is that stack overflowing — typically unbounded recursion.
Trade-offs. The stack is synchronous and single-threaded. Anything long-running on it blocks rendering and input, which is the root of every INP problem.
Follow-ups. What's on the stack when a promise resolves? Why does a stack trace lose frames across an await?
More examples.
// 1 — the stack explains why this catch never fires
try {
setTimeout(() => { throw new Error('boom'); }, 0);
} catch (e) {
// unreachable: the callback runs later, on a fresh, empty stack
}
// 2 — but this one does, because await resumes the same logical frame
try {
await Promise.reject(new Error('boom'));
} catch (e) {
// caught
}
// 3 — recursion depth is finite; depth is roughly 10k frames in V8
const depth = (n = 0) => depth(n + 1);
try { depth(); } catch (e) { e.name; } // 'RangeError'
Trace it — what order do the logs print?
function a() { console.log('a'); b(); console.log('a done'); }
function b() { console.log('b'); c(); console.log('b done'); }
function c() { console.log('c'); }
a();
Answer: a · b · c · b done · a done.
Each call pushes a frame and the caller is suspended, not finished. b done
cannot print until c pops, and a done cannot print until b pops. The
"done" lines print in reverse call order — that is the unwinding.
2.2 Hoisting and the temporal dead zone
What it is. Declarations are registered during the creation phase. var is initialised to undefined; function declarations are initialised fully; let, const and class are registered but uninitialised — reading them before the declaration throws. That window is the temporal dead zone.
Example.
console.log(a); // undefined
console.log(b); // ReferenceError — TDZ
var a = 1;
let b = 2;
foo(); // works — declaration hoisted whole
bar(); // TypeError: bar is not a function
function foo() {}
var bar = function () {};
Trade-offs. The TDZ turns a class of silent undefined bugs into loud errors, at the cost of ordering discipline.
Follow-ups. Why does the TDZ exist? Are function declarations hoisted inside blocks? What does typeof b do in the TDZ? (It throws — the one case where typeof is not safe.)
More examples.
// 1 — the loop-variable classic
for (var i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 3 3 3
for (let i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 0 1 2
// 2 — const protects the binding, not the value
const user = { name: 'Ada' };
user.name = 'changed'; // fine — the object is mutable
// user = {}; // TypeError: assignment to constant variable
// 3 — a function declaration inside a block is block-scoped in modules
if (true) { function f() {} }
// f is not reliably visible outside the block — use a const arrow instead
Predict the output
console.log(typeof x);
console.log(typeof y);
var x = 1;
let y = 2;
Answer: 'undefined', then a ReferenceError.
typeof is famously safe on undeclared identifiers — but not inside the
temporal dead zone. y is declared, just not yet initialised, so the engine
throws rather than reporting 'undefined'. This is the one case where typeof
can fail, and it is a common interview trap.
2.3 Scope, closures and the scope chain
What it is. A closure is a function plus a reference to the lexical environment it was created in. Lookup walks the scope chain outward to the global scope.
Example.
for (var i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 3 3 3 — one binding
for (let i = 0; i < 3; i++) setTimeout(() => console.log(i)); // 0 1 2 — per-iteration binding
Trade-offs. Closures are the most common memory leak in long-lived SPAs: a listener closing over a DOM subtree keeps it alive after unmount. Guard with cleanup functions and AbortController.
Follow-ups. How would you implement once, memoise, or a private counter with a closure? How do closures cause detached-DOM leaks?
More examples.
// 1 — private state, no class needed
function counter() {
let n = 0; // unreachable from outside
return { inc: () => ++n, value: () => n };
}
// 2 — run-once, a closure over a boolean
const once = fn => {
let done = false, result;
return (...args) => done ? result : (done = true, result = fn(...args));
};
// 3 — memoise: the cache lives in the closure, not in global scope
const memo = fn => {
const cache = new Map();
return x => cache.has(x) ? cache.get(x) : (cache.set(x, fn(x)), cache.get(x));
};
Why does this leak, and what is the one-line fix?
function attach(node) {
const bigData = new Array(1e6).fill(node.textContent);
node.addEventListener('click', () => console.log(bigData.length));
}
Answer: the listener closes over bigData and over node. Even after the
node is removed from the DOM, the listener keeps both alive — a detached DOM node
plus a million-element array, retained for the life of the page.
The fix is to tie the subscription to a lifetime:
const ac = new AbortController();
node.addEventListener('click', handler, { signal: ac.signal });
// later: ac.abort(); — removes the listener and releases the closure
Note the cause is not "closures are leaky" — it is that nothing ever removed the subscription. The closure is just what made the retention large.
2.4 The event loop
What it is. One turn: take one macrotask → run it to completion → drain the entire microtask queue → run requestAnimationFrame callbacks → style, layout, paint → requestIdleCallback if time remains.
- Macrotasks:
setTimeout,setInterval, I/O, DOM events,postMessage,MessageChannel. - Microtasks: promise continuations,
queueMicrotask,MutationObserver.
Example.
console.log('1 sync');
setTimeout(() => console.log('5 macrotask'), 0);
Promise.resolve().then(() => console.log('3 microtask'));
queueMicrotask(() => console.log('4 microtask'));
console.log('2 sync');
// 1 sync, 2 sync, 3 microtask, 4 microtask, 5 macrotask
Trade-offs. A microtask that queues another microtask starves rendering forever; a setTimeout loop does not (nested timeouts clamp to 4ms after five levels). Microtasks give ordering guarantees at the cost of being unyielding.
Follow-ups. Predict the output of a mixed setTimeout/promise/async snippet. Why does Promise.resolve().then(spin) freeze the tab? Where does requestAnimationFrame run relative to microtasks?
More examples.
// 1 — await is microtask scheduling, so this prints 1 4 2 3
async function f() { console.log(2); await null; console.log(3); }
console.log(1); f(); console.log(4);
// 1 · 2 · 4 · 3 — the body runs synchronously up to the first await
// 2 — rAF runs after microtasks, before paint
Promise.resolve().then(() => console.log('micro'));
requestAnimationFrame(() => console.log('frame'));
// micro · frame
// 3 — yielding so input can be handled between chunks
async function process(items) {
for (const [i, item] of items.entries()) {
work(item);
if (i % 100 === 0) await new Promise(r => setTimeout(r, 0));
}
}
Order these eight logs
console.log('1');
setTimeout(() => console.log('2'), 0);
Promise.resolve().then(() => { console.log('3'); return Promise.resolve(); })
.then(() => console.log('4'));
queueMicrotask(() => console.log('5'));
(async () => { console.log('6'); await null; console.log('7'); })();
console.log('8');
Answer: 1 · 6 · 8 · 3 · 5 · 7 · 4 · 2
1,6,8are synchronous. The async IIFE body runs immediately up to itsawait.- The microtask queue then drains in the order things were queued:
3, then5, then7. 4is late because returning a promise from.thencosts two extra microtask ticks to unwrap — this is the detail almost everyone misses.2is a macrotask, so it runs on the next turn, last.
2.5 Promises and async/await
What it is. await x splits the function: everything after it becomes a microtask continuation. async functions always return a promise.
Example. Serial vs parallel vs bounded concurrency:
for (const id of ids) out.push(await fetch(url(id))); // N × latency
const out = await Promise.all(ids.map(id => fetch(url(id)))); // 1 × latency, N sockets
async function mapLimit(items, limit, fn) { // bounded
const out = [], running = new Set();
for (const [i, item] of items.entries()) {
const p = Promise.resolve(fn(item, i)).then(r => { running.delete(p); return r; });
out.push(p); running.add(p);
if (running.size >= limit) await Promise.race(running);
}
return Promise.all(out);
}
Trade-offs. all fails fast (one rejection discards the rest); allSettled never rejects, right for dashboards and telemetry; any returns the first fulfilled and rejects with AggregateError; race returns the first settled and is the timeout idiom. Unbounded Promise.all over a large list floods the server and the browser's connection pool.
Follow-ups. How do you add a timeout to a fetch? How do you cancel in-flight work? What is an unhandled rejection and how do you report it?
More examples.
// 1 — timeout any promise, using race
const withTimeout = (p, ms) => Promise.race([
p,
new Promise((_, rej) => setTimeout(() => rej(new Error('timeout')), ms)),
]);
// 2 — retry with exponential backoff and jitter
async function retry(fn, tries = 3) {
for (let i = 0; i < tries; i++) {
try { return await fn(); }
catch (e) {
if (i === tries - 1) throw e;
const wait = 2 ** i * 100 + Math.random() * 100; // jitter avoids a thundering herd
await new Promise(r => setTimeout(r, wait));
}
}
}
// 3 — genuinely sequential, when each step needs the previous result
const result = await steps.reduce(
(acc, step) => acc.then(step),
Promise.resolve(initial),
);
What does this log, and why is it a bug?
const results = [];
[1, 2, 3].forEach(async (n) => {
results.push(await double(n));
});
console.log(results.length);
Answer: 0.
forEach ignores the returned promise, so all three callbacks are started and
abandoned — console.log runs before any of them resume. await inside a
callback does not make the outer function wait.
const results = await Promise.all([1, 2, 3].map(double)); // 3
The general rule: forEach is not async-aware. Use map + Promise.all, or a
for…of loop if you genuinely need them sequential.
2.6 this, call/apply/bind, arrow functions
What it is. this is resolved at call time by four rules in priority order: new → explicit call/apply/bind → method receiver → default (undefined in strict mode and modules, globalThis otherwise). Arrow functions have no this binding and close over the enclosing scope.
Example. const f = obj.method; f() loses the receiver; obj.method.bind(obj) or () => obj.method() keeps it.
Trade-offs. Class-field arrows bind per instance — convenient but they cost memory per instance and are not on the prototype, so they cannot be overridden or spied on easily.
Follow-ups. What is this in a standalone function in a module? In a setTimeout callback? In an event handler? Implement bind yourself.
More examples.
// 1 — the receiver is lost the moment you detach the method
const obj = { n: 1, get() { return this.n; } };
const f = obj.get;
obj.get(); // 1 — rule 3, receiver is obj
f(); // throws — rule 4, this is undefined in a module
// 2 — array callbacks take an explicit thisArg
[1, 2].map(function () { return this.k; }, { k: 9 }); // [9, 9]
// 3 — prototype method vs class field, and what each costs
class A {
onClick() {} // on A.prototype — one copy, needs binding
onTap = () => {}; // per instance — pre-bound, costs memory per object
}
Which of these four log 10?
const o = {
n: 10,
regular() { return this.n; },
arrow: () => this?.n,
nested() { return [1].map(function () { return this?.n; })[0]; },
nestedArrow() { return [1].map(() => this.n)[0]; },
};
Answer: regular() and nestedArrow() return 10.
regular— rule 3, the receiver iso.arrow— defined at module top level, so it closed over module scope, wherethisisundefined. The object literal creates no scope.nested— the innerfunctionis called bymapwith no receiver, so rule 4 applies andthisisundefined.nestedArrow— the arrow closes overnestedArrow'sthis, which iso.
The takeaway: an arrow is only useful for this when it is nested inside a
function that already has the this you want.
2.7 Prototypes and classes
What it is. Every object has an internal prototype link (Object.getPrototypeOf(obj) === Ctor.prototype). Property lookup walks that chain. class is syntax over it: methods land on the prototype, class fields on the instance.
Example. Array.prototype.map is found by walking from the array instance to Array.prototype. obj.hasOwnProperty('x') checks only the own object; 'x' in obj walks the chain.
Trade-offs. Prototype sharing saves memory; prototype mutation at runtime deoptimises V8's hidden classes and should be avoided in hot paths. Monkey-patching built-in prototypes breaks other libraries.
Follow-ups. Difference between __proto__ and prototype? How does instanceof work? How would you implement inheritance without class?
More examples.
// 1 — instanceof is a walk, not a type tag
class A {} class B extends A {}
const b = new B();
b instanceof A; // true — A.prototype is on the chain
Object.getPrototypeOf(B.prototype) === A.prototype; // true
// 2 — own vs inherited
const o = Object.create({ inherited: 1 });
o.own = 2;
'inherited' in o; // true — `in` walks the chain
o.hasOwnProperty('inherited'); // false — own properties only
Object.keys(o); // ['own'] — own and enumerable
// 3 — a null-prototype object has no inherited methods at all
const dict = Object.create(null);
dict.toString; // undefined — safe as a plain string map
Why does this break, and what is the safer check?
const data = JSON.parse('{"hasOwnProperty": 1}');
data.hasOwnProperty('x'); // TypeError: not a function
Answer: the parsed object shadows the inherited method with a number, so calling it fails. Any object built from untrusted input can do this.
Safe forms, in order of preference:
Object.hasOwn(data, 'x'); // modern, clearest
Object.prototype.hasOwnProperty.call(data, 'x'); // the classic borrow
The same reasoning is why Object.create(null) is the right choice for a
lookup map — there is nothing on the chain to shadow or to collide with.
2.8 Memory and garbage collection
What it is. V8 uses generational mark-and-sweep: a scavenged young space plus a mark-compact old space. Objects surviving two scavenges are promoted. Anything reachable from a root is retained.
Example. Leak taxonomy: detached DOM nodes, forgotten timers and listeners, unbounded module-level caches, closures over large objects. WeakMap/WeakSet keys do not retain; WeakRef + FinalizationRegistry for caches; AbortController for listener and fetch lifecycle.
Trade-offs. Weak collections prevent leaks but make eviction non-deterministic — you cannot rely on a finalizer running.
Follow-ups. How do you find a leak in DevTools? (Heap snapshot → interact → second snapshot → "objects allocated between 1 and 2" → sort by retained size → follow the retainer path.) What does a sawtooth memory graph that never returns to baseline mean?
More examples.
// 1 — a WeakMap lets the key be collected; a Map does not
const meta = new WeakMap();
meta.set(node, { seen: true }); // when `node` dies, the entry dies with it
// 2 — the three subscriptions that outlive a component
const id = setInterval(tick, 1000); // must clearInterval
window.addEventListener('resize', onResize); // must removeEventListener
const sub = store.subscribe(onChange); // must unsubscribe
// one controller can cover the listeners:
const ac = new AbortController();
window.addEventListener('resize', onResize, { signal: ac.signal });
// 3 — an unbounded module-level cache is a leak with good intentions
const cache = new Map(); // grows forever
const bounded = new Map(); // evict by size or time instead
Two snapshots, same page, 40 MB apart. What do you look at first?
Answer: the Comparison view, filtered to objects allocated between snapshot 1 and snapshot 2, sorted by retained size — not shallow size.
Shallow size is the object itself; retained size is everything that would be freed if it went away. A leak is usually a small object retaining a large graph, so sorting by shallow size hides it.
Then follow the retainer path from the suspect up to a GC root. That path is
the answer: it names the exact reference that must be released. A path ending in
a module-level Map, an array of handlers, or a detached DOM node is the
overwhelmingly common shape.
Quick sanity check before any of that: in the Performance panel, record an interaction loop and look for a sawtooth that never returns to its starting baseline. If memory does return to baseline, you have churn, not a leak.
2.9 The patterns asked by name
| Pattern | What it is | Typical use |
|---|---|---|
| Debounce | Run after a quiet period | Search-as-you-type |
| Throttle | At most once per interval | Scroll, resize, mousemove |
| Event delegation | One listener on a container, dispatch via event.target |
1,000-row tables |
| Capture vs bubble | Root→target, then target→root; {capture: true} |
Intercepting before a child handles it |
passive: true |
Promises not to preventDefault |
Lets scroll run on the compositor |
structuredClone |
Deep clone with cycles, Map, Date |
Replaces JSON.parse(JSON.stringify()) |
| Generators | Lazy sequences via Symbol.iterator |
Pagination, async streams |
Proxy/Reflect |
Intercept property access | Vue 3 reactivity, MobX |
Follow-ups. Implement debounce with a leading edge and cancel(). Why does JSON.parse(JSON.stringify(x)) lose undefined, Date and cycles? When is stopImmediatePropagation needed over stopPropagation?
3. TypeScript
More examples.
// 1 — debounce with a cancel, the version worth memorising
function debounce(fn, ms) {
let t;
const wrapped = (...a) => { clearTimeout(t); t = setTimeout(() => fn(...a), ms); };
wrapped.cancel = () => clearTimeout(t);
return wrapped;
}
// 2 — throttle: trailing edge preserved
function throttle(fn, ms) {
let last = 0, timer;
return (...a) => {
const now = Date.now(), wait = ms - (now - last);
if (wait <= 0) { last = now; fn(...a); }
else { clearTimeout(timer); timer = setTimeout(() => { last = Date.now(); fn(...a); }, wait); }
};
}
// 3 — event delegation: one listener for any number of rows
table.addEventListener('click', (e) => {
const row = e.target.closest('[data-hotel-id]');
if (row && table.contains(row)) select(row.dataset.hotelId);
});
Search-as-you-type: debounce or throttle, and what else?
Answer: debounce — you want one request after the user stops typing, not a steady stream while they type. Throttle is for continuous streams you must sample (scroll, resize, pointermove).
But debounce alone is not enough. Three more things belong in a real implementation:
- Cancel the in-flight request when a newer keystroke supersedes it —
AbortController— otherwise a slow early response can overwrite a fast later one. - Guard against out-of-order responses even so: tag each request and drop any response that is not the latest.
- Mark the update as non-urgent with
useDeferredValueorstartTransition, so re-filtering a large list never blocks the keystroke itself.
The race condition in point 2 is the one candidates usually miss, and it is a real bug class: the user sees results for a query they already edited away.
3.1 Structural typing and branded types
What it is. TypeScript compares types by shape, not by name. Two unrelated types with the same members are interchangeable.
Example. To get nominal behaviour for identifiers:
type HotelId = string & { readonly __brand: 'HotelId' };
type CityId = string & { readonly __brand: 'CityId' };
const asHotelId = (s: string) => s as HotelId;
// passing a CityId where a HotelId is expected is now a compile error
Trade-offs. Branding costs a cast at the boundary and slightly noisier types, and buys you a whole class of ID-mixup bugs caught at compile time.
Follow-ups. Why does TypeScript use structural typing? Where does structural typing surprise people? (Excess property checks apply to object literals only.)
3.2 any vs unknown vs never
What it is. any disables checking and propagates silently. unknown is the safe top type — assignable from anything, assignable to nothing without narrowing. never is the bottom type: no value inhabits it.
Example. never gives exhaustiveness for free:
function render(s: FetchState) {
switch (s.status) {
case 'loading': return spinner();
case 'error': return alert(s.error.message);
case 'success': return list(s.data);
default: { const _never: never = s; return _never; } // errors if a variant is added
}
}
Trade-offs. Banning any outright creates friction at untyped third-party boundaries; the usual policy is unknown plus a parse, with any allowed only behind an explicit lint suppression and a comment.
Follow-ups. Why is unknown safer than any? What does never mean as a return type? How do you type a function that always throws?
3.3 Discriminated unions
What it is. A union of object types sharing a literal-typed discriminant field, which lets the compiler narrow by checking that field.
Example.
type FetchState<T> =
| { status: 'idle' }
| { status: 'loading' }
| { status: 'error'; error: Error }
| { status: 'success'; data: T };
Trade-offs. More verbose than { loading: boolean; error?: Error; data?: T }, but it makes impossible states unrepresentable — you can never have loading: true with data present.
Follow-ups. Model a booking flow's states. How does narrowing work with in, typeof and custom type guards (x is Foo)?
More examples.
// 1 — the discriminant does not have to be called "status"
type Shape =
| { kind: 'circle'; r: number }
| { kind: 'rect'; w: number; h: number };
const area = (s: Shape) => s.kind === 'circle' ? Math.PI * s.r ** 2 : s.w * s.h;
// 2 — a custom type guard narrows across function boundaries
function isError<T>(s: FetchState<T>): s is Extract<FetchState<T>, { status: 'error' }> {
return s.status === 'error';
}
// 3 — a booking funnel where impossible states cannot be constructed
type Booking =
| { step: 'dates' }
| { step: 'rooms'; dates: DateRange }
| { step: 'pay'; dates: DateRange; room: Room }
| { step: 'done'; confirmation: string };
// you cannot reach 'pay' without a room — the type enforces the order
Why is the boolean version worse?
interface State<T> { loading: boolean; error?: Error; data?: T }
Answer: it permits combinations that cannot happen, and forces every reader to
defend against them. { loading: true, error: e, data: d } type-checks perfectly,
so each consumer invents its own precedence rule — and they disagree.
A discriminated union makes illegal combinations unrepresentable, and gives you
exhaustiveness for free: add a 'refetching' variant and every switch that
forgot it fails to compile. With the boolean shape, adding a state means hunting
down every if (loading) by hand and hoping you found them all.
3.4 Generics, conditional and mapped types
What it is. Generics parameterise types; conditional types branch on assignability with infer; mapped types transform every key of a type, optionally remapping keys with as.
Example.
type Unwrap<T> = T extends Promise<infer U> ? Unwrap<U> : T;
type Getters<T> = { [K in keyof T as `get${Capitalize<string & K>}`]: () => T[K] };
// Getters<{ name: string }> === { getName: () => string }
type PolymorphicProps<E extends React.ElementType> =
{ as?: E } & Omit<React.ComponentPropsWithoutRef<E>, 'as'>;
Trade-offs. Deep recursive conditional types are expressive but slow the compiler and produce unreadable errors. A platform rule of thumb: if the type needs a comment to explain it, prefer an explicit overload.
Follow-ups. Implement DeepPartial. What are Partial, Required, Pick, Omit, Record, ReturnType built from? Why do large unions blow up compile time?
3.5 Variance and where TypeScript is unsound
What it is. Function parameters are contravariant and returns covariant. TypeScript deliberately accepts some unsound patterns for ergonomics.
Example. Known unsound spots: arrays are covariant (Dog[] assignable to Animal[], then you can push a Cat); method parameters are bivariant unless strictFunctionTypes applies; index access without noUncheckedIndexedAccess claims T where T | undefined is true; as assertions bypass everything.
Trade-offs. Full soundness would reject large amounts of idiomatic JavaScript. Enabling noUncheckedIndexedAccess is correct but noisy on existing code.
Follow-ups. Where is TypeScript deliberately unsound, and why? What does strictFunctionTypes change?
3.6 Runtime validation at the boundary
What it is. Types are erased at compile time, so anything crossing a boundary — API responses, URL params, localStorage, postMessage, feature flags — must be parsed, not asserted.
Example.
const Hotel = z.object({ id: z.string(), price: z.number(), currency: z.string().length(3) });
type Hotel = z.infer<typeof Hotel>; // one source of truth
const hotel = Hotel.parse(await res.json()); // throws on drift, at the edge
Trade-offs. Zod adds bundle weight (Valibot is lighter, ArkType faster) and runtime cost per parse. Validate at boundaries only, not on every internal call.
Follow-ups. Where exactly do you validate? How do you keep frontend types in sync with the backend? (Generate from OpenAPI/GraphQL/protobuf in CI so a breaking change fails the build.)
More examples.
// 1 — one schema, one type, parsed once at the edge
const Hotel = z.object({ id: z.string(), price: z.number() });
type Hotel = z.infer<typeof Hotel>;
const hotel = Hotel.parse(await res.json());
// 2 — safeParse when a bad payload should degrade, not crash
const r = Hotel.safeParse(raw);
if (!r.success) { report(r.error); return fallback; }
// 3 — the boundaries people forget
const flags = Flags.parse(JSON.parse(localStorage.getItem('flags') ?? '{}'));
const params = Search.parse(Object.fromEntries(new URLSearchParams(location.search)));
window.addEventListener('message', (e) => Msg.parse(e.data)); // postMessage is a boundary too
The backend renames a field. Where should that break?
Answer: at build time in CI, not at runtime in a user's browser.
Three layers, and a mature setup has all three:
- Generated types — derive frontend types from the backend's OpenAPI or GraphQL schema in CI. A rename now fails the frontend build, before merge.
- Runtime parsing at the boundary — because generated types still assume the deployed server matches the schema you generated from. It may not, during a rollout.
- Contract tests — the consumer publishes what it expects; the provider's CI replays it. This catches the change in the backend's pipeline, which is where it is cheapest to fix.
Types alone are not enough: they vanish at runtime and describe what you were promised, not what arrived. Parsing alone is not enough either: it finds the problem in production. You want the schema to be the shared artifact.
3.7 Declaration merging and module augmentation
What it is. Interfaces with the same name in the same scope merge; declare module extends a third-party library's types.
Example.
declare module 'styled-components' {
export interface DefaultTheme extends AppTheme {}
}
Trade-offs. Powerful for design-system theming, but global augmentation is invisible action-at-a-distance and can conflict between packages.
Follow-ups. How do you type a library that ships no types? What goes in exports conditions (import, require, types, browser) when publishing a package?
3.8 TypeScript at monorepo scale
| Problem | Lever |
|---|---|
Slow tsc |
Project references + composite + incremental; build only changed projects |
| Editor lag | skipLibCheck, shallower conditional types, smaller unions |
| Slow CI | Type-check per package in parallel; transpile with esbuild/SWC separately from tsc --noEmit |
| Published types drift | Generate .d.ts from source; run publint and arethetypeswrong in CI |
| Breaking a shared type | @deprecated JSDoc, a codemod, an overload kept for one minor |
Strictness ratchet. You cannot flip strict: true on a legacy codebase in one PR. Enable one flag at a time (noImplicitAny, then strictNullChecks), write existing failures to a baseline file, fail CI only on new violations, and burn the baseline down.
Follow-ups. How would you migrate a 500k-line JS codebase to TypeScript? What do you do about the long tail of @ts-expect-error?
4. Browser internals
4.1 The critical rendering path
What it is. The sequence from bytes to pixels: network → HTML parse to DOM → CSS parse to CSSOM → render tree → layout → paint → composite.
Example. Script loading changes the path: a classic <script> blocks the parser; defer runs after parsing, before DOMContentLoaded, in document order; async runs whenever it arrives, out of order. CSS is render-blocking by default — the browser will not paint without the CSSOM.
Trade-offs. Inlining critical CSS removes a round trip but bloats the HTML and loses caching. Inlining a script can hurt, because the preload scanner can no longer discover and fetch subresources ahead of the blocked parser.
Follow-ups. What blocks first paint? Where should <script> go and why? What is the preload scanner? What's the difference between DOMContentLoaded and load?
4.2 Layout, paint, composite — what each change costs
What it is. A style change re-enters the pipeline at a stage determined by the property changed.
| Change | Triggers | Properties |
|---|---|---|
| Layout | layout + paint + composite | width, height, top, margin, font-size, display |
| Paint | paint + composite | color, background-color, box-shadow, border-radius |
| Composite | composite only | transform, opacity, filter (on a promoted layer) |
Example. Animate transform: translateX() rather than left, and opacity rather than visibility. Promote with will-change: transform immediately before the animation and remove it after.
Trade-offs. Each promoted layer costs GPU memory; leaving will-change on many elements causes layer explosion and can make things slower than not promoting at all.
Follow-ups. Which properties are cheap to animate and why? What does promoting a layer actually do? How would you debug a janky animation in DevTools? (Rendering panel → paint flashing, layer borders, FPS meter.)
More examples.
/* 1 — same visual motion, three very different costs */
.a { left: 100px; } /* layout → paint → composite */
.b { background-color: red; } /* paint → composite */
.c { transform: translateX(100px); } /* composite only */
/* 2 — promote only for the duration of the animation, then release */
.card.animating { will-change: transform; }
/* 3 — stop a widget's reflows escaping into the host page */
.widget { contain: layout paint; content-visibility: auto;
contain-intrinsic-size: 0 420px; }
It animates smoothly in isolation but janks in the real page. Why?
Answer: three candidates, in the order worth checking.
- Layer explosion —
will-changeleft permanently on many elements. Each layer costs GPU memory, and past a threshold the compositor spends more time managing layers than it saves. Check Rendering → Layer borders. - It is not actually composited — something in the real page forces it back
to the main thread: an ancestor filter, a
box-shadowanimating alongside, or a property you assumed was cheap. - Main-thread contention — the animation is fine, but long tasks elsewhere starve the frame budget.
The test that separates them: deliberately block the main thread with a long task while the animation runs. If it keeps going, it is genuinely composited and your problem is elsewhere. If it freezes, it never was.
4.3 Reflow and layout thrashing
What it is. Reading a geometry property while style is dirty forces a synchronous layout. Interleaving reads and writes in a loop produces O(n) forced reflows.
Example.
// BAD — read/write/read/write, n forced synchronous layouts
boxes.forEach(b => { b.style.width = b.offsetWidth + 10 + 'px'; });
// GOOD — batch reads, then batch writes
const widths = boxes.map(b => b.offsetWidth);
boxes.forEach((b, i) => { b.style.width = widths[i] + 10 + 'px'; });
Forcing properties: offsetTop/Left/Width/Height, scrollTop/Height, clientWidth/Height, getComputedStyle(), getBoundingClientRect(), focus(), scrollIntoView().
Trade-offs. A read/write scheduler (FastDOM-style) fixes it generically but adds a frame of latency and indirection. ResizeObserver/IntersectionObserver deliver measurements off the critical path and are usually the better answer.
Follow-ups. Why is getBoundingClientRect() expensive? How does DevTools surface a forced reflow? Where would you put the fix — the component or the framework?
4.4 Main thread vs compositor thread
What it is. The main thread runs JavaScript, style, layout and paint. The compositor thread assembles layers, often on the GPU, and can scroll and animate independently.
Example. Escape hatches from the main thread:
| Tool | Use |
|---|---|
| Web Worker | Parsing large JSON, search indexing, image processing. No DOM access; postMessage structured clone, or zero-copy via Transferable/SharedArrayBuffer |
requestIdleCallback |
Non-urgent work in spare frame time |
scheduler.postTask |
Explicit priorities: user-blocking, user-visible, background |
| Yielding | Chunk work and await scheduler.yield() between chunks so input is handled |
content-visibility: auto |
Skip layout and paint for offscreen subtrees — large win on long lists |
contain: layout paint size |
Bound reflow scope inside a widget |
Trade-offs. Workers buy parallelism at the cost of serialisation and a more complex API (Comlink helps). content-visibility can break find-in-page and anchor scrolling if contain-intrinsic-size is wrong.
Follow-ups. What is a long task and why does 50ms matter? What can't a worker do? How do you move a 300ms JSON parse off the main thread?
4.5 Observers and platform APIs
| API | Use it for |
|---|---|
IntersectionObserver |
Lazy loading, infinite scroll, impression tracking — no scroll listeners |
ResizeObserver |
Element-level resize, chart sizing — no window.resize |
MutationObserver |
Watching third-party DOM; delivered as a microtask |
PerformanceObserver |
LCP, CLS, INP, long tasks, resource timing — the basis of RUM |
BroadcastChannel |
Cross-tab sync: logout everywhere, cart updates |
AbortController |
One cancellation token for fetch, listeners and subscriptions |
| View Transitions / Navigation API | Native SPA transitions and routing |
OffscreenCanvas |
Canvas rendering inside a worker |
Trade-offs. Observers are asynchronous and batched — better for performance, but you cannot read a value synchronously right after a change.
Follow-ups. Implement infinite scroll without a scroll listener. How does AbortController clean up both a fetch and its listeners? What are the rootMargin and threshold options for?
More examples.
// 1 — infinite scroll with no scroll listener at all
const io = new IntersectionObserver(([e]) => e.isIntersecting && loadMore(),
{ rootMargin: '400px' }); // prefetch before it is visible
io.observe(sentinel);
// 2 — impression tracking: fire once, when half the card has been seen
new IntersectionObserver((es) => es.forEach((e) => {
if (e.isIntersecting) { track(e.target.dataset.id); io2.unobserve(e.target); }
}), { threshold: 0.5 });
// 3 — one controller cancels a fetch and its listeners together
const ac = new AbortController();
fetch(url, { signal: ac.signal });
el.addEventListener('click', onClick, { signal: ac.signal });
ac.abort(); // both gone
Why is IntersectionObserver faster than a scroll handler?
Answer: a scroll handler runs on the main thread, on every scroll event,
and almost always calls getBoundingClientRect() — which forces a synchronous
layout. Scrolling is exactly when you can least afford that.
IntersectionObserver computes intersections off the main thread, in the
compositor, and only calls you when a threshold is actually crossed. Your callback
runs a handful of times rather than hundreds, and it receives already-computed
geometry, so there is no forced reflow.
The same argument applies to ResizeObserver versus a window.resize handler,
and it is why { passive: true } exists for the scroll listeners you genuinely
cannot avoid: it promises not to preventDefault, so the compositor need not wait
for your handler before scrolling.
4.6 The browser as a security boundary
What it is. Site isolation puts each site in its own renderer process, so a compromised renderer cannot read another site's memory.
Example. This is why SharedArrayBuffer requires COOP and COEP headers, and why window.opener access from target="_blank" links is restricted (rel="noopener").
Trade-offs. Process-per-site costs memory — a visible issue on low-end devices, which matters for emerging-market traffic.
Follow-ups. Why do Spectre mitigations require COOP/COEP? What is the same-origin policy and what does "origin" mean exactly? (scheme + host + port.)
5. Rendering architectures and React
5.1 CSR, SSR, SSG, ISR, streaming, islands, edge
What it is. Where the HTML comes from, and when.
| Strategy | HTML from | Best for | LCP | INP risk | Cost |
|---|---|---|---|---|---|
| CSR | Empty shell + JS | Logged-in dashboards, internal tools | Poor | Medium | Cheapest infra, worst first load, SEO workarounds |
| SSR per request | Server, per request | Personalised, price-sensitive, SEO pages | TTFB-dependent | High (hydration) | Server CPU per request |
| SSG | Build time | Marketing, city landing pages, docs | Excellent | Low | Build time grows with page count |
| ISR / on-demand revalidate | Build + background refresh | Large slow-changing catalogues | Excellent | Low | Needs a revalidation and purge story |
| Streaming SSR + Suspense | Server, in chunks | Fast shell, slow data | Best perceived | Medium | Streaming-capable runtime |
| Islands / partial hydration | Server HTML + selective JS | Content pages with a few interactive bits | Excellent | Lowest | Framework lock-in (Astro, Qwik) |
| Edge SSR | CDN edge node | Geo-personalised, low TTFB | Excellent | Medium | Limited runtime, cold starts, DB distance |
Example. A hotel search page is best segmented by volatility: static marketing pages SSG; hotel detail ISR; search shell streamed from the edge; prices client-fetched because they must never be cached.
Trade-offs. One global strategy is almost always wrong for a marketplace. SSR improves LCP and SEO but shifts cost to servers and adds hydration cost; SSG is fastest but cannot personalise.
Follow-ups. How would you render a hotel search page and why? What breaks if you cache an SSR page at the CDN? When is CSR the right answer?
5.2 Hydration
What it is. The client re-runs the component tree over server-rendered HTML to attach event listeners and rebuild state.
Example. The ladder of fixes, cheapest first:
- Ship less — server-render static parts and don't hydrate them.
- Selective/progressive hydration — hydrate above-the-fold and interactive islands first, the rest on
IntersectionObserveror on interaction. - Streaming + Suspense boundaries, so hydration happens in chunks rather than one long task.
- Resumability (Qwik) — serialise listener state into HTML, no hydration pass at all.
- Server Components — whole subtrees never reach the client.
Trade-offs. Hydration is why a server-rendered page can look ready and not respond — the "uncanny valley" that wrecks INP. Avoiding it entirely usually means framework lock-in.
Follow-ups. What causes a hydration mismatch? (Date.now(), locale formatting, window access during render, random ids, UA-dependent markup.) How do you fix a genuinely client-only value? What does React do when markup doesn't match?
More examples.
// 1 — the mismatch nobody expects: locale and clock differ server vs client
<span>{new Date().toLocaleDateString()}</span>
// 2 — the two-pass escape hatch for genuinely client-only values
const [mounted, setMounted] = useState(false);
useEffect(() => setMounted(true), []);
return <span>{mounted ? localTime() : null}</span>;
// 3 — hydrate a heavy island only once it is visible
const Map = lazy(() => import('./Map'));
<Suspense fallback={<MapSkeleton />}>{visible && <Map />}</Suspense>
LCP is 1.8s but INP is 500ms on a server-rendered page. Diagnose it.
Answer: the hydration uncanny valley. The HTML painted fast — hence the good LCP — but the page is not yet interactive, so taps queue behind hydration.
Confirm it: in the field, check whether bad INP clusters in the first seconds after load. In the lab, look for one long task immediately after FCP.
Fixes, cheapest first:
- Ship less — server-render static regions and never hydrate them.
- Break the single long task — Suspense boundaries hydrate in chunks instead of one uninterruptible pass.
- Defer non-critical islands to
IntersectionObserveror first interaction. - Move subtrees to Server Components so their code never reaches the client.
What will not help: optimising LCP further. The failing metric is caused by JavaScript execution, not resource loading — saying that distinction out loud is most of the answer.
5.3 Fiber and reconciliation
What it is. A fiber is a plain object describing a unit of work with child, sibling and return pointers. React keeps two trees — current and work-in-progress. The render phase is interruptible; the commit phase is one synchronous pass.
Example. Diffing heuristics: different element type → unmount the whole subtree; same type → update props in place; list children matched by key. So index-as-key breaks on reorder — a sorted list of checkboxes keeps the wrong checked state because state attaches to position, not identity.
Trade-offs. O(n) heuristic diffing instead of an optimal O(n³) tree diff: fast, but it relies on you giving stable keys and not changing component types.
Follow-ups. Why must render be pure? Why does StrictMode double-invoke in dev? What happens in commit vs render? When is useLayoutEffect correct over useEffect?
More examples.
// 1 — index keys attach state to a slot, not to an item
{items.map((it, i) => <Row key={i} />)} // sort → checkbox state stays behind
{items.map((it) => <Row key={it.id} />)} // sort → state travels with the row
// 2 — changing the element type unmounts the whole subtree
{editing ? <input value={v} /> : <Input value={v} />} // remounts, focus lost
// 3 — a key change is the deliberate way to reset state
<Form key={userId} /> // switching user clears every field inside
A sorted table keeps the wrong rows checked. Walk through it.
Answer: the rows are keyed by array index.
Before the sort React holds fibers keyed 0,1,2, and fiber 1 carries
checked: true. After sorting, the data at index 1 is a different hotel — but
the key is still 1, so React matches the old fiber to the new data and
reuses its state. The box stays checked, now for the wrong hotel.
With key={hotel.id} React matches by identity: it reorders the existing fibers
rather than reusing them positionally, and the checked state travels with its row.
The rule generalises: index keys are safe only when a list is append-only and never reordered, filtered or sorted. Since that is a claim about future code as well as current code, most teams simply ban them.
5.4 Hooks
What it is. Hooks are a linked list on the fiber, resolved by call order — which is why they cannot be called conditionally.
| Hook | The nuance |
|---|---|
useState |
Updates are queued and batched; setX(x => x+1) avoids stale closures |
useEffect |
A synchronisation primitive, not a lifecycle. Cleanup runs before every re-run and on unmount |
useLayoutEffect |
Runs before paint — only to measure and correct; warns in SSR |
useMemo / useCallback |
A hint, not a guarantee; has its own cost. Largely unnecessary with React Compiler |
useRef |
Mutable box that doesn't trigger renders; the latest-value escape hatch |
useReducer |
State machines, and when next state depends on prior state |
useSyncExternalStore |
Keeps external stores tear-free under concurrent rendering |
useTransition / useDeferredValue |
Mark updates interruptible so typing stays responsive |
useId |
SSR-safe ids — required for accessible design-system components |
Trade-offs. Most useEffects that derive state should be plain computation during render; effects that sync to external systems are the legitimate use.
Follow-ups. Why can't hooks be conditional? When does useEffect run relative to paint? Write a useDebounce. What is a stale closure and how do you avoid it?
5.5 Concurrent rendering
What it is. React 18+ assigns lanes/priorities: discrete input is urgent, transitions are not. Automatic batching now covers promises and timeouts, not just event handlers.
Example. startTransition(() => setFilters(next)) keeps the input responsive while a large result list re-filters.
Trade-offs. A component can render without committing, so side effects in render became actively dangerous; and tearing becomes possible with external mutable stores, which is why useSyncExternalStore exists.
Follow-ups. What is tearing? What's the difference between useTransition and useDeferredValue? What does Suspense actually suspend on?
More examples.
// 1 — keep typing responsive while a large list re-filters
const [q, setQ] = useState('');
const deferred = useDeferredValue(q); // list renders from the stale value
const rows = useMemo(() => filter(all, deferred), [all, deferred]);
// 2 — mark a navigation as interruptible and show pending state
const [isPending, startTransition] = useTransition();
startTransition(() => setRoute(next));
// 3 — an external store that cannot tear under concurrent rendering
const width = useSyncExternalStore(
(cb) => { window.addEventListener('resize', cb); return () => window.removeEventListener('resize', cb); },
() => window.innerWidth,
() => 1024, // server snapshot
);
What is tearing, and why can't a plain module variable be the fix?
Answer: tearing is one commit showing two different values of the same external state — the top half of the screen rendered before a change, the bottom half after.
It became possible with concurrent rendering because React can now pause mid-render, let other work run, and resume. If a module variable or an external store mutates during that pause, components rendered before and after the pause read different values, and React commits the inconsistent mix.
A plain module variable has no way to tell React it changed mid-render, so React
cannot detect the inconsistency or restart. useSyncExternalStore exists exactly
for this: it gives React a getSnapshot to re-read and compare, so React can
notice the store moved during the render pass and redo the work consistently.
Component state (useState) is immune, because React controls it and keeps a
consistent snapshot per render pass.
5.6 Server Components
What it is. Components that run only on the server, never ship their code to the client, can read data directly, and serialise their output into an RSC payload (a stream, not HTML). Client Components are marked 'use client' and form the boundary.
Example. A hotel detail page can fetch and render description, amenities and policies as RSC, with only the date picker and booking widget as Client Components.
Trade-offs. Smaller client bundles and no client-side data waterfall, versus a new mental model, a server runtime requirement, harder incremental adoption, and payload size for deeply nested trees.
Follow-ups. What can't a Server Component do? How do props cross the boundary? (They must be serialisable.) How does this differ from SSR?
5.7 State management
| Need | Tool | Why |
|---|---|---|
| Server data | TanStack Query / SWR / RTK Query | Caching, dedup, stale-while-revalidate, retries, cancellation. Most "global state" is this |
| URL-shaped state | The URL | Filters, dates, occupancy must be shareable and back-button safe |
| Small client state | Zustand / Jotai / Context | Context re-renders all consumers — split by concern or use selectors |
| Complex workflows | Redux Toolkit / XState | Middleware, devtools, explicit state machines for a booking funnel |
| Form state | React Hook Form | Uncontrolled inputs avoid a render per keystroke |
Trade-offs. Keep server cache separate from client state; pick the narrowest tool; give teams one blessed default so twelve squads don't ship twelve patterns.
Follow-ups. Why is Context not a state manager? Where does search-filter state belong and why? How do you prevent a context change re-rendering the whole tree?
5.8 React vs signals
What it is. React re-renders subtrees and diffs a virtual DOM. Vue 3, Svelte 5 and Solid use fine-grained reactivity (signals) that updates only the bound DOM nodes.
Example. In Solid, a component function runs once; the signal read inside a JSX expression creates a direct subscription to that text node.
Trade-offs. Signals do less work at runtime but need a compiler or proxy layer and have a smaller ecosystem. React Compiler attacks the same problem from the other direction by auto-memoising. Web Components remain the interop layer when you must ship a widget into partner sites you don't control.
Follow-ups. When would you not choose React? What does React Compiler change about how you write components? Why is the virtual DOM not inherently fast?
6. Web protocols
6.1 What happens when you type a URL
What it is. URL parse → HSTS check → DNS resolution → TCP handshake (or QUIC's single round trip) → TLS handshake, ALPN negotiating h2/h3 → HTTP request → CDN or origin response → parse and render.
Example. Counting round trips is the point: a cold cross-origin image host costs DNS + TCP + TLS before the first byte, which is exactly what preconnect removes.
Trade-offs. Every extra origin costs a connection setup; that's why domain sharding stopped being an optimisation and became a cost.
Follow-ups. Where does the browser cache fit in this sequence? What does HSTS preload change? Why does preconnect help more on mobile?
6.2 DNS
What it is. Hierarchical name resolution: browser cache → OS cache → resolver → root → TLD → authoritative, each hop honouring a TTL.
Example. dns-prefetch resolves a third-party hostname early; preconnect goes further and completes TCP + TLS too.
Trade-offs. A low TTL buys fast failover and costs more lookups. A CNAME chain to your CDN adds a lookup on cold connections.
Follow-ups. What's the difference between dns-prefetch and preconnect? How many preconnect hints should a page have? (A handful — they compete for bandwidth.)
6.3 TCP vs QUIC
What it is. TCP is a reliable, ordered, connection-oriented byte stream with a three-way handshake. QUIC runs over UDP, builds in TLS 1.3, and multiplexes independent streams.
Example. On a lossy mobile network, one lost TCP segment stalls every HTTP/2 stream on that connection; with QUIC only the affected stream stalls.
Trade-offs. QUIC moves congestion control into userspace (more CPU, faster iteration) and some corporate networks block or throttle UDP.
Follow-ups. What is head-of-line blocking, at which layer? Why does connection migration matter when a phone switches from Wi-Fi to cellular?
6.4 TLS
What it is. The handshake negotiates cipher and protocol version, ALPN selects h2/h3, the server presents a certificate chain validated against a trust store, and session keys are derived.
Example. TLS 1.3 completes in one round trip; session resumption and 0-RTT reduce it further.
Trade-offs. 0-RTT data is replayable, so it must only carry idempotent requests. Certificate pinning improves security but causes outages when rotation is mishandled.
Follow-ups. What does Strict-Transport-Security do and why does preload matter? What breaks with mixed content? What is ALPN for?
6.5 HTTP versions
| HTTP/1.1 | HTTP/2 | HTTP/3 | |
|---|---|---|---|
| Transport | TCP | TCP | QUIC over UDP |
| Concurrency | \~6 connections per origin | Multiplexed streams on one connection | Multiplexed, independent streams |
| Head-of-line blocking | At the application layer | Removed at app layer, remains at TCP layer | Removed |
| Headers | Plaintext, repeated | HPACK compression | QPACK compression |
| Handshake | TCP + TLS, 2–3 RTT | TCP + TLS | 1 RTT, 0-RTT on resume |
Example. Under HTTP/1.1, sprite sheets, concatenated bundles and domain sharding were correct. Under HTTP/2 sharding is an anti-pattern and many small cacheable chunks are fine.
Trade-offs. H2 Server Push is dead (removed from Chrome) — the replacement is 103 Early Hints plus preload.
Follow-ups. Does HTTP/2 remove the need for bundling? (No — compression ratio and request overhead still favour reasonable chunk sizes.) Why did Server Push fail?
6.6 HTTP methods and idempotency
What it is. Safe means no state change; idempotent means repeating has the same end effect as doing it once.
| Method | Safe | Idempotent | Notes |
|---|---|---|---|
| GET, HEAD, OPTIONS | Yes | Yes | Cacheable; never use GET for a mutation — it gets prefetched and logged |
| PUT | No | Yes | Full replacement |
| DELETE | No | Yes | Second call returns 404/204, end state matches |
| POST | No | No | Create or trigger — the retry hazard |
| PATCH | No | Not necessarily | A merge patch can be; "increment by 1" is not |
Example. A booking POST carries a client-generated Idempotency-Key; the server stores key → response for a window, so a retry after a timeout returns the original result instead of double-booking.
Trade-offs. Idempotency is an API contract, not a client property. Retries without a circuit breaker and jittered backoff turn a blip into a retry storm.
Follow-ups. How do you make a POST retry-safe? Where do you put retry policy — the component, the data layer, or the gateway? What does Retry-After do?
More examples.
# 1 — the same booking submitted twice, deduplicated by the server
POST /bookings
Idempotency-Key: 6f3a1c2e-9b7d-4e11-ae02-1d9c4f8b2a77
# a retry with the same key returns the original response, no second booking
# 2 — honour the server's own backoff instruction
HTTP/1.1 429 Too Many Requests
Retry-After: 30
// 3 — retry reads freely, mutations only with a key
const retryable = (method, hasKey) =>
['GET', 'HEAD', 'PUT', 'DELETE'].includes(method) || hasKey;
Payment request times out. The user sees a spinner. What now?
Answer: a timeout tells you the response was lost — not that the request was. The charge may well have succeeded. So the one thing you must not do is blindly retry a bare POST.
The correct sequence:
- Retry the same request with the same idempotency key. The server either replays the original response (it did complete) or processes it for the first time (it did not). Either way you end up with exactly one charge.
- If retries are exhausted, reconcile rather than guess — poll
GET /bookings?idempotencyKey=…to discover the true state. - Never show "failed" on a timeout. Show "we're confirming this" and resolve it from the server's answer, because a false failure causes the user to try again manually, which is the worst outcome.
The general principle worth stating: idempotency is a property the server provides; the client can only supply the key that makes it possible.
6.7 Realtime transports
| Long polling | SSE | WebSocket | |
|---|---|---|---|
| Direction | Request/response | Server → client | Bidirectional |
| Protocol | HTTP | HTTP (text/event-stream) |
Upgrade from HTTP, then its own framing |
| Reconnect | Manual | Built in, with Last-Event-ID |
Manual |
| Proxy/CDN friendliness | Best | Good | Often needs special config |
| Use for | Fallback | Price pushes, notifications, progress | Chat, collaborative editing, trading |
Example. Subscribe per visible item rather than per result set — drive subscriptions from IntersectionObserver so offscreen cards unsubscribe; coalesce bursts into one flush per animation frame.
Trade-offs. WebSockets are stateful, which complicates load balancing, scaling and deploys; SSE rides normal HTTP infrastructure but is one-way and limited to text.
Follow-ups. How do you reconnect without hammering the server? (Exponential backoff with jitter, a sequence number to detect gaps, full refetch after a gap.) How do you scale a WebSocket fleet? How do you handle backpressure?
6.8 API shapes: REST, GraphQL, BFF
What it is. REST exposes resources; GraphQL exposes one endpoint with a client-specified query; a BFF is a server owned by the frontend team that aggregates and trims payloads for one surface.
Example. For a travel marketplace, a BFF per surface (web, iOS, Android) keeps the mobile payload small without forcing every backend service to change.
Trade-offs. GraphQL removes over-fetching and client-side waterfalls, and costs you CDN cacheability (POST by default), N+1 risk on the server, and the need for query-cost limits and persisted queries. REST stays trivially cacheable.
Follow-ups. How do you cache a GraphQL response at the edge? Cursor vs offset pagination, and why does offset break on a live inventory? How do you version an API the frontend depends on?
7. Caching, storage and offline
7.1 HTTP cache headers
What it is. Cache-Control sets the policy; ETag/Last-Modified enable conditional revalidation; Vary defines the cache key's extra dimensions.
| Directive | Meaning | Use for |
|---|---|---|
max-age=31536000, immutable |
Never revalidate for a year | Content-hashed static assets |
no-cache |
May store, must revalidate before use | HTML documents |
no-store |
Never written to disk | Authenticated or PII responses |
private / public |
Whether shared caches may store it | private for anything user-specific |
stale-while-revalidate=60 |
Serve stale instantly, refresh in background | Content APIs that tolerate seconds of staleness |
s-maxage |
CDN-only lifetime, overrides max-age |
Short browser TTL, long edge TTL |
Example.
GET /app.3f9a1c.js → Cache-Control: public, max-age=31536000, immutable
GET / → Cache-Control: no-cache
ETag: "v842"
next request: If-None-Match: "v842" → 304 Not Modified
Trade-offs. private vs public is a correctness decision, not a performance one — getting it wrong serves one user's booking page to another. Over-varying (e.g. on User-Agent) destroys hit rate.
Follow-ups. What's the difference between no-cache and no-store? What exactly does a 304 save? Strong vs weak ETags? Why must Vary: Accept-Encoding be set?
More examples.
# 1 — hashed asset: cache hard, forever, never revalidate
GET /app.3f9a1c.js
Cache-Control: public, max-age=31536000, immutable
# 2 — the document: always revalidate, cheap when unchanged
GET /
Cache-Control: no-cache
ETag: "v842"
→ If-None-Match: "v842" → 304 Not Modified (headers only, no body)
# 3 — short browser TTL, long edge TTL, instant-but-fresh
Cache-Control: max-age=0, s-maxage=600, stale-while-revalidate=3600
# 4 — correctness with compression and locales
Vary: Accept-Encoding, Accept-Language
Users report seeing another user's name in the header. One header is wrong.
Answer: a personalised response was served with Cache-Control: public (or
simply no private), so a shared cache — the CDN, or a corporate proxy — stored
one user's HTML and served it to the next.
The fix is Cache-Control: private, no-store on anything user-specific, but the
more useful answer is the architecture that makes this impossible:
- Keep the cached document anonymous. Cache the shell publicly and fill in personalisation client-side, or at the edge from the session cookie.
- If you must vary at the edge, put the discriminator in the cache key
explicitly rather than relying on
Varywith a cookie, which is easy to get subtly wrong and destroys hit rate. - Treat
publicas something you opt into deliberately, never a default.
This is worth rehearsing as an incident story: it is high-severity, it is a one-line cause, and the prevention is a policy rather than a patch.
7.2 The cache-busting pattern
What it is. Immutable, long-lived, content-hashed assets plus a short-lived or revalidated HTML document that points at them.
Example. index.html is no-cache; it references app.3f9a1c.js which is immutable. A deploy changes the hash, so no purge is needed and no user ever gets a mismatched chunk.
Trade-offs. You still must handle a chunk 404 for users who loaded old HTML before a deploy — catch ChunkLoadError, retry with a cache-bust, then force a reload.
Follow-ups. What happens to a user who's been sitting on a tab for an hour when you deploy? How do you roll back without breaking in-flight sessions? (Keep the previous N builds' assets on the CDN.)
7.3 CDN and the cache layer stack
What it is. A response passes through: browser memory cache → browser disk cache → Service Worker cache → preload cache → CDN edge → CDN shield/regional → origin.
Example. A shield tier collapses many edge misses into one origin request, which protects the origin during a cache purge or a traffic spike.
Trade-offs. Caching personalised HTML at the edge requires the personalisation to be in the cache key, which fragments the cache. The usual resolution is to cache an anonymous shell and personalise client-side or via an edge function.
Follow-ups. What should never be cached at the CDN? How do you invalidate — purge by URL, surrogate key, or just change the hash? What does a cache HIT/MISS/STALE header tell you when debugging?
7.4 Browser storage
| Mechanism | Size | Sync? | Sent to server | Use for |
|---|---|---|---|---|
| Cookie | \~4KB | n/a | Yes, every request | Session/auth (HttpOnly, Secure, SameSite) |
localStorage |
\~5MB | Synchronous — blocks the main thread | No | Small prefs, flag cache. Never tokens |
sessionStorage |
\~5MB | Synchronous | No | Per-tab state, multi-step funnel |
| IndexedDB | Quota-based | Async | No | Offline data, search index, request queue |
| Cache Storage | Quota-based | Async | No | Service Worker asset and response cache |
| In-memory | n/a | n/a | No | Access tokens, anything secret-adjacent |
Example. Reading localStorage inside a scroll handler is a genuine performance bug — it is synchronous disk I/O on the main thread.
Trade-offs. Cookies travel on every request, adding bytes to every call; localStorage is convenient but XSS-readable and synchronous; IndexedDB is the only option at size but has an awkward API (use idb).
Follow-ups. Where would you store an auth token and why? What is storage quota and what happens when it's exceeded? How do you sync state between tabs? (BroadcastChannel or a storage event.)
7.5 Cookie attributes
What it is. HttpOnly blocks JavaScript access; Secure restricts to HTTPS; SameSite controls cross-site sending; Domain/Path set scope; Max-Age/Expires set lifetime.
Example. Set-Cookie: sid=...; HttpOnly; Secure; SameSite=Lax; Path=/; Max-Age=1209600
Trade-offs. SameSite=Strict is the safest and breaks inbound links from email and partner sites (the user arrives logged out). Lax is the modern default and allows top-level GET navigation. None requires Secure and is needed for genuine third-party contexts.
Follow-ups. Why does SameSite=Lax mostly solve CSRF? What is a cookie-prefix (__Host-)? How do cookies behave across subdomains, and what breaks when several apps share an apex domain?
7.6 Service Workers and offline
What it is. A proxy worker sitting between the page and the network, with its own cache, lifecycle and no DOM access.
Example. Lifecycle is install → waiting → activate; skipWaiting and clients.claim take over immediately. Strategy by resource type:
| Resource | Strategy |
|---|---|
| Hashed static assets | Cache-first |
| HTML documents | Network-first with a cached fallback |
| Content APIs (descriptions, reviews) | Stale-while-revalidate |
| Prices, availability, payments | Network-only |
Trade-offs. A Service Worker is the most dangerous thing a frontend team deploys — a bad one is sticky and can serve broken assets indefinitely. Always ship a kill switch (remote flag that makes it unregister) and know Clear-Site-Data. skipWaiting risks mixing an old page with a new worker.
Follow-ups. How do you update a Service Worker safely? What is Background Sync for? How would you make a booking flow survive a tunnel? (Draft in IndexedDB, queued submit with an idempotency key, explicit offline UI — never an optimistic success.)
8. Security
8.1 Same-origin policy and CORS
What it is. An origin is scheme + host + port. The same-origin policy blocks a document from reading cross-origin responses. CORS is the server's mechanism to relax that — it is not a defence.
Example. A non-simple request triggers a preflight:
OPTIONS /api/book Origin: https://www.example.com
Access-Control-Request-Method: POST
Access-Control-Request-Headers: content-type, authorization
→ Access-Control-Allow-Origin: https://www.example.com
Access-Control-Allow-Credentials: true
Access-Control-Max-Age: 600
Trade-offs. Allow-Credentials: true cannot be combined with Origin: * — you must echo a specific, validated origin. Blindly reflecting Origin is a common vulnerability. A high Max-Age cuts preflights but delays policy changes.
Follow-ups. What makes a request "simple" (no preflight)? Does CORS protect your API? (No — curl ignores it; authorization is the server's job.) Why can a <img> or <form> go cross-origin but fetch cannot read the response?
8.2 XSS
What it is. Executing attacker-controlled script in your origin.
| Type | How | Defence |
|---|---|---|
| Stored | Malicious input saved then rendered (a hotel review) | Encode on output, sanitise on input |
| Reflected | Payload in the URL echoed into the page | Context-aware escaping |
| DOM-based | Client-side sink: innerHTML, document.write, eval, location, dangerouslySetInnerHTML |
Avoid sinks; DOMPurify; Trusted Types |
Example. React escapes interpolated values, so the real React vectors are exactly four: dangerouslySetInnerHTML, href={userInput} with a javascript: URL, spreading unvalidated props onto an element, and server-rendered JSON injection (an unescaped </script> inside a serialised payload).
Trade-offs. Sanitising HTML is a losing arms race compared with not rendering HTML at all; Trusted Types converts a code-review problem into a browser-enforced one, at the cost of a migration.
Follow-ups. How would you safely render user-authored rich text? Why is encodeURIComponent not enough in an HTML attribute? What does Trusted Types enforce?
More examples.
// 1 — the four real React vectors
<div dangerouslySetInnerHTML={{ __html: userHtml }} /> // sink
<a href={userUrl}>link</a> // javascript: URL
<Component {...untrustedProps} /> // can inject onError etc.
<script>{`window.__DATA__ = ${JSON.stringify(data)}`}</script> // </script> breakout
// 2 — safe versions
<a href={/^https?:\/\//.test(userUrl) ? userUrl : '#'}>link</a>
const safe = JSON.stringify(data).replace(/</g, '\\u003c'); // escape the breakout
// 3 — when you genuinely must render HTML
import DOMPurify from 'dompurify';
<div dangerouslySetInnerHTML={{ __html: DOMPurify.sanitize(userHtml) }} />
A hotel review renders fine but fires an alert. Where did it get in?
Answer: it is stored XSS, and the review text reached a sink somewhere that bypasses React's escaping. The three places to look, in order:
dangerouslySetInnerHTML— because the product wanted bold text in reviews.- Server-rendered JSON — the review was interpolated into a
<script>tag and contained</script>, which closes the tag early and starts markup. - A third-party widget given the raw string and writing it with
innerHTML.
Note what is not the cause: {review.text} in JSX. React escapes that.
The systemic fix is not "sanitise harder" — sanitisers are an arms race. It is
Trusted Types (require-trusted-types-for 'script'), which makes the browser
throw on any assignment to a dangerous sink unless the value passed through a
registered policy. That converts a code-review problem into a runtime guarantee,
which is the difference between a fix and a control.
8.3 Content Security Policy
What it is. A response header restricting which sources of script, style, images and frames may load and execute.
Example. A strict modern policy uses nonces plus strict-dynamic, not a host allowlist:
Content-Security-Policy:
script-src 'nonce-{random}' 'strict-dynamic' https: 'unsafe-inline';
object-src 'none'; base-uri 'none'; require-trusted-types-for 'script';
report-uri /csp-report
'unsafe-inline' and https: here are fallbacks for old browsers that modern browsers ignore once a nonce is present.
Trade-offs. Host allowlists are routinely bypassable via JSONP endpoints on allowed CDNs. Nonces are incompatible with full-page CDN caching (each response needs a fresh nonce) — which pushes you to hashes or edge-injected nonces. Inline styles from runtime CSS-in-JS and tag managers are the usual blockers.
Follow-ups. How do you roll out CSP without breaking the site? (Content-Security-Policy-Report-Only, collect violations for two weeks, then enforce.) Why is strict-dynamic better than an allowlist? How does CSP interact with a CDN?
8.4 CSRF
What it is. An attacker-controlled page causes the victim's browser to make an authenticated write using cookies it sends automatically.
Example. Defences, in order of preference: SameSite=Lax/Strict cookies, a synchroniser or double-submit token, and checking Origin/Sec-Fetch-Site server-side.
Trade-offs. CSRF defences are only needed when you authenticate with cookies — a bearer token in an Authorization header is not sent automatically, so it is not CSRF-exposed (but is XSS-exposed instead).
Follow-ups. CORS vs CSRF in one sentence each. Does SameSite=Lax fully solve CSRF? (Not for top-level GET state changes — another reason GET must stay safe.)
8.5 Security headers
| Header | Purpose |
|---|---|
Strict-Transport-Security |
Force HTTPS; preload covers the first visit |
X-Content-Type-Options: nosniff |
Stop MIME-confusion attacks |
Referrer-Policy: strict-origin-when-cross-origin |
Stop leaking paths and query strings |
Permissions-Policy |
Disable geolocation/camera for embedded third parties |
Cross-Origin-Opener-Policy / -Embedder-Policy |
Process isolation; required for SharedArrayBuffer |
Cross-Origin-Resource-Policy |
Block cross-origin reads of your resources |
X-Frame-Options / CSP frame-ancestors |
Clickjacking |
Follow-ups. Which of these would you set first on a new service, and why? What breaks when you turn on COEP?
8.6 Authentication vs authorization
What it is. Authentication is "who are you"; authorization is "what may you do."
Example. Hiding an admin button is a UX affordance. The server must independently enforce the permission on the endpoint.
Trade-offs. Fetching a capabilities object lets the UI render correctly without duplicating policy, at the cost of an extra round trip and a cache-invalidation question when roles change.
Follow-ups. Where must authorization be enforced? How do you handle a user whose permissions change mid-session?
8.7 OAuth, OIDC and JWT
What it is. OAuth 2.0 is a delegated authorization framework; OpenID Connect is the authentication layer on top of it (the id_token); JWT is merely a token format. An OAuth implementation need not use JWTs.
Example. For a browser app, use Authorization Code with PKCE. The implicit flow is deprecated because it leaked tokens into URL fragments, history and referrers. Validate state (CSRF on the callback) and nonce (id_token replay).
A JWT is header.payload.signature, base64url — encoded, not encrypted. Server-side validation must check the signature, pin the expected alg (reject none and algorithm confusion), and check iss, aud, exp, nbf.
Trade-offs. Stateless JWTs scale well but cannot be revoked before expiry — so either keep access-token lifetimes very short or maintain a denylist and give up some statelessness. Opaque tokens with introspection are the opposite trade.
Follow-ups. Access vs refresh tokens? How does rotation with reuse detection work? (Each refresh issues a new token and invalidates the old; if an old one reappears, revoke the whole family.) What stops ten parallel 401s triggering ten refreshes? (A single-flight lock.)
8.8 Token storage
What it is. The choice between localStorage, cookies and memory for auth material.
Example. The defensible architecture: short-lived access token in memory only, refresh token in an HttpOnly; Secure; SameSite=Strict cookie scoped to the refresh path, silent refresh on 401.
Trade-offs. localStorage is XSS-readable; cookies are CSRF-exposed but can be HttpOnly. The honest caveat: under XSS you lose either way, because the attacker can simply call your API from the victim's session. That is why CSP and Trusted Types matter more than storage choice.
Follow-ups. Why is "just use HttpOnly cookies" not a complete answer? What survives a page refresh with in-memory tokens? (Nothing — you silently refresh on load.)
More examples.
// 1 — access token in memory only, never persisted
let accessToken = null; // dies on refresh, by design
// 2 — refresh via an HttpOnly cookie, single-flight so N 401s cause one refresh
let inflight = null;
const refresh = () => (inflight ??= fetch('/auth/refresh', { credentials: 'include' })
.then(r => r.json())
.finally(() => { inflight = null; }));
// 3 — the cookie that carries it
// Set-Cookie: rt=…; HttpOnly; Secure; SameSite=Strict; Path=/auth/refresh; Max-Age=1209600
“Just use HttpOnly cookies, then XSS can’t steal the token.” Rebut it.
Answer: HttpOnly stops the attacker reading the token. It does not stop
them using it.
With script execution in your origin, the attacker simply calls your API from the victim's own session. The browser attaches the cookie automatically, same-origin checks pass, and every request looks legitimate. They do not need the token's value — they have something better: the ability to act as the user, from the user's browser, for as long as the page is open.
So HttpOnly is worth having (it prevents exfiltration, which enables offline
and later abuse), but it is a mitigation, not a boundary.
The honest framing: under XSS you have lost, whatever the storage choice.
Storage decisions trade one exposure for another — localStorage is XSS-readable,
cookies are CSRF-exposed. Neither is a defence against script injection. That is
why CSP and Trusted Types matter more than the storage debate, and saying so is
what separates a considered answer from a memorised one.
8.9 Supply chain and third parties
What it is. Risk arriving through dependencies and third-party scripts rather than your own code.
Example. Controls: committed lockfile, npm ci only, Renovate/Dependabot with grouped patch automerge, npm audit/Socket in CI, ignore-scripts plus a private registry proxy, and provenance/sigstore where available. For third-party tags: Subresource Integrity hashes, crossorigin, sandboxed iframes, and a review gate.
Trade-offs. SRI breaks when the vendor updates the file — which is the point, but it means vendor-hosted "always latest" scripts cannot use it. Sandboxing a tag often breaks the feature the business wanted.
Follow-ups. What is dependency confusion and how do scoped private packages prevent it? A marketing tag on your checkout page — what's the blast radius? (Full compromise.) How would you audit what third-party scripts are actually on the page?
8.10 Privacy and compliance on the frontend
What it is. Consent, data minimisation and scope reduction as frontend responsibilities.
Example. Consent management must gate analytics before it loads; PII never goes in URLs (they land in logs and Referer); session-replay tools need field masking; and a payment iframe or hosted fields keeps card data out of your DOM, which removes you from most PCI scope.
Trade-offs. Hosted payment fields reduce compliance scope and cost you styling control and a slightly worse checkout UX.
Follow-ups. How do you keep analytics working under a strict CSP? What would you mask in session replay on a booking form?
9. Performance
9.1 Core Web Vitals
What it is. Three field metrics, measured at the 75th percentile of real users.
| Metric | Good | Needs work | Measures | Top causes |
|---|---|---|---|---|
| LCP | ≤ 2.5s | ≤ 4.0s | Render time of the largest in-viewport element | Slow TTFB, render-blocking CSS/JS, lazy-loaded hero, unoptimised images |
| INP | ≤ 200ms | ≤ 500ms | Worst interaction latency: input delay + processing + presentation | Long tasks, heavy handlers, large re-render trees, hydration |
| CLS | ≤ 0.1 | ≤ 0.25 | Sum of unexpected layout shift scores | Images without dimensions, injected banners, late fonts, dynamic content above the fold |
Example. Supporting diagnostics: TTFB (server + network), FCP (first paint of anything), TBT (sum of long-task time over 50ms), TTI. TBT is the lab proxy; INP is the field metric.
Trade-offs. p75 hides the tail — always look at p95 too, and segment by device class and country, because a global average flatters you while low-end Android users suffer.
Follow-ups. Why did INP replace FID? What's the difference between lab and field data? How do you attribute an INP regression to a specific interaction? (event timing entries and the Long Animation Frames API.)
More examples.
// 1 — attribute a regression to a deploy, not just a route
onLCP(m => beacon({ v: m.value, build: __BUILD_ID__, route, device, country }));
// 2 — find WHICH element is the LCP, in the field
new PerformanceObserver((l) => {
const e = l.getEntries().at(-1);
beacon({ lcp: e.startTime, el: e.element?.tagName, url: e.url });
}).observe({ type: 'largest-contentful-paint', buffered: true });
// 3 — attribute INP to the actual interaction
new PerformanceObserver((l) => l.getEntries()
.filter(e => e.duration > 200)
.forEach(e => beacon({ inp: e.duration, type: e.name, target: e.target?.id })),
).observe({ type: 'event', durationThreshold: 200, buffered: true });
p75 LCP is 2.4s — green. Users still complain. What did you miss?
Answer: p75 globally is an average of populations, and it hides the segments that are actually suffering.
Slice the same data four ways before believing it:
- Device class — a mid-range Android may sit at p75 of 5s while desktop drags the aggregate down. Most travel traffic in emerging markets is low-end Android.
- Country / network — a distant origin or a 3G tail behaves nothing like the office.
- Route — the home page may be fine while search, where the money is, is not.
- p95, not just p75 — the complaining users are, by definition, in the tail.
Then check you are measuring the right thing at all: if complaints are about responsiveness rather than loading, LCP is simply the wrong metric and INP is where to look.
The general lesson: a green aggregate is a hypothesis, not a conclusion. Always ask "green for whom?"
9.2 Measurement
What it is. Lab tools give deterministic, repeatable numbers; field tools (RUM) give the truth about real users.
Example.
import { onLCP, onINP, onCLS } from 'web-vitals';
const send = m => navigator.sendBeacon('/rum', JSON.stringify({
name: m.name, value: m.value, id: m.id,
route: currentRoute(), buildId: __BUILD_ID__, // attribute regressions to a deploy
conn: navigator.connection?.effectiveType,
}));
onLCP(send); onINP(send); onCLS(send);
Trade-offs. Lighthouse CI is good for gating because it is deterministic, and blind to real device and network diversity. RUM is true but noisy and needs volume before a signal appears — hence synthetic monitoring as the early-warning layer.
Follow-ups. How do you read a Performance panel trace? (Long tasks, forced reflow warnings, the main-thread flame chart, Recalculate Style and Layout blocks.) What is CrUX? Why beacon on visibilitychange rather than unload?
9.3 Loading optimisation
What it is. Shortening the critical path from request to largest paint.
Example. In rough order of payoff:
- Critical CSS inline, rest deferred; per-route CSS extraction.
- Resource hints:
preconnectto the image/API origin,preloadthe LCP image and the critical font,fetchpriority="high"on the hero andlowbelow the fold,prefetchthe next likely route. - Images: AVIF with WebP fallback, responsive
srcset/sizes, explicitwidth/heightoraspect-ratio,loading="lazy"only below the fold,decoding="async", CDN resizing on the fly. - Fonts:
font-display: swap(oroptional),preloadone critical woff2, subset glyph ranges,size-adjustmetric overrides so the swap doesn't shift layout. - JavaScript: route-level splitting, then component-level for heavy widgets, then shrink what remains.
- Third parties: facade pattern (a fake chat button that loads the real widget on click), iframe sandboxing, or Partytown.
Trade-offs. Preload competes with itself — more than about three hints and you delay the thing you cared about. loading="lazy" on the LCP image is a classic self-inflicted regression.
Follow-ups. What's the difference between preload, prefetch and preconnect? Why can inlining a script hurt? How would you optimise a hotel listing page with forty images?
9.4 Compression
What it is. Brotli (br) beats gzip by roughly 15–20% on text assets; gzip is the universal fallback. Negotiated via Accept-Encoding, which makes Vary: Accept-Encoding mandatory.
Example. Pre-compress static assets at build time with Brotli quality 11 and serve the .br file; use quality 4–6 for dynamic responses so you don't trade TTFB for bytes.
Trade-offs. Never re-compress already-compressed payloads (JPEG, AVIF, WebP, woff2, video) — you burn CPU and can grow the file. Compressing a response containing a secret alongside reflected user input is the BREACH class.
Follow-ups. Why state bundle budgets in compressed bytes? How does compression interact with CDN caching? Why not Brotli images?
9.5 Runtime performance
What it is. Keeping interactions under the INP threshold once the page is loaded.
Example. The standard levers: virtualise long lists (react-window, TanStack Virtual); useDeferredValue/startTransition for filtering while typing; debounce input and throttle scroll; move big JSON parses to a worker; animate on the compositor via CSS or the Web Animations API; memoise deliberately rather than everywhere.
Trade-offs. Memoisation has its own cost (comparison plus retained references) and useMemo everywhere makes code harder to change for no measured gain — prefer state colocation and context splitting first. React Compiler changes this calculus.
Follow-ups. A 10,000-row price calendar is janky — what do you do? Why doesn't React.memo help when you pass an inline object or arrow? What is a long task and how do you break one up?
9.6 Budgets and enforcement
What it is. A per-route limit on transferred bytes and lab metrics, enforced in CI.
Example. Search route: 170KB gzip JS, 50KB CSS, LCP image ≤ 150KB, Lighthouse LCP ≤ 2.5s. Enforce with size-limit or bundlesize; fail the PR and post the delta plus a treemap link as a bot comment.
Trade-offs. Hard budgets block legitimate features; the usual resolution is a documented exception process with an owner and an expiry, not a silently raised threshold.
Follow-ups. Where do you set the budget — total repo or per route? Who owns a budget breach? How do you stop budgets drifting upward over a year?
9.7 A performance debugging workflow
What it is. A repeatable sequence from symptom to systemic fix.
Example.
- Quantify in RUM: which metric, which p, which segment, since when.
- Correlate with build ids and releases to find a candidate cause.
- Reproduce in the lab with matching throttling and device class.
- Trace in the Performance panel; find the long task or the blocking resource.
- Attribute to a module — bundle analyser treemap, or the LoAF script attribution.
- Fix, then verify in the lab, then confirm in RUM at p75 and p95.
- Guardrail: a budget, a lint rule, or a CI check so the class of regression cannot return.
Trade-offs. Steps 1–2 are often skipped in favour of jumping to the lab — which finds a problem, rarely the problem.
Follow-ups. Walk me through a real regression you diagnosed. How did you know the fix worked? What stops it recurring?
10. Bundling and build systems
10.1 Module formats
What it is. CommonJS uses dynamic, synchronous require resolved at runtime. ESM uses static import/export, hoisted, with live bindings, resolved before execution.
Example. That static structure is what makes tree shaking, top-level await and import maps possible. CJS cannot be reliably tree-shaken because require('./x')[name] is only knowable at runtime.
Trade-offs. Interop is the pain: "type": "module", dual exports maps, no __dirname in ESM, and packages that ship both. A platform team's job is to publish correct exports conditions (import, require, types, browser, default) and run publint/arethetypeswrong in CI.
Follow-ups. Why can't CJS be tree-shaken? What is the dual-package hazard? What does an import map solve?
10.2 What a bundler does
What it is. Resolve imports from the entry points → load and transform each module → build the module graph → optimise (tree shake, scope hoist, minify, split chunks) → emit content-hashed chunks plus a manifest.
Example. The manifest is what the server or HTML uses to map a route to its chunks — and what makes "build once, promote many" possible.
Trade-offs. More aggressive optimisation means slower builds; most teams run full optimisation only in production builds and transpile-only in dev.
Follow-ups. What is scope hoisting and why does it help? What's in a source map and why hidden-source-map in production?
10.3 Tree shaking
What it is. Dead-code elimination over the ESM module graph.
Example. It works only when modules are ESM, the bundler can prove no side effects, and the package declares "sideEffects": false (or lists the files that do have them).
Trade-offs. Silent breakers: re-exporting a whole barrel (export * from './everything'), CJS dependencies, class-property mutation at module scope, and transpiling to CJS before bundling. Barrel files are the number-one cause of bloated frontend bundles — the fix is to ban deep barrels, import from paths, or generate per-component entry points.
Follow-ups. Why did importing one icon pull in 400KB? What does sideEffects actually tell the bundler? How do you prove a dependency is tree-shakeable?
10.4 Code splitting and chunking
What it is. Breaking the graph into chunks loaded on demand.
Example.
const SearchPage = React.lazy(() => import('./SearchPage')); // route level
const Map = React.lazy(() => import('./Map')); // heavy widget
onMouseEnter={() => import('./CheckoutPage')} // prefetch on intent
Chunk strategy: a framework chunk (react, react-dom) that changes rarely so it stays cached, a commons chunk for modules used by ≥N routes, and per-route chunks.
Trade-offs. Over-splitting costs request overhead and deepens the waterfall; under-splitting ships code nobody runs. Splitting also introduces the chunk-404-after-deploy failure, which needs a retry-then-reload handler.
Follow-ups. Where do you split first? How do you avoid a loading spinner on every navigation? Why does a shared chunk's hash changing invalidate everything downstream?
More examples.
// 1 — route split with prefetch on intent, so splitting costs nothing visible
const Checkout = lazy(() => import('./Checkout'));
<Link onMouseEnter={() => import('./Checkout')} onFocus={() => import('./Checkout')} />
// 2 — survive a deploy that removed the chunk you are asking for
window.addEventListener('vite:preloadError', () => location.reload());
// webpack equivalent: catch ChunkLoadError, retry with a cache-bust, then reload
// 3 — keep the framework chunk stable so it stays cached across deploys
// splitChunks: { cacheGroups: { framework: { test: /[\\/]node_modules[\\/](react|react-dom)/ } } }
You split aggressively and it got slower. How?
Answer: several ways, and they compound.
- Waterfall depth. Chunk A imports B imports C. Each level costs a round trip, and the browser cannot discover C until B has arrived. Splitting trades bytes for round trips, and on high-latency mobile that trade can lose.
- Lost compression ratio. Many tiny chunks compress worse than one larger one — shared dictionary, per-response overhead. Below roughly 20KB a chunk often is not worth its own request.
- Duplicated shared code. Without a sensible
commonsgroup, the same module gets inlined into several route chunks. - Spinner-per-navigation. Technically faster first load, worse felt performance, because now every click waits.
The fix is not "split less" but "split along the right seams": route level first, then genuinely heavy conditional widgets, with the framework in its own long-lived chunk — and prefetch on intent so the user never waits for a split you chose.
10.5 Tool comparison
| Tool | Engine | Strength | Choose when |
|---|---|---|---|
| Webpack 5 | JS | Largest plugin ecosystem, Module Federation, filesystem cache | Large legacy monorepos, federation-heavy setups |
| Vite | esbuild (dev) + Rollup (prod) | Native-ESM dev server, near-instant HMR | Default for new apps |
| Rspack / Turbopack | Rust | Webpack-compatible (Rspack), 5–10× faster | Migrating a big Webpack build without a rewrite |
| esbuild | Go | Extremely fast transform and bundle | Libraries, tooling, dev transforms |
| Rollup / tsup | JS | Cleanest library output, multiple formats | Publishing design-system packages |
| SWC / Babel | Rust / JS | Transform only | SWC for speed, Babel when you need AST plugins or codemods |
Follow-ups. Why is Vite's dev server fast? (Serves source as native ESM, transforms on demand, pre-bundles deps with esbuild.) Why does Vite still use Rollup for production?
10.6 HMR
What it is. Swapping a module in the running app and propagating the update up the import graph to the nearest boundary that accepts it.
Example. If a change to a context provider full-reloads the page, it is because the boundary is effectively the root — usually a module with side effects or a root-level export change.
Trade-offs. HMR preserves state, which speeds iteration but can mask bugs that only appear on a fresh mount.
Follow-ups. Why does our HMR full-reload every time? What state survives an HMR update and what doesn't?
10.7 Build speed at scale
| Lever | Effect |
|---|---|
| Persistent filesystem cache (Webpack 5, Vite) | 5–10× on warm local builds |
| Remote build cache (Turborepo, Nx Cloud, Bazel) | CI reuses teammates' and previous runs' outputs |
| Affected-only builds from the dependency graph | PR CI proportional to the change, not the repo |
| Transpile-only in dev, type-check in a parallel job | Removes tsc from the hot loop |
hidden-source-map in prod |
Large emit-time saving, still uploadable to Sentry |
| Babel → SWC, Webpack → Rspack | Usually the single biggest step change |
Trade-offs. Remote caching requires deterministic builds — timestamps, absolute paths and non-pinned tool versions all poison the cache.
Follow-ups. Our CI is 35 minutes; get it to 10. What makes a build non-deterministic? How would you measure build time as an SLO?
11. Platform architecture
11.1 Styling approaches
| Approach | Runtime cost | Pros | Cons |
|---|---|---|---|
| Plain CSS + BEM | None | Simple, cacheable | Discipline-dependent, no scoping guarantee |
| CSS Modules | None | Real scoping, zero runtime | Needs CSS variables for dynamic theming |
| Tailwind / utility | None | Tiny shipped CSS, no naming debates | Verbose markup, needs token discipline |
| Runtime CSS-in-JS (styled-components, Emotion) | Per-render style computation | Colocated, dynamic props | Runtime overhead, breaks with RSC, hurts INP at scale |
| Zero-runtime CSS-in-JS (vanilla-extract, Linaria, Panda) | None | Type-safe tokens, extracted at build | Build complexity, less dynamic |
Trade-offs. The defensible modern recommendation is CSS variables for tokens plus a zero-runtime or utility layer: it survives Server Components, costs nothing at runtime, and makes theming a CSS concern rather than a React concern.
Follow-ups. Why does runtime CSS-in-JS hurt INP? How would you theme without re-rendering? What breaks when you use Emotion inside a Server Component?
11.2 Design tokens
What it is. Three tiers: primitives (--blue-500: #1a73e8), semantic (--color-action-primary: var(--blue-500)), component (--button-bg-primary: var(--color-action-primary)).
Example. Components consume only tiers 2 and 3. Themes swap the tier-1→tier-2 mapping via a data-theme attribute, so a theme switch is one DOM attribute and zero React re-renders. Tokens live in one source (JSON / Figma Tokens) and generate CSS, TS types, iOS and Android outputs via Style Dictionary.
Trade-offs. Three tiers is more indirection than a small team needs; it pays off the moment you have a second theme, a white-label partner, or a native platform.
Follow-ups. How do you add dark mode without touching every component? How do you stop teams using raw hex values? (Lint rule plus a Stylelint config.)
11.3 Component API design
What it is. The rules that make a shared library usable by hundreds of engineers without tickets.
Example.
- Composition over configuration. A
<Select>with 40 props is unmaintainable; compound components (Select.Trigger,Select.Option) let consumers compose. Headless primitives (Radix, React Aria, Ark) give behaviour and a11y while you own styling. - Support controlled and uncontrolled —
value+onChange, anddefaultValue. - Bounded escape hatches — allow
className/styleon the root, forward refs, spread rest props onto the right element, exposeas/asChildfor polymorphism; don't expose internals you'll want to change. - Accessibility is not a prop — keyboard interaction, focus management and
aria-*live inside the component. - No business logic — no fetching, flags or analytics coupling inside a primitive.
Trade-offs. Headless libraries save you months of a11y work and add a dependency you must track. Escape hatches improve adoption and make future refactors harder — which is why they should be explicit and documented, not accidental.
Follow-ups. Design the API for a date-range picker. How do you handle a team that needs a one-off variant? How do you forward a ref through a polymorphic component?
11.4 Publishing shared packages
What it is. Internal versioned packages as the distribution mechanism for platform code.
Example. The breaking-change playbook:
- Additive change first — new prop, old behaviour unchanged by default.
- Deprecate with
@deprecatedJSDoc, a dev-only console warning, and a lint rule. - Ship a codemod (
jscodeshift/ts-morph) and run it across the monorepo yourself. - Publish the major with a migration guide and a compatibility layer for one minor cycle.
- Track "packages on latest major" on a dashboard and chase the tail personally.
Trade-offs. Semver discipline slows you down and is the only thing that makes a shared library trustworthy. Steps 3 and 5 are the ones most teams skip, and they are where adoption actually happens.
Follow-ups. How do you prevent two versions of React in a consumer's bundle? (Peer dependencies plus a duplicate check in CI.) How do you version design tokens?
11.5 Monorepo vs polyrepo
| Monorepo | Polyrepo | |
|---|---|---|
| Cross-cutting change | One atomic PR | N coordinated PRs and a release dance |
| Dependency versions | Single version policy, enforced | Drift is the default |
| CI cost | Needs affected-graph tooling or it explodes | Naturally bounded |
| Release independence | Requires tooling (changesets) | Free |
| Ownership | CODEOWNERS paths |
Repo boundaries |
| Tooling investment | High, centralised, pays back at scale | Low per repo, duplicated N times |
Trade-offs. Choose by how often changes cross boundaries. If a typical feature touches the design system and two apps, monorepo — and you must fund the build tooling. If teams genuinely ship independently against stable contracts, polyrepo plus published packages is cheaper.
Follow-ups. How do you keep monorepo CI fast? (pnpm, affected-only, remote cache.) What's the failure mode of each? (Unowned shared code — identical in both.)
11.6 Micro-frontends
What it is. Independently deployable frontend applications composed into one user-facing surface.
| Approach | Mechanism | Pros | Cons |
|---|---|---|---|
| Build-time packages | npm versions in one host | Simple, optimal bundles | No independent deploy |
| Module Federation | Runtime remoteEntry.js + shared scope |
True independent deploy, shared singletons | Version negotiation, runtime failures, hard debugging |
| Import maps + native ESM | Browser-level resolution | Standards-based | Caching care needed, immature tooling |
| iframes | Hard isolation | Perfect CSS/JS isolation, security boundary | Routing, sizing, a11y, duplicated runtime |
| Web Components | Custom elements per team | Framework-agnostic, works in partner sites | Styling/slotting friction, weak SSR story |
| Server-side composition (ESI, Tailor, Podium) | Edge stitches HTML fragments | Great first load, SEO-safe | Needs edge infra; interactivity still needs a plan |
Example. With Module Federation the host declares remotes, the remote declares exposes, and both declare shared with singleton: true and requiredVersion for React — two React copies is the canonical production incident.
Trade-offs. Micro-frontends are justified only when teams need independent deploy cadence and the org boundary is real. The price: duplicated dependencies, version skew, cross-app inconsistency, harder E2E testing, worse aggregate performance, and a much heavier platform team.
Follow-ups. When would you not use micro-frontends? What must stay centralised regardless? (Tokens and components, auth, analytics and experiments, routing contract, error reporting, perf budget.) How do you roll back one remote?
11.7 Many apps, one domain
What it is. Several independently deployed apps served under one hostname via path-based routing at the CDN, reverse proxy or edge function.
Example. /search/* → origin A, /account/* → origin B, via CloudFront behaviours or an edge function. The details that matter:
- Asset path collisions — each app emits under its own prefix (
/search/_assets/...) or chunk names clash at the edge. SetpublicPath/baseper app. - SPA fallback per behaviour —
/account/settings/billingmust rewrite to that app'sindex.html, not a global one. - Cookie scope — apex-domain cookies are shared by all apps; scope non-auth cookies by path and be explicit about who may write the session cookie.
- Cache keys and HTML TTL — hashed assets immutable, each app's HTML invalidated independently.
- Shared chrome — a common header either duplicates (drift) or is federated/edge-included (coupling).
- Failure isolation — if app B's origin is down, app A must still serve.
Trade-offs. Edge routing isolates apps and adds infrastructure and routing-rule complexity that someone must own.
Follow-ups. How do you deploy one app without touching the others? What happens to a user navigating between two apps — full page load or client-side? What breaks with relative asset paths?
11.8 The frontend data layer
What it is. One place that owns request behaviour: auth headers, retries, caching, serialisation, error normalisation and cancellation.
Example. Instead of every component calling fetch differently, a query layer standardises cache keys, loading and error states, dedup, and AbortController wiring for stale searches.
Trade-offs. Not every request should retry — reads yes, non-idempotent mutations only with an idempotency key. Blanket retry policy is how a blip becomes an outage.
Follow-ups. Where should retry policy live? How do you cancel a stale search? How do you normalise errors from three different backends?
11.9 Developer experience
What it is. Treating the local loop and CI as a product with metrics.
Example. The four numbers worth quoting: cold install, dev server start, HMR round-trip, CI wall-clock p95. The five-step adoption playbook for any platform change: prove it on one team with before/after numbers → make it opt-in and trivial → provide the codemod → warn, then enforce (dev warning → CI warning → CI error, with announced dates) → dashboard the tail and close it by pairing.
Trade-offs. Enforcement too early creates resentment; too late and you have permanent drift. The dates are the contract.
Follow-ups. How do you measure DX? (PR cycle time, first-run CI pass rate, flake rate, time-to-first-PR, cache hit rate, plus a survey.) How do you get 12 teams to adopt something without a mandate?
12. Testing
12.1 The shape: trophy, not pyramid
What it is. For frontend the useful distribution is a thin base of static analysis, a few unit tests, the bulk in integration/component tests, and a small high-value E2E layer.
| Layer | Tool | Scope | Count | Runs in |
|---|---|---|---|---|
| Static | TypeScript, ESLint | Types, patterns | n/a | Pre-commit + CI |
| Unit | Vitest / Jest | Pure functions, reducers, hooks | Hundreds | Seconds, on save |
| Component / integration | Vitest + Testing Library, Playwright CT | A component with children, real DOM, mocked network | Thousands | 1–5 min in CI |
| Contract | Pact, or schema diff vs OpenAPI/GraphQL | Frontend↔backend payload shape | Per endpoint | CI, both repos |
| Visual regression | Chromatic / Percy / Playwright screenshots | Rendered pixels per state | Per story | CI on PR |
| E2E | Playwright, Cypress | Critical journeys on a real build | 20–50, not 500 | CI on merge + scheduled |
| Performance | Lighthouse CI, size-limit | Budgets per route | Per route | CI on PR |
| A11y | axe-core in component tests | WCAG violations | Every component | CI on PR |
Trade-offs. The classic pyramid came from backend services where unit tests catch most defects. Most frontend bugs are integration bugs — wrong props, wrong state transition, wrong fetch shape — so the mass belongs one layer up.
Follow-ups. Why not more E2E? (Slow, flaky, expensive to maintain, and they fail for reasons unrelated to the change.) What would you test at each layer for a booking form?
12.2 Testing behaviour, not implementation
What it is. Query the DOM the way a user perceives it — by role and accessible name — rather than by class or internal structure.
Example.
// good — survives refactors, and asserts accessibility as a side effect
await userEvent.click(screen.getByRole('button', { name: /search/i }));
expect(await screen.findByRole('list', { name: /results/i })).toBeInTheDocument();
// bad — couples the test to markup
wrapper.find('.btn-primary').simulate('click');
Trade-offs. Role-based queries are slower to write and occasionally force you to fix the markup first — which is the point. Test ids are a legitimate fallback for things with no accessible identity, used sparingly.
Follow-ups. Why is getByRole preferred? When is a test id acceptable? Why are large DOM snapshots worse than no test? (They get rubber-stamped on failure.)
More examples.
// 1 — query the way a user perceives the UI
await userEvent.click(screen.getByRole('button', { name: /search/i }));
expect(await screen.findByRole('list', { name: /results/i })).toBeInTheDocument();
// 2 — assert the absence of a loading state, not an implementation detail
await waitForElementToBeRemoved(() => screen.queryByRole('status'));
// 3 — the escape hatch, used sparingly and deliberately
screen.getByTestId('price-cell'); // only when there is no accessible identity
The test passes but the feature is broken in production. Classic causes?
Answer: the test asserted the implementation rather than the behaviour.
The usual shapes:
- Mocked at the module boundary.
jest.mock('./useSearch')means the real data layer — the part that broke — never ran. Mock at the network boundary with MSW so the component exercises its real hooks, cache and error handling. - Queried by test id. The element exists but is
aria-hidden, covered, or disabled. A real user could not click it;getByRolewould have failed. - No assertion on the async settled state. The test asserted the spinner, which always appears, then finished before the failure surfaced.
- The bug is in integration. Both units pass in isolation; the contract between them changed. This is precisely why the mass of frontend tests belongs at the integration layer rather than the unit layer.
The diagnostic question worth asking of any test: if the feature broke, would this test fail? A surprising number of green tests cannot answer yes.
12.3 Mocking and MSW
What it is. Intercepting at the network boundary rather than the module boundary, so the component under test runs its real data layer.
Example.
const server = setupServer(
http.get('/api/search', () => HttpResponse.json({ hotels: [...] })),
http.get('/api/prices', () => new HttpResponse(null, { status: 500 })), // error path
);
Trade-offs. Mocking the hook instead tests nothing but your mock. MSW handlers are another artefact that can drift from the real API — which is what contract tests are for.
Follow-ups. What's the difference between a stub, a mock and a fake? How do you test a loading state deterministically? How do you keep mocks honest?
12.4 Contract testing
What it is. Verifying that the shape the frontend expects matches what the backend actually produces, without running both end to end.
Example. Consumer-driven contracts (Pact) publish the frontend's expectations; the provider's CI replays them. Alternatively, generate types from OpenAPI/GraphQL in CI so a breaking schema change fails the frontend build.
Trade-offs. Pact adds a broker and process overhead; schema-generated types are cheaper but only catch shape changes, not semantic ones.
Follow-ups. Who owns a broken contract? How do you ship a backward-incompatible API without breaking the web client?
12.5 Visual regression and accessibility testing
What it is. Screenshot diffing per component state, and automated WCAG rule checks.
Example. Every design-system component ships a Storybook story per state, which doubles as both the visual-regression fixture and the axe test target — across themes and both text directions.
Trade-offs. Visual diffs catch what assertions can't and generate noise from font rendering and animation; you need deterministic rendering (frozen clock, disabled animations, pinned browser). axe catches roughly a third of a11y issues — claiming automation covers accessibility is a red flag.
Follow-ups. How do you stop visual tests being flaky? What does axe not catch? (Focus order, meaningful alt text, correct heading structure, keyboard traps in practice.)
12.6 Flaky tests
What it is. Tests that pass and fail on identical code, destroying trust in the suite.
Example. A programme that actually works:
- Quarantine, don't ignore — a detected flaky test moves to a suite that still runs but doesn't block, with an auto-filed ticket and an owner from
CODEOWNERS. - Measure flake rate per test and per suite, and publish it. A test over the threshold is deleted, not retried forever.
- Ban blanket retries at suite level; allow one retry, recorded as a flake signal.
- Attack root causes: auto-waiting locators instead of fixed timeouts, per-test data seeded via API not UI, frozen time and network, containerised environment identical to CI.
- Gate the gate — if
main's suite pass rate drops below \~99%, the pipeline is the incident.
Trade-offs. Deleting a flaky test loses coverage and is often still correct — a test nobody trusts provides zero coverage already while costing wall-clock and attention.
Follow-ups. What causes flakiness in E2E specifically? How do you detect a flaky test automatically? (Re-run on main on a schedule and diff outcomes.)
12.7 Coverage
What it is. The proportion of code executed by the test suite.
Example. Enforce coverage on changed lines in a PR rather than a global percentage. Use mutation testing (Stryker) on critical modules — pricing, currency, date handling — to check whether the tests actually assert anything.
Trade-offs. A global target produces coverage theatre: tests that execute code without asserting. Changed-lines coverage avoids both that and the untested-new-code problem.
Follow-ups. Is 100% coverage a good goal? What does mutation testing measure that line coverage doesn't?
12.8 A testing strategy for a large frontend
What it is. The platform-level answer: own the infrastructure, not the individual tests.
Example. Publish the test utilities as a package — a renderWithProviders wrapper carrying theme, i18n, router and query client; MSW handlers generated from the API schema; a Playwright fixture that logs in via API and seeds data. Every team then starts from the right setup, and improvements propagate with a version bump.
Trade-offs. A shared harness becomes a bottleneck if the platform team is the only one who can change it — hence a contribution model and clear extension points.
Follow-ups. How do you raise testing standards across twelve teams? What do you do about a team that writes no tests? How do you decide what gets an E2E test?
13. Operations
13.1 Pipeline shape
What it is. The sequence of checks on a PR and the promotion path after merge.
Example. On every PR: cached install → lint, type-check and unit/component tests in parallel → build affected packages → bundle-size diff comment → Lighthouse CI on key routes → visual regression → a11y checks → a preview URL per PR. On merge: build once into an immutable artifact and promote that same artifact through environments.
Trade-offs. "Build once, promote many" requires runtime configuration rather than build-time inlining, which is slightly more work and is the only way to guarantee that what you tested is what you shipped.
Follow-ups. Why not rebuild per environment? What goes in a PR check vs a nightly job? How do preview deploys change code review?
13.2 Progressive delivery
| Mechanism | Gives you | Watch out for |
|---|---|---|
| Feature flags | Deploy decoupled from release; kill switch without a rollback | Flag debt — require an expiry date and a cleanup ticket |
| Canary by percentage | Regression caught on 1% of traffic | Needs per-cohort metrics or there's no signal |
| Blue/green at the CDN | Instant switch and rollback | Both versions' chunks must stay fetchable |
| Server-side experiment assignment | No flicker, no blocking script | Must be in the cache key or you serve the wrong variant |
| Staged rollout by geo/device | Limited blast radius | Observability must segment the same way |
Example. An experiment platform that doesn't regress performance: assign at the edge or in the BFF, pass the assignment into SSR, fire the exposure event only when the component actually renders, add sample-ratio-mismatch alerting, auto-expire experiments, and run a per-experiment performance budget check.
Trade-offs. Client-side assignment is far easier to ship and works with static caching, at the cost of flicker, CLS and INP. Server-side is correct and couples experiments to the render path and the cache key.
Follow-ups. How do you A/B test without a flash of the wrong variant? How do you stop flags accumulating? How would you roll back a frontend change in 60 seconds?
13.3 Frontend rollback
What it is. Reverting the served version without breaking users who are mid-session.
Example. Users holding old HTML will request old chunks. Keep the previous N builds' assets on the CDN (immutable hashed assets make this free) and handle ChunkLoadError with a cache-busted retry then a reload.
Trade-offs. Flipping a feature flag is faster and safer than a deploy rollback, but only covers code that was behind a flag — which is an argument for flagging more than feels necessary.
Follow-ups. What's different about rolling back a frontend vs a backend service? What breaks if you purge the CDN on every deploy?
13.4 Observability
What it is. Knowing what real users experience, and being able to attribute it.
Example.
- RUM —
web-vitalsplus custom timings (time-to-search-results, time-to-first-price), beaconed withsendBeacon. Segment by route, device class, country, connection and build id — the build id is what lets you blame a deploy. - Error tracking — Sentry with source maps uploaded in CI, release tagging, grouping rules, and a per-release error-rate gate. Filter browser-extension and third-party noise or the signal drowns.
- Tracing — propagate a trace header from the browser (OpenTelemetry browser instrumentation) so a slow page can be followed into backend spans. This answers "is it us or the API?"
- Synthetic monitoring — the booking funnel from several regions, as the early warning before RUM volume accumulates.
Trade-offs. RUM is truthful and lagging; synthetic is fast and artificial. You need both, and the cost is two systems to maintain and reconcile.
Follow-ups. How do you attribute an LCP regression to a specific release? What do you do about errors from browser extensions? How much RUM do you sample and why?
13.5 SLOs and error budgets
What it is. A target for a user-facing metric over a window, with a budget for failing it.
Example. "p75 LCP on search ≤ 2.5s over 28 days"; "JS error rate ≤ 0.5% of sessions"; "booking funnel success ≥ 99.5%". Attach a written policy: burn the budget and feature work pauses for reliability work.
Trade-offs. SLOs only work if someone will actually honour the policy; an SLO with no consequence is a dashboard. Setting them too tight produces alert fatigue.
Follow-ups. What would you set as the three SLOs for a frontend platform team? What's the difference between alerting on a threshold and on burn rate?
13.6 Incident handling
What it is. Detect → mitigate → diagnose → review.
Example. Alert on SLO burn rate rather than raw spikes; mitigate before diagnosing (flip the flag, roll back the CDN pointer); diagnose afterwards; run a blameless postmortem whose action item removes the class of failure, not just the instance.
Trade-offs. Mitigating first sometimes destroys the evidence — so capture the trace, the build id and a HAR before you roll back.
Follow-ups. Walk me through a frontend incident you owned. What was the prevention item? How do you know it worked?
14. Accessibility and internationalization
14.1 WCAG essentials
What it is. POUR — perceivable, operable, understandable, robust — with WCAG 2.2 AA as the usual bar.
Example. The concrete AA requirements you'll be asked about: 4.5:1 contrast for body text (3:1 for large text and UI components), a visible focus indicator, 24×24 CSS px minimum target size (44×44 is the stricter AAA/mobile guidance), no information conveyed by colour alone, and 200% zoom without loss of content.
Trade-offs. Meeting contrast ratios constrains brand palettes; the resolution is to bake compliant pairs into semantic tokens so product teams cannot pick a failing combination.
Follow-ups. What's the difference between AA and AAA? How do you handle a brand colour that fails contrast?
14.2 Semantic HTML and ARIA
What it is. Native elements carry role, keyboard behaviour and focus for free. ARIA only adds semantics — it adds no behaviour.
Example. <button> gives focusability, Enter/Space activation and the right role; <div onClick> gives a bug report. The first rule of ARIA is don't use ARIA.
Trade-offs. Custom components sometimes need ARIA because no native element exists (combobox, tabs, tree) — and then you own every keyboard interaction yourself.
Follow-ups. When is aria-label wrong? What does role="presentation" do? Why is aria-hidden on a focusable element a bug?
14.3 Keyboard and focus management
What it is. The hard part of accessibility: what has focus, where it goes, and what is announced.
Example. Focus trap inside a modal, return focus to the trigger on close, aria-live="polite" region for async results ("24 hotels found"), skip links, logical tab order, and no positive tabindex.
Trade-offs. aria-live announcements that fire on every keystroke are worse than none — debounce them and announce results, not progress.
Follow-ups. How do you make an autocomplete accessible? (role="combobox", aria-expanded, aria-controls, aria-activedescendant — the hardest standard pattern.) Where does focus go after deleting a row?
14.4 Testing accessibility
What it is. A layered approach, because no single method is sufficient.
Example. axe-core in component tests catches roughly a third of issues; add keyboard-only test runs, a screen-reader pass (VoiceOver/NVDA) on critical flows, and periodic audits with disabled users.
Trade-offs. Automation is cheap and shallow; manual testing is expensive and real. The platform play is to enforce axe on every design-system component so product teams inherit compliance.
Follow-ups. What can't axe detect? How would you stop accessibility regressing across twelve teams?
14.5 Internationalization architecture
| Concern | Approach |
|---|---|
| Message catalogues | ICU MessageFormat for plurals, gender and select; keys namespaced by feature; never concatenate sentences |
| Loading | Split translations per route and locale; load the active locale only — shipping 40 locales to every user is a common bundle bug |
| Dates, numbers, currency | Native Intl.DateTimeFormat, Intl.NumberFormat, Intl.RelativeTimeFormat, Intl.PluralRules; Temporal as it lands |
| Wire format | Send ISO/UTC; format on the client in the user's locale and timezone |
| Currency | Never string-template a price — locale controls symbol position, grouping and decimals; conversion server-side with a rate timestamp |
| Timezones | Hotel check-in dates are local calendar dates, not instants — a genuine travel-domain trap |
| Pseudo-localisation | A build-time pseudo-locale ([!!! Ṡéárçh !!!]) to catch hardcoded strings and layout breakage in CI |
| Text expansion | German and Finnish run 30–50% longer than English — never fix widths to English text |
| RTL | CSS logical properties (margin-inline-start, inset-inline), dir="rtl" on <html>, mirrored icons, visual regression in both directions |
| SEO | hreflang alternates per locale, localised URLs, correct lang attribute |
| Workflow | Extract keys in CI, push to the TMS, fail the build on missing keys for launched locales, fall back to English with a logged warning rather than showing the key |
Trade-offs. ICU is verbose and is the only sane way to handle languages with six plural forms. Server-side locale detection is accurate and fragments the CDN cache; client-side is cacheable and causes a flash.
Follow-ups. How do you format a price for 40 markets? Why can't you concatenate "Showing " + n + " results"? How would you support RTL in an existing design system? What breaks when a translator writes a string 60% longer?
15. Applying it
15.1 A system-design framework
Use the same eight steps every time, and announce them at the start.
- Clarify and scope (5 min). Users and devices, locales and markets, logged-in or anonymous, SEO needed, scale (QPS, catalogue size), and what is explicitly out of scope.
- Define success metrics (2 min). Two user metrics (p75 LCP, INP), one business metric (search→book conversion), one engineering metric (build time or deploy frequency).
- Sketch the architecture (8 min). Client, CDN/edge, BFF, backend services, third parties. Label arrows with protocol and cache policy.
- Pick the rendering strategy (5 min). Segment the page by volatility.
- Component hierarchy and state (8 min). Mark which state is URL, server cache, local, global. Name the data-fetching and Suspense boundaries.
- Data layer and contract (5 min). Endpoint shapes, pagination, error model, optimistic updates, cache keys and invalidation.
- Cross-cutting concerns (7 min). Budget, a11y, i18n, security headers, analytics and experiments, error/empty/loading states, offline.
- Trade-offs and next steps (5 min). Two rejected alternatives and why. End with the metric you'd watch after launch.
15.2 Worked scenarios
Hotel search results page. Streaming SSR at the edge for shell and first cards; filters, dates and occupancy in the URL; results in a server cache keyed by the normalised filter object; prices fetched separately and never cached; virtualised list with content-visibility: auto; AVIF with fixed aspect ratios; fetchpriority="high" on the first card image; map library loaded on interaction; filter changes in startTransition with AbortController cancellation; cursor pagination because offset breaks on live inventory. Rejected: client-side filtering (only works for small result sets) and a single global rendering strategy.
Image delivery pipeline. One high-res master; an image CDN derives variants from URL parameters (?w=800&fm=avif&q=70), cached immutably at the edge; a design-system <Image> component generates srcset/sizes and enforces width/height, loading, decoding and fetchpriority; LQIP inside a fixed aspect-ratio box so CLS stays 0. Guardrails: CI fails any raw <img>, a budget on LCP image bytes, RUM reports the LCP element type.
Design system rollout to 12 teams. Phase 0 audit — script the codebase to count distinct button implementations and colour values; that number creates the mandate. Phase 1 tokens only, adopted by codemod. Phase 2 the ten highest-traffic primitives on headless libraries with a11y and visual regression from day one. Phase 3 migration via codemods plus pairing, with a per-team adoption dashboard. Phase 4 enforcement via lint rules. Risks: the library becoming a bottleneck (contribution model, documented escape hatches) and version skew in a federated setup (singleton shared scope, compatibility matrix).
Offline-capable booking flow. Service Worker with precached shell, network-first HTML, network-only for prices; funnel state in IndexedDB keyed by a draft booking id; a request queue with Background Sync and a client-generated idempotency key so a retry cannot double-book; explicit offline UI rather than optimistic success; on reconnect, re-validate price and availability server-side and show a diff; a remote kill switch for the worker.
Real-time price updates. SSE for one-way pushes (simpler, auto-reconnect) or WebSocket if you need client→server too; subscribe per visible card via IntersectionObserver so offscreen items unsubscribe; coalesce bursts into one startTransition flush per animation frame; exponential backoff with jitter plus a sequence number to detect gaps and refetch; animate price changes without moving scroll position, and require confirmation if the price changed between selection and payment.
Experimentation platform. Covered in 13.2 — server-side assignment, variant in the cache key, typed variant names generated from the experiment registry, exposure on render, SRM alerting, auto-expiry, per-experiment performance budget.
15.3 Sources and further reading
MDN Web Docs · web.dev · Refactoring Guru · Awesome Scalability