PROTOCOLS-FUN . REQUEST PATH

Anatomy of a Web Request

The stages of an HTTPS page load, from URL parse to render.

section
keystroke, enter
00URL PARSEbrowser
scheme :// host : port / path ? query # fragment
#fragmentnever sent
schemepicks the default port
hostwhat gets resolved
querymay or may not be keyed
▾check the HSTS preload list before anything touches the network
01HSTS PRELOADbrowser
A list compiled into the browser binary. Consulted before DNS, before any packet.
preload hithttp rewritten to https
preload missplaintext, then upgrade
▾hand the request to the service worker, if one claims this scope
02SERVICE WORKERbrowser
Author-written JavaScript with a fetch handler. It runs before the HTTP cache, not after it.
fetch handlerauthor's code decides
cache 1Cache API
Cache APIHIT: zero network
▾worker passes through, or none is registered
03BROWSER CACHEbrowser
Two distinct caches, and the pair people mean when they say clear your cache.
freshnessmax-age / Expires
revalidationETag / If-None-Match
cache 2memory cache
memory cacheHIT: done
cache 3HTTP disk cache
disk cacheHIT: done, or 304
▾cache miss: the host name now has to become an address
04DNS RESOLUTIONbrowser + os + network
A chain of caches of its own before any authoritative server is asked.
browser DNS cachefirst stop
OS resolver cachethen the stub
recursive resolverISP, or DoH / DoT
root, TLD, authoritativethree questions, three answers
CNAME into the CDNthe CDN takes over here
the packet leaves the machine
05TRANSPORTnetwork
TCP's three-way handshake, or QUIC folding transport and crypto into one flight.
TCPSYN, SYN-ACK, ACK
QUICtransport + crypto together
▾negotiate keys, identity and protocol
06TLSnetwork
TLS 1.3. One round trip, or zero on resumption, and the zero has a replay problem.
ClientHelloSNI, ALPN, groups, key_share
JA3 / JA4 fingerprintcomputed here
0-RTT resumptionreplay unsafe
ECHencrypts the SNI
OCSP staplingsaves a fetch
cache 4proxy cache
corporate / ISP proxymostly dead
▾the connection terminates at a CDN edge, not at the origin
07EDGE POPcdn
The edge does the security work first, then decides whether it already has the answer.
scrubbingabsorb L3/L4
TLS terminationconnection ends here
WAFsignatures + anomaly score
bot managementJA4, HTTP/2 frames, behaviour
cache key constructiondecides what counts as the same request
cache 5edge cache
edge PoPHIT: served here
▾edge miss goes to a parent tier, not to the origin
08PARENT / SHIELDcdn
Tiered distribution. Most explainers skip this layer, which is why origin load surprises people.
tiered distributionfewer nodes, longer TTL
cache 6shield cache
shield tierHIT: origin idle
origin leg may ride an overlay path, not default BGP
09ORIGINyour infrastructure
Only reached when every cache above it missed.
load balancerpicks a backend
reverse proxynginx / Varnish
app serverruns the code
cache 7reverse proxy cache
nginx / VarnishHIT
cache 8application cache
Redis / MemcachedHIT
cache 9database
buffer pool / query cachethe floor
▾the response now travels back up every layer it came down
10RESPONSE RETURNevery layer
On the way back, each layer independently decides whether it is allowed to keep a copy.
Cache-Controlthe main instrument
Varyjoins the key
Surrogate-ControlCDN only, then stripped
▾bytes arrive, and the browser starts building a page
11RENDERbrowser
Parse to pixels. A fast network can still produce a slow page here.
HTML parse to DOMincremental
preload scannerruns ahead
CSS is render-blockingCSSOM gates first paint
sync script is parser-blockingasync / defer release it
layout, paint, compositegeometry, pixels, layers
LCP / CLS / INPwhat is measured

Resolution walks a chain of caches to an authoritative server. For a CDN-fronted site that server returns a CNAME into the CDN's zone, and the CDN's nameservers choose the edge in that second lookup.

the chain, and where it stops early
04resolution orderfirst hit wins
browser cachechrome://net-internals
OS resolver cachesystem wide
hosts filestatic override
stub resolverformats the query
recursive resolverISP, or DoH / DoT
rootpoints at the TLD
TLDpoints at the zone
authoritativehands off to the CDN
the CNAME chain a real lookup walks
name www.example.com CNAME example.com.cdn-provider.net CNAME a1234.dscx.akamaiedge.net A 203.0.113.42
Addresses here use the RFC 5737 documentation range. The TTL on the last hop is typically seconds to a couple of minutes, which is what gives the mapping system room to move traffic.
two mechanisms, and the summary everyone gets wrong
DNS mappingAkamai
selectionin the DNS answer
blind spotsees the resolver
ECSpasses a client prefix
end user mappingmaps the client
AnycastCloudflare, Fastly
selectionby BGP routing
blind spotAS hops, not latency
route changecan move a live flow
simplicityno DNS trickery
the summary that is too clean. “Akamai uses DNS, Cloudflare uses anycast” is the usual line, and it is only half right. Akamai does select edges through DNS mapping, but the nameservers you ask are themselves fronted by IP anycast, with the mapping decision made behind that by servers picked to sit near your resolver. Both vendors use anycast; they use it at different layers. Cloudflare anycasts the edge that serves your content, Akamai anycasts the nameservers that tell you which edge to use. The real distinction is where the per-client decision is made: in a computed DNS answer, or in the routing table.
the leg you do not think about

An edge miss goes to the origin over whatever path BGP picks. BGP minimises the number of networks traversed and ignores latency and congestion, so CDNs offer overlay routing that measures paths across their own network and forwards over the fastest one (Akamai SureRoute, for example).

SureRoute only runs when caching cannot. It applies to content that cannot be cached. When the edge recognises an object as cacheable it disables the path racing, because a cache hit beats any route. So the feature that sounds like it should help everything is, by design, only for the requests that had to go all the way back.

Connection setup costs round trips, and round-trip time is bounded by distance. The ClientHello that starts encryption also identifies the client's TLS library.

round trips before the first byte of response
pathtransportcryptototal before requestnote
TCP + TLS 1.2123The old baseline, and why connection reuse mattered so much.
TCP + TLS 1.3112TLS 1.3 removed a full round trip by sending a speculative key share in the first flight.
QUICcombined1Transport and crypto negotiate together over UDP.
QUIC 0-RTTresumed0Application data rides in the first flight. Replayable, so idempotent requests only.
the ClientHello, in the clear

The first message of a TLS handshake is unencrypted, since the two sides have no shared key yet. Hover any field.

0x0303 cipher_suites[18] server_name alpn supported_groups key_share supported_versions GREASE
version and framing negotiated capability privacy sensitive deliberately ignored
the handshake identifies the client. Taken together, the cipher list, the extension list, the curves and their order describe the TLS library and build with enough precision to fingerprint it. That is JA3, and its successor JA4. It identifies the stack, not the person. That is enough to separate a real browser from a scraper wearing a browser's User-Agent. The same value turns up again on the DEFENCE tab.
0-RTT, and the reason it is fenced off
replay is not a theoretical risk. Data sent in a 0-RTT flight is encrypted under a key derived from a previous session, with nothing in it that binds it to this connection. An observer can capture that flight and send it again. For a GET this is harmless. For anything that changes state it is not, which is why the rule is that 0-RTT carries idempotent requests only, enforced by the application and not by the protocol.

Nine caches can hold a copy of a response, each with its own key and expiry rules. Fixing a stale page means finding which layer holds the old copy, and each one is cleared differently.

the nine layers
#layerkeyed byinvalidated bytypical TTL
1Browser memory cacheURL, per tabtab close, navigationseconds to minutes
2Service worker Cache APIwhatever the worker's code decidesthe worker's code, explicitlyauthor controlled
3Browser HTTP disk cacheURL + headers named in Varymax-age, revalidation, evictionminutes to a year
4Corporate / ISP proxyURL + Varys-maxage, no-storemostly moot under TLS
5CDN edge PoPthe cache key: URL plus selected headers, cookies and query parametersTTL, purge or invalidate APIseconds to days
6CDN parent / shieldsame key, fewer nodessame as edgeusually longer than edge
7Origin reverse proxyconfigured keyTTL, purge, banminutes to hours
8Application cachewhatever the application inventsapplication logic, TTL, explicit evictionseconds to hours
9Databasequery plan, buffer poolwritesimplicit
Each row is a different cache, with its own key and its own way of being cleared. A browser “clear cache” reaches rows 1 and 3. It does not touch the service worker, it does not touch the CDN, and it certainly does not touch your application's Redis. Debugging staleness means naming the layer first.
cache-control, and the pair that gets confused
CCresponse directivesRFC 9111
max-agefreshness, everyone
s-maxageshared caches only
public / privatewho may store it
no-cache vs no-storenot the same thing
must-revalidateno stale on error
immutablenever revalidate
stale-while-revalidateserve now, refresh after
stale-if-errorstale beats a 5xx
the cache key, and what joins it
KEYkey constructionwhere poisoning lives
The default is method plus URL. Everything beyond that is configuration.
Varyrequest headers join the key
query parametersinclude, exclude, normalise
cookiesusually a mistake
unkeyed inputthe poisoning primitive
Surrogate-Control and CDN-Cache-Control are aimed at the CDN rather than the browser. Surrogate-Control is consumed and stripped by any CDN that honours it. CDN-Cache-Control is newer and is specified to survive: a CDN that does not use it passes it through, and one that does may still forward it so a second CDN downstream can read it. They let the edge and the browser be given entirely different policies for the same object, and they are invisible in a browser devtools panel, which makes them a recurring source of “but the header says”.

A CDN only protects traffic that passes through it. The matrix below is organised by where each attack is stopped.

attack matrix
attackstopped wheremechanism
Volumetric L3/L4 DDoSedge networkAnycast dispersion plus raw scrubbing capacity. You absorb it, you do not filter it.
L7 HTTP floodedgeRate controls per IP, per token, per path.
Credential stuffing, scrapingedge bot managementTLS fingerprint (JA3/JA4), HTTP/2 frame fingerprint, behavioural scoring, JS challenge, proof of work.
SQLi, XSS, traversal, RCEWAFSignature match plus anomaly scoring, usually an OWASP-derived ruleset.
Cache poisoningcache key designAn unkeyed input that still changes the response gets stored and served to everyone. Defence is strict key hygiene and normalisation.
Cache deceptioncache rulesRequest /account.php/x.css; a CDN caching by extension stores private data as a public asset. Defence is never caching on extension alone, and checking Content-Type.
Request smugglingedge to origin boundaryHTTP/1.1 Content-Length and Transfer-Encoding desync between two parsers. Native HTTP/2 end to end removes the ambiguity; an edge that downgrades to HTTP/1.1 on the origin leg reopens it.
Origin bypassorigin ACLIf the origin address is reachable directly, every control above is optional. Defence is allowlisting edge addresses, or mTLS.
DNS hijackregistrar, DNSSECRegistrar lock, and DNSSEC where it is deployed.
Subdomain takeoverDNS hygieneA dangling CNAME pointing at a deprovisioned service someone else can now claim.
a CDN in front of an origin is exactly the two-parser condition request smuggling needs. Smuggling requires two HTTP implementations that disagree about where one request ends and the next begins. An edge that terminates the client connection and forwards over a reused connection to the origin is that arrangement by construction. That is the normal arrangement, not an exotic one. HTTP/2 removes the CL/TE ambiguity only where it runs the whole way to the app server. Most edges still speak HTTP/2 to the client and downgrade to HTTP/1.1 behind it, and that downgrade is its own live attack surface, so treat “we use HTTP/2” as a question about the origin leg rather than an answer.
the one that is routinely forgotten. Every control on this page lives at the edge, and the edge is only in the path because DNS says so. If your origin answers on a public address, an attacker who finds it skips the WAF, the bot management, the rate limits and the scrubbing in a single step. Historical DNS records, certificate transparency logs and a careless mail server header are all common ways to find it. Allowlist the edge, or require mTLS, and treat a directly reachable origin as the vulnerability it is.
why the fingerprint from stage 06 matters here

A client can send any User-Agent string. The cipher and extension order its TLS library sends is fixed by the library build and much harder to change. A mismatch between the two is the detection signal; evading it means matching the fingerprint at every layer.

Everything from here runs locally in the browser. A fast network can still produce a slow page at this stage.

the critical rendering path
11parse to compositebrowser main thread
HTML parsebuilds the DOM
preload scannerfetches ahead
CSSrender-blocking
scriptparser-blocking
render treeDOM + CSSOM
layoutgeometry
paint, compositepixels, then layers
what is actually measured
metricmeasuresgoodthe usual cause when it is bad
LCPLargest Contentful Paint: when the biggest element in the viewport is drawn< 2.5sA slow origin response, or a hero image the preload scanner never saw.
CLSCumulative Layout Shift: how much content moves after it is drawn< 0.1Images without dimensions, and fonts or ads injected above existing content.
INPInteraction to Next Paint: the delay from an interaction to the frame that reflects it< 200msLong tasks holding the main thread so no frame can be produced.