Pagedeck

3. One canonical spelling for a path, minted once

Date: 2026-08-25

Status

Accepted. Implemented by #85 for the route table, and by #30 for the redirect and headers manifest, which normalizes every authored path through the same function rather than minting a second spelling. The three this was written as binding on have all since landed, and all three are bound by it: #37 (output trees), #40 (per-locale sitemaps — sitemap.ts consumes Page.output and calls the function for the one path it composes itself, its file's own) and #88 (href — templates lower onto #87's segment list rather than encoding of their own). Binding on any later emitter that writes a path.

This section states the decision's reach and is maintained as work lands, the way ADR-0001's names its implementors. What else an ADR may edit after the fact is convention rather than written rule, and the convention has run both ways: #148 edited ADR-0004's Status and its Consequences, while #133 rewrote ADR-0001's and this file's Consequences and left both Status sections alone. This edit takes the narrower of the two and touches Status only.

Amended by #607: a redirect target is looked up as a file before the policy spells it. planRouting first spells a redirect's to without a trailing slash, by normalizeOutputPath under "never", and looks that spelling up in the files the build emits into the rule's tree. If it names one, that spelling is kept. Only a target that names no file is spelled by the site's policy, as a page. This is not a second opinion about the spelling: every step is still canonicalizePath's, through the one normalizer, and an emitted file has one spelling, its own, which the policy was never an address style for. Before #607, "always" spelled /sitemap.xml as /sitemap.xml/, which names no file.

Context

A URL has many spellings. /café, /caf%C3%A9 and /caf%c3%a9 are one address to a visitor, and so are /~a and /%7Ea. Spec §11 asks that URL normalization mean "S3/CloudFront resolution quirks can't create duplicate-content URLs", and until #85 the route table still admitted every one of those pairs as two rows.

Leaning on the URL parser does not settle it. Measured on Node 24.18.1, new URL(route, origin).pathname encodes non-ASCII as UTF-8 escapes, but does not uppercase an escape's hex digits, does not decode an escape of an unreserved character, and passes a malformed escape straight through. That is one of RFC 3986 §6.2.2's three normalizations, so the parser leaves two of the three duplicate-content pairs above intact.

The spelling is not a local question. It has to agree across the route table, the links a page emits, its canonical tag, the per-locale sitemaps, the redirect and headers manifest, the edge config, and the S3 object key — café.html and caf%C3%A9.html are different keys to S3, and CloudFront forwards the encoded path. Any emitter that re-spells a path locally is free to disagree with all the others, and the failure it produces is a page that ships at an address nothing links to. Nothing detects that: each emitter is internally consistent, and the route table it disagrees with is a different file.

Decision

A path is spelled by canonicalizePath (packages/core/src/pages.ts) and nowhere else. Every emitter that writes a path — an output tree, a sitemap entry, a manifest key, a rendered href — calls it, or consumes a Page whose path and output already went through it. No emitter percent-encodes, decodes, cases, or re-joins a path of its own.

The canonical form is the RFC 3986 normalized URI form, §6.2.2, in three steps: non-ASCII becomes UTF-8 percent-escapes, every escape's hex digits go uppercase, and an escape of an unreserved character (ALPHA / DIGIT / "-" / "." / "_" / "~") decodes back to the character.

Only unreserved characters decode. / is reserved, so %2F stays escaped and /a%2Fb remains one segment. This is not a detail of the encoder — it is what keeps the function from inventing or destroying a path segment.

The function normalizes spelling, not structure. It does not resolve dot segments: called directly, canonicalizePath("/a/%2E%2E/b") gives /a/../b, because %2E is an escape of an unreserved character and decoding it is step three. A caller assembling a path by hand therefore resolves dot segments before calling it. Inside normalizeRoute that is already true — removeDotSegments runs first and reads the escaped spellings itself — which is why the route table never sees a /a/../b.

A path that is not a path is refused, as a ConfigError. A malformed escape (/a%2, /a%zz) is not an address two clients agree on; a query or a fragment (/a?b, /a#b) is not part of a path at all — ? and # are no pchar, and encodeURI passes both through, so refusing them is what stops a caller getting a query back from the function that is supposed to have settled the spelling; and a lone surrogate (/a<U+D800>) is no character at all, so encodeURI throws URIError on one rather than spelling it (#131). None of the four can reach a path except from a route callback, which makes it wiring rather than content — exit 2 under rule 7 of docs/error-messages.md. The tested order is the delimiter, then the escape, then the surrogate, in unusableReason's order: /x?token=…%2 cut at its % would quote the token (rule 6).

Non-ASCII paths are supported. No transliteration requirement: whether a site wants /café or /cafe is a content decision, and the build spells the result one way regardless of how the author typed it.

Consequences

Two spellings of one URL now land on one row of the route table, so a second entry claiming it is reported as the collision it is, with both claimants named, instead of shipping as two pages.

A path prefix is spelled, but its trailing slash is not a spelling. A header rule's prefix (#30) is a matcher, not an address: /docs/ scopes a rule to what is under /docs, and /docs also matches /docsearch. Running the trailing-slash policy over one does not canonicalize it, it changes what it matches — under "never" the slash is stripped and the rule silently widens onto a sibling the author scoped out, in all three of #33's compile targets at once, since a prefix is precisely the matcher they agree on. So normalizeOutputPrefix shares every step of the spelling with normalizeOutputPath — the leading slash, dot segments, slash runs, and canonicalizePath itself — and stops before the policy. That is one normalizer with two endings, not the second normalizer this decision rejects: the alternative is a caller applying the policy and then patching the slash back on, which is a second opinion about the spelling in the place this decision says there may not be one.

canonicalizePath runs inside normalizeRoute, after dot segments are resolved and slash runs collapsed, before the trailing-slash policy is applied. The order matters in one direction only: dotSegment reads %2E off the raw spelling to recognize /a/%2e%2e/b as /b, so it must keep running first. A refactor that reorders the two has a test in packages/core/src/pages.test.ts waiting for it.

The refusal is detected in collectPages's first pass, by unusableReason, rather than at the throw inside canonicalizePath. Both read the same MALFORMED_ESCAPE, PATH_DELIMITER and LONE_SURROGATE predicates, so they cannot drift about what "not a path" means; the first pass is what enumerates every unusable route in a run (rule 5) and classifies it (rule 7). The throw inside canonicalizePath is the guarantee an emitter calling it directly gets.

The function is idempotent, which is what lets a path be canonicalized again without a caller having to track whether it already was — collectPages normalizes output a second time after prefixing a locale, and that stays a no-op on the spelling.

#87 lets a route callback return the segments of a path rather than the finished string, and percent-encodes each of them on its own before joining them. That happens before this function sees the result, and stays per-segment: it is what stops a slug containing / from inventing a segment, which unusableReason cannot catch because it reads the finished route. unusableSegmentReason is the third reader of LONE_SURROGATE, and the only reader of that one predicate alone: encodeURIComponent throws on a lone surrogate too, so the segment form needs its own guard, taken before anything is encoded. The other two predicates do not apply to it — a delimiter or a malformed escape in a segment is encoded, not refused. Templates are #88's, and they lower onto the same segment list, so nothing gains a second place where a path is spelled.

Alternatives considered

Canonicalize to the decoded spelling (/café). Rejected. A URI path is ASCII by definition, so /café is an IRI. It is a spelling to show a human, not one an emission layer can write to an S3 key or a manifest.

Use new URL() as the canonicalizer. Rejected, on measurement. It performs one of the three normalizations and passes malformed escapes through, so it leaves the pairs this decision is about unmerged.

Refuse non-ASCII paths and require transliteration. Rejected as the framework overreaching on a content decision, and unnecessary — step one spells them unambiguously.

Normalize per emitter, at emission. Rejected. It is the failure mode in the Context above: every emitter is free to disagree, and nothing checks that they do not.