Site search
@pagedeck/search indexes the text of every page a build renders and writes the
index as static JSON files beside the pages. A search island reads those files
in the reader's browser. Nothing runs on a server, and a page that does not
render the island ships none of its JavaScript.
The package has two halves:
@pagedeck/searchgivesdefineSearch(), the adapter a site declares asbuild.search. It runs inpagedeck build.@pagedeck/search/islandis the search box, and@pagedeck/search/queryis the query code the island calls. They run in the browser.
Declaring the index
import { defineSearch } from "@pagedeck/search";
// In the site's config, inside `build`:
search: defineSearch(),defineSearch() takes no options. The tokenizer runs twice, at build time over
the pages and in the browser over what the reader types, and the shard size is
written into the files the browser reads. An option could make the two halves
disagree, and a query that finds nothing looks the same as a page that does not
contain the word. So there is nothing to configure.
A site that does not declare build.search gets no index, and its build output
is the same as before the field existed.
What the index contains
For each page, three fields:
- The title: the
titlethe site'sheadcallback returns for the page. A page with no title is indexed without one. See Page head. - The headings: the text of each
<h1>to<h6>in the page. - The body: the rest of the page's text.
The text comes from the page's own rendered tree, which is what the build puts
inside <main>. The document around it is not read: not the <head>, not the
script tags, and not the chrome a site declares with build.chrome, which the
build writes before and after <main>. A header or footer declared as chrome is
therefore never indexed. See
Page body.
Inside the page's tree, the index leaves out:
- The contents of
<script>,<style>and<template>elements. - An element with the
hiddenattribute, and everything inside it. Any value hides,hidden="false"included, as it does in a browser. - An element with
aria-hidden="true", and everything inside it. Other values, such asaria-hidden="false", do not hide. - An element that carries the attribute named by
SEARCH_ATTRIBUTEwith the value"ignore", and everything inside it. See the next section. - Comments and the doctype.
An inline style="display:none" does not hide text from the index, and
neither does a class that hides an element. The index reads the page's markup
and not its stylesheets.
A heading inside a left-out element opens no section. Its text is not indexed.
Keeping template text out of the index
Some text inside <main> is not the page's own words: a navigation that lists
every page's title, a previous and next pager, a label that is the same on
every page. If it is indexed, a search for one page's name finds every page
that links to it.
SEARCH_ATTRIBUTE, exported from @pagedeck/core/tree, is the name of the attribute
that keeps such an element out of the index. Give it the value "ignore". The
element and everything inside it are left out of the index, and nothing else
changes: readers and assistive technology get the element as before. This site
spreads it onto its navigation, its pager and its source line:
import { SEARCH_ATTRIBUTE } from "@pagedeck/core/tree";
export const UNINDEXED = { [SEARCH_ATTRIBUTE]: "ignore" } as const;
<nav className="fw-docnav" aria-label="Documentation" {...UNINDEXED}>Only the value "ignore" has this effect. An empty value or any other value is
indexed as usual.
The attribute is for a site's templates, not for content. unescapedHtml
from @pagedeck/core/tree, the way rich text reaches a page, removes the attribute
from content whatever its value. A rich-text field therefore cannot hide its
own text, or the rest of the page, from search. The same removal applies to
text: a page rendered through unescapedHtml, like this one, cannot show the
attribute's name. That is why this page names it through the constant.
Content that reaches a page another way, such as a site's own
dangerouslySetInnerHTML, is not changed.
How text becomes terms
- Text is lowercased and normalized to Unicode NFC.
- A term is a run of Unicode letters, combining marks and numeric characters
(digits, and characters such as
²andⅫ). Anything else separates terms. - Nothing is stemmed and no stop words are dropped.
loadersandloaderare different terms. The query side covers part of this with a prefix match; see "How a query is answered" below. - Chinese, Japanese and Thai text is not split into words. A run of those characters is one term.
Where the index is written
The build writes the index into the output directory, one directory per locale:
/search/<locale>/index.json: the format version and the range of terms each shard holds. It is the first file a query fetches./search/<locale>/documents.json: one record per page, with its path, the URL it is served at, and its title./search/<locale>/terms-0000.json,terms-0001.jsonand so on: the shards. Each holds a sorted range of terms and, for each term, the pages that contain it, how often, and in which fields.
A shard holds about 64 KiB of JSON: a term goes into a new shard when adding it would take the current one past 64 KiB. A single term whose entries are larger than that gets a shard of its own, and that shard is larger.
On a site that publishes to more than one domain, each domain's output tree
gets its own /search/ directory with the locales of the pages in that tree.
The index files are ordinary output files. They are hashed and recorded in the
build manifest, so a deploy uploads and prunes them with the pages. If a page or
another file of the build is already at one of these paths, the build fails and
names the path. A page at /search itself is fine: it is written to
/search/index.html, which is not an index path.
On an incremental build, a locale directory is written again, whole, only when one of its pages was rendered again or removed. The other directories are carried over unchanged.
The search island
Declare the island in the site's components by its package specifier, with
hydrate: "idle". Hydrating fetches no index file, so the island can be ready
before the reader reaches it without costing a request:
// In `components`:
search: { path: "@pagedeck/search/island", hydrate: "idle" },Then place it as a node of the page's entry tree, as the sites in this repository do:
// SEARCH is the island's registry name, "search".
{
component: SEARCH,
props: {
locale: "en",
label: COPY.search.label,
emptyLabel: COPY.search.emptyLabel,
},
},The island takes three props, and all three are required:
| Prop | What it is |
|---|---|
locale |
Which locale's index to search: the <locale> directory under /search/. |
label |
The text of the input's <label>, and the accessible name of the results list. |
emptyLabel |
The message shown when a query finds nothing. |
The island renders a search input with role="combobox" and, when a query
matches, a list with role="listbox" whose options are links. Each result
links to the URL the page is served at, and shows the page's title, or its path
when the page has no title.
- The arrow keys move through the results while focus stays in the input.
- Enter follows the active result.
- Escape closes the list and keeps the query in the input.
- A query that finds nothing shows
emptyLabelin arole="status"paragraph.
If an index file cannot be fetched, the island reports the error to the
browser console and shows emptyLabel. The message names the file and the
status.
Fold strategy does not change an idle island. See
Fold strategy.
Loading on focus
The island fetches nothing when it hydrates. The requests start when the reader uses it:
- Focus fetches
index.json, the shard ranges, and no shard. A reader who focuses the box and leaves has paid for that one small file. - Typing fetches only the shards whose range can hold the words typed. A word outside every range fetches no shard at all.
- The first match fetches
documents.json, once.
Every file is fetched at most once per page view. Two keystrokes that need the same shard make one request.
The index files are requested from the origin the page is served from, at
/search/<locale>/. A site with a Content Security Policy needs connect-src
to allow that origin. This site's policy sets connect-src 'self'.
How a query is answered
- The query is split into terms by the same rules as the index.
- Every term must be on a page for the page to match.
- The last term matches as a prefix, because the reader may still be typing it. The terms before it must match exactly.
- For each term, a page scores the number of times the term occurs on it, multiplied by the weights of the fields it occurs in, added together: 10 for the title, 4 for a heading and 1 for the body. The page's score is the sum over the query's terms. A title match weighs more than a body match, but a page that uses a word often in its body can still rank above a page with the word once in its title.
- Results with equal scores are listed in path order.
An index written in a format version the query code does not read is refused with an error, not guessed at.