Write a loader
A loader is the framework's whole interface to wherever your content lives. If you can list your content and detect what changed in it, you can write one.
The contract
A loader implements two methods, and may implement a third.
interface Loader<T> {
syncAll(writer: CollectionWriter<T>): SyncResult | Promise<SyncResult>;
syncSince(
writer: CollectionWriter<T>,
cursor: number,
): SyncResult | Promise<SyncResult>;
fetchOne?(id: EntryId): T | undefined | Promise<T | undefined>;
}The writer is collection-scoped, so a loader cannot write outside its own namespace:
interface CollectionWriter<T> {
upsert(entry: { locale: string; path: string; data: T }): void;
delete(id: { locale: string; path: string }): void;
}Nothing you write through the writer is applied while your method runs. The framework buffers the calls, validates every upsert against the collection's schema once you return, and then replays the whole batch in a single transaction together with the cursor. A loader that throws leaves nothing behind, because no transaction was ever opened.
Write syncAll
Fetch everything, upsert it, report what you wrote.
async syncAll(writer) {
const changed = [];
for (const item of await fetchEverything()) {
const id = { locale: item.language, path: item.slug };
writer.upsert({ ...id, data: item });
changed.push(id);
}
return { changed, deleted: [], cursor: Date.now() };
}cursor is yours. The framework stores the number and hands it back to
syncSince; it never interprets it. A revision number, an epoch millisecond
and a page offset are all fine.
Write syncSince
The same shape, filtered by the watermark you were given.
async syncSince(writer, cursor) {
const changed = [];
for (const item of await fetchChangedSince(cursor)) {
const id = { locale: item.language, path: item.slug };
writer.upsert({ ...id, data: item });
changed.push(id);
}
return { changed, deleted: [], cursor: Date.now() };
}Bias the watermark towards syncing too much rather than too little. A cursor that is slightly stale costs one delta that re-reads a few unchanged entries; a cursor that is slightly ahead costs an entry that is never synced again.
Report deletions
deleted names the entries your source says are gone. A source that remembers
its own removals — a delivery API with a removals endpoint — fills it from
there.
A source that remembers nothing cannot. A directory, a static export or a JSON dump leaves no trace of what it no longer holds, so walking all of it is the only removal report there is. Say so, and the framework removes every entry of the collection your sync did not report as changed:
return { changed, deleted: [], authoritative: true, cursor: Date.now() };Only from syncAll. A syncSince reports a delta, so it never mentions the
entries that did not change — returning authoritative from one is refused,
because pruning against it would delete the unchanged remainder of the
collection.
Entry ids are identifiers, not paths
EntryId is { locale, path }, and both fields are identifiers. A sync refuses
every upsert whose id holds a .., a leading /, a backslash or a NUL —
spelled literally or percent-encoded — because a loader is free to build a file
path or a URL out of an id, and one that resolves away structure would let
whoever supplies it read outside your content.
The whole sync fails when one id is bad, and every offending id is named in the report. Deletes are exempt: a delete files nothing under an id, and a store written before this rule existed may hold a row only a delete can remove.
Declare a schema
A collection's schema field is required. Give it a Standard Schema — any
implementation — and every entry your loader writes is validated before it is
stored, with failures named by locale, path and field. Give it false and
nothing is checked. There is no way to skip the decision.
The schema's output type and your loader's entry type are two types, and they
do not have to agree. A collection is Collection<TOut, TIn>: TOut is the
entry as it is stored and read, TIn is what your loader produces, and TIn
defaults to TOut for the loader that already produces what the site reads. So
a schema may narrow what a loader hands over — turning an open record of
frontmatter into a closed one — and every page, query and extractor sees the
narrowed shape. Your loader is written against its own type and never learns
about the site's.
One exception, and the types enforce it: a collection that declares
schema: false stores what the loader hands over unchecked, so there is nothing
to narrow through and the two types are one. Such a collection cannot claim a
read type its loader does not produce.