The whole database ships as one FlatBuffers file, slackwater.tcdb, built from schemas/database.fbs. It is the format the JS module reads, the browser build fetches, and native apps can bundle: readers touch only the bytes they access, so an identity scan reads ids, names, and coordinates without decoding constituents, and a lookup by id reads one station's constituents without decoding anything else. The public API is unchanged and synchronous.
Any platform that needs station data otherwise has to parse all of it into memory to read any of it: object literals, JSON catalogs, and Codable records all decode the whole database up front, whether the consumer is a Node process, a browser, or an iOS widget ranking nearby stations. On the JS side that parse costs ~118 MB of V8 heap and OOMs memory-constrained devices (see signalk-tides#103); on a phone it is hundreds of milliseconds of decode before the first read. A file that readers access in place avoids the whole class of problem:
- Node reads it with
readFileSyncinto aBuffer, which is external memory, off the heap. - Browsers fetch it into an
ArrayBuffer. - Native apps can memory-map it; a widget touching one station faults in a few pages.
Per-record JSON decode was never the cost (4 ms for 200 records); the cost is decoding everything to read anything. FlatBuffers fixes that with zero-copy access, and one schema generates readers for TypeScript, Swift, and Kotlin.
- One root table with
version, astationsvector sorted byid(the FlatBufferskey, so lookup is a binary search inside the buffer), name tables for constituent and datum names, and separate tide/current route vectors sorted by slug. Stationcarries identity (id, name, kind, type, coordinates, timezone, locality, region, ISO region code, country, ISO country code, continent, context, cities, aliases), prediction data (constituents, datums, chart datum, tide/current offsets, epoch), and provenance (source, license, disclaimers).- Constituents are a struct vector: a
ushortindex into the root name table plus twofloatvalues — 12 bytes per constituent, contiguous, versus ~28 for a table per constituent. Float32 holds seven significant digits; sources publish three decimal places. Datums use the same struct-plus-name-table pattern. - Identical
sourceandlicensetables are written once and shared; repeated strings (timezones, countries, epochs) are deduplicated with shared strings. Station.attributionis the station's complete display notice — creator credit, licence and its URI, and an indication of modification — assembled by the builder fromsourceandlicenseso a reader in any language displays it without keeping its own table of sources or its own reading of what a licence requires. It is per station because licence varies within a single source, and a shared string, so the file holds one copy per distinct notice.- A
Currentsub-table andKindenum keep current stations in the samestationsvector and id lookup. It carries directions, mean flow, tide references, subordinate-current offsets, magnitude notes, and tide-derived rules. tide_routesandcurrent_routesmap stable slugs to one or more provider ids and retain former paths for redirects. Route lookup is a binary search; the full list is decoded only when requested.- Quality evaluation (
quality.json) rides along: the gate —acceptedandscore— is inline onStationso an identity scan can filter and rank from head pages, and the detail (factors, issues, reason, redundant) is aQualitysub-table written at the tail with the other lookup data. The file carries all stations, rejected ones included; readers apply theacceptedfilter. file_identifier "TCDB",file_extension "tcdb".
The finished file is laid out in two bands: every station table together at the head, and every station's lookup data together at the tail. The head band holds everything a scan needs — ids, names, coordinates, and the quality gate (accepted and score are inline scalars on the station table, not in the tail-side Quality detail) — so ranking all 8,000+ stations by distance and filtering to accepted ones touches only the head pages. The tail band holds what only a per-station lookup reads: constituents, datums, and the quality detail (factors, issues, reason), faulting in only when someone looks up that station. Identity is ~100 bytes per station and constituents ~1,300, so an interleaved layout would spread identity across thirteen times as many pages and an identity scan would fault in essentially the whole file.
The schema can't express this; the builder has to produce it deliberately. FlatBuffers writes buffers back to front — whatever is built first lands at the highest addresses — so buildDatabase (src/database/builder.ts) builds in two passes: first every station's constituents, datums, and quality tables (landing together at the tail), then every station table (landing together at the head). The natural refactor, one loop building each station's vectors right before its table, produces the interleaved layout — and nothing else fails when that happens. test/database.test.ts guards it by asserting the lowest prediction-data offset sits past the highest station-table offset.
src/stations.ts opens the buffer once, materializes identity fields into plain objects (allStations), and attaches lazy getters for harmonic_constituents, datums, epoch, and current details that decode one station's data from the buffer on access. Subordinate stations resolve their reference station's harmonics and datums; their own offsets still apply. No caching — a persistent cache on module-level objects would pull the data back onto the heap.
The quality gate comes from the same file: station.quality carries accepted and score eagerly (read inline during the identity scan) with lazy getters for the detail, qualityMap indexes those objects by id, and the stations export filters allStations on accepted. The module does not bundle quality.json; it stays in the repo as the artifact packages/stations/evaluate-quality.ts writes and the build embeds.
The bytes come from a per-build source behind the #slackwater.tcdb subpath import — each default-exports the bytes:
- Node (
src/database/bytes.node.ts):readFileSyncinto an off-heapBuffer. - Browser (
src/database/bytes.browser.ts):fetch(new URL("../generated/slackwater.tcdb", import.meta.url)); bundlers that understandnew URL(..., import.meta.url)copy the asset and rewrite the URL. - Workers (
src/database/bytes.worker.ts, selected by theworkerd/workerexport conditions): Cloudflare Workers can't construct file URLs fromimport.meta.urland disallowfetchduring module evaluation, so the database is inlined intodist/workeras a base64 literal by a build-time macro and decoded at module evaluation. The bundle is ~4.5 MiB compressed, which needs a plan with the 10 MiB script limit — free plans (3 MiB) have never fit this database in any format. The smoke test guards the compressed size so data growth surfaces at build time rather than at a consumer's deploy.
Both bundles resolve ../generated/slackwater.tcdb to one shared copy at dist/generated/slackwater.tcdb.
near/nearest/bbox use a bundled KDBush geo index (~66 KB, eager); search() builds a MiniSearch text index lazily on first call. Both are generated at build time from the raw data, sorted by id so index positions match the buffer's stations vector.
npm run build:
generate(scripts/generate-database.ts) — runsflatcto generate the TypeScript accessors intosrc/generated/fbs/, then buildssrc/generated/slackwater.tcdbfromdata/**/*.json(all git-ignored). Apretesthook runs it too.flatccomes from mise (.mise.toml).tsdown— buildsdist/node,dist/browser, anddist/worker(all ESM), resolving#slackwater.tcdbper build.copy-database— copies the file todist/generated/.tsc --noEmit— type-checks src and the tests/tools against the schemas.smoke(scripts/smoke.mjs) — imports all three built entries, checks a reference and a subordinate station resolve prediction data, and asserts the browser and worker bundles have nonode:fsand that the worker bundle neither fetches during module evaluation (fetch is poisoned for its import) nor usesimport.meta.url.
Releases attach the file as slackwater-<date>.tcdb alongside the TCD files.
buildDatabase(stations, { routes }) is exported so downstream generators can write filtered catalogues through the same builder and read them with the same generated readers.
kdbush and geokdbush are ESM-only packages (no require export), so a CJS build can't require() them without a double-wrapped-default interop bug (KDBush.from is not a function). All first-party consumers use ESM, so the package ships ESM only. Modern Node still lets require() load the ESM entry (require-of-ESM); older CJS-only tooling would need to import() it.