|
| 1 | +--- |
| 2 | +applyTo: "scripts/services/**/*.js" |
| 3 | +--- |
| 4 | + |
| 5 | +# Talk & speaker data – business logic |
| 6 | + |
| 7 | +All data access and domain logic lives in `scripts/services/`. Blocks consume these services and |
| 8 | +should **not** parse the raw JSON themselves. There are two data sources, both published by EDS: |
| 9 | + |
| 10 | +- **`/query-index.json`** – flat list of every published page with its metadata (one row per page). |
| 11 | +- **`/<year>/schedule-data.json`** – per-edition spreadsheet export describing the timetable. |
| 12 | + |
| 13 | +## Query index (`QueryIndex.js`, `QueryIndexItem.js`) |
| 14 | + |
| 15 | +`getQueryIndex()` fetches `/query-index.json` once and **caches the promise** (singleton). Tests |
| 16 | +reset it with `clearQueryIndexCache()`. Each raw row is mapped onto a `QueryIndexItem` |
| 17 | +(`Object.assign(new QueryIndexItem(), row)`), and the placeholder `default-meta-image.png` is |
| 18 | +nulled out. |
| 19 | + |
| 20 | +`QueryIndexItem` is the canonical page/metadata record. Fields are strings as stored; helper |
| 21 | +methods parse them: |
| 22 | +- `getKeywords()` / `getRobots()` → `parseCSVArray` (comma-separated). |
| 23 | +- `getTags()` → `parseJsonArray` (JSON array, falls back to CSV). |
| 24 | +- `getSpeakers()` → `parseCSVArray` of the `speakers` field. |
| 25 | + |
| 26 | +Speaker-specific fields: `affiliation`, `twitter`, `speaker-alias`, `uptoyear`. |
| 27 | +Talk-specific field: `speakers` (speaker **names** or speaker **document-names**). |
| 28 | + |
| 29 | +### Page-type detection is path-based |
| 30 | + |
| 31 | +`QueryIndex` classifies items purely by their `path` using regexes: |
| 32 | +- site root: `/^\/\d\d\d\d\/$/` |
| 33 | +- speaker page: `/^\/speakers\/.*$/` |
| 34 | +- talk page: `/^\/\d\d\d\d\/schedule\/.+$/` |
| 35 | + |
| 36 | +Key methods: |
| 37 | +- `getItem(path)` – exact path lookup. |
| 38 | +- `getAllSiteRoots()` – yearly editions, newest first. |
| 39 | +- `getAllTalks()` – all talk pages, sorted year-desc then title-asc. |
| 40 | +- `getTalkSpeakerNames(siteRoot)` – distinct sorted speakers of **main** talks (`schedule/<x>`). |
| 41 | +- `getLightningTalkSpeakerNames(siteRoot)` – speakers of **lightning** talks |
| 42 | + (`schedule/<x>/<y>`), **minus** anyone already in the main-talk list. |
| 43 | +- `getTalksForSpeaker(speakerItem)` – talks whose `speakers` include the speaker's title or |
| 44 | + document-name. |
| 45 | + |
| 46 | +### Speaker variant resolution (important domain rule) |
| 47 | + |
| 48 | +A speaker can appear at the same name across years with different affiliation/details. |
| 49 | +`getSpeaker(pathOrName, siteRootPath)`: |
| 50 | +1. If given a URL/path under `/speakers/`, returns that exact item. |
| 51 | +2. Otherwise matches speaker items whose `title` **or** document-name equals the input. |
| 52 | +3. Disambiguates multiple matches via `getMatchingSpeakerVariant`: sorts by `uptoyear` (items |
| 53 | + **without** `uptoyear` sort last = "current"), and picks the first variant whose `uptoyear` |
| 54 | + is absent or `>= requested year`. |
| 55 | + |
| 56 | +When you need the right speaker for a given edition, always pass the `siteRootPath` so this logic |
| 57 | +applies — don't pick the first match yourself. |
| 58 | + |
| 59 | +## Schedule data (`ScheduleData.js`, `ScheduleDay.js`, `ScheduleEntry.js`) |
| 60 | + |
| 61 | +`getScheduleData(url, forceReload)` fetches a yearly `schedule-data.json` and **joins it with the |
| 62 | +query index** to produce `ScheduleDay[]` → `ScheduleEntry[]`. |
| 63 | + |
| 64 | +Spreadsheet columns → `ScheduleEntry`: `Day`, `Track`, `Start`, `End`, `Entry` (title), |
| 65 | +`Duration`, `FAQ` (Q&A minutes), `Type`, `Speakers`. |
| 66 | +- Valid `Type`s: `day`, `talk`, `break`, `other`, `other_rating`. `day` rows are dropped from |
| 67 | + entries (only used implicitly). |
| 68 | +- `Start`/`End` are Excel/Sheets serial numbers → real `Date`s via |
| 69 | + `convertSheetDateValue` (UTC). Times are formatted UTC (`utils/datetime.js`). |
| 70 | +- A row is **invalid and skipped** if day/start/end/title/duration are missing/zero or the type is |
| 71 | + not recognised. |
| 72 | + |
| 73 | +### Talk rows resolve against the query index |
| 74 | + |
| 75 | +For `Type === 'talk'`, the `Entry` value is a reference to the talk detail page: |
| 76 | +- absolute URL/path → used directly; otherwise treated as a document-name under |
| 77 | + `/<year>/schedule/<ref>`. |
| 78 | +- If no matching query-index item exists, **the entry is dropped**. |
| 79 | +- The entry's `title` is taken from the index item (with `removeTitleSuffix`), and if the sheet |
| 80 | + has no `Speakers`, speakers are inherited from the index item. |
| 81 | + |
| 82 | +`ScheduleData.getTalkEntry(path)` finds the entry for a talk detail page (used by talk-detail |
| 83 | +blocks to show time/duration). |
| 84 | + |
| 85 | +`ScheduleDay` aggregates its entries' min `start` / max `end`. Parallel tracks are represented by |
| 86 | +the `track` number (track 1..n share a start time); grouping into parallel rows is done in the |
| 87 | +`schedule` block, not here. |
| 88 | + |
| 89 | +## Talk archive (`TalkArchive*.js`) |
| 90 | + |
| 91 | +`getTalkArchive()` builds a `TalkArchive` from the query index. It projects talks into lightweight |
| 92 | +`TalkArchiveItem`s (arrays already parsed) and **drops talks with no speakers**. |
| 93 | + |
| 94 | +- `TalkArchiveFilter` – `tags` / `years` / `speakers` (AND across categories, OR within a |
| 95 | + category). Serialised to/from the URL hash via `buildHash()` / `getFilterFromHash(hash)` using |
| 96 | + `category=val1,val2/...`. Only `tags`, `years`, `speakers` are valid categories. |
| 97 | +- `TalkArchive.applyFilter(filter)` recomputes `filteredTalks` and invalidates the full-text index. |
| 98 | +- `getFilteredTalksFullTextSearch(text)` lazily builds a `TalkArchiveFullTextIndex` over the |
| 99 | + currently filtered talks. The index is deliberately simplistic: it concatenates |
| 100 | + title/description/keywords/tags/speakers, lowercases, and does substring matching. |
| 101 | +- Filter option lists: `getTagFilterOptions()` / `getSpeakerFilterOptions()` (asc), |
| 102 | + `getYearFilterOptions()` (desc) — all distinct & sorted. |
| 103 | + |
| 104 | +## Link handling (`Link.js`, `LinkHandler.js`) |
| 105 | + |
| 106 | +`rewriteUrl` / `decorateAnchor` strip the host from internal adaptTo() URLs (so preview links stay |
| 107 | +on preview and live stays on live), open external links in a new tab, and mark `.pdf`/`.zip` links |
| 108 | +as downloads. `decorateAnchors(container)` is applied during `decorateMain`. |
| 109 | + |
| 110 | +## When extending the data layer |
| 111 | + |
| 112 | +- Add new domain logic here as a service/method with JSDoc, not inline in a block. |
| 113 | +- Reuse `utils/path.js` (`isUrlOrPath`, `getPathName`, `getDocumentName`, `getYearFromPath`) and |
| 114 | + `utils/metadata.js` (`parseCSVArray`, `parseJsonArray`, `removeTitleSuffix`) rather than new regexes. |
| 115 | +- Fetch via `utils/fetch.js` cache helpers. |
| 116 | +- Add a matching test under `test/scripts/services/` with sample JSON in `test/test-data/`. |
0 commit comments