Static site for Open Bible Stories (unfoldingWord), built with Astro. Astro is used only for shared layouts — the output is plain HTML/CSS/JS with no client framework.
npm install
npm run dev # dev server at http://localhost:4321/
npm run build # builds the site into dist/
npm run preview # serves the built dist/
Requires Node 18.20+ (Astro's minimum).
The site is localized into 16 languages — the same setup as churchbased.bible. English lives at /, every other locale at /{lang}/ (e.g. /es/translate/). Locales: en, es, fr, hi, ru, ar, zh, sw, pt, id, vi, bn, ur, fa, my, nl — defined in src/i18n/config.ts.
src/i18n/en/*.json— the English source strings, one file per page plusui.json(nav, footer, shared strings).src/i18n/{lang}/*.json— translations. Missing keys fall back to English automatically (deep merge insrc/i18n/content.ts), so a partially translated locale still renders completely.src/components/pages/*.astro— the page markup, one component per page, rendered once per locale;src/pages/[lang]/holds the non-English routes.src/components/LanguageSwitcher.astro— the language dropdown in the header.- Non-Latin scripts (and Cyrillic) get self-hosted font packs, same policy as the Latin faces:
scripts/build-font-css.mjs(run automatically before dev/build) emits one stylesheet per script intopublic/assets/fonts/from the@fontsourcepackages, andBase.astrolinks only the current page's pack (seefontHrefFor()insrc/i18n/config.ts). The font-family overrides instyles.cssare keyed on<html data-script="…">, whichBase.astrosets from the locale'sscript(marketing pages) or the detected script of the content language (/l/{code}/hubs); ar/ur/fa render RTL. - Sixteen packs, not seven. The marketing locales need arabic, nastaliq, devanagari, bengali, myanmar, han and cyrillic; the content languages also publish in Odia (17 languages), Gujarati (14), Gurmukhi, Tamil, Telugu, Kannada, Malayalam, Lao and Ethiopic. Those 38 languages used to detect as script
otherand linked no pack at all, so their hubs and story pages rendered in whatever face the visitor's OS had — for most of these scripts, empty boxes.npm run check:scriptsfails when any published language has no pack, so the next new script gets noticed rather than shipped unreadable. Adding one means: a@fontsourcepackage and aPACKSentry inbuild-font-css.mjs, anhtml[data-script="…"]rule instyles.css, the value inFONT_PACKSand thescriptunion insrc/data/catalog.ts, and the Unicode block indetectScript()infetch-catalog.mjs. - The legal pages (
/license/,/privacy/,/terms-of-use/) and the 404 page are English-only — no/{lang}/variants, no switcher. - The common questions live on the contact page, above the contact button (
ContactPage.astro, strings infaq.json): eight questions whose answers open with the answer, each with a stable anchor slugged from the English question so/es/contact/#can-i-copy-printand/contact/#can-i-copy-printare the same anchor and a retranslation cannot break an inbound link. They sit before the button on purpose — most of what people write in to ask is answered there, so a reader gets their answer without waiting on a reply. There is deliberately no separate/faq/route: two pages carrying the same eight questions would be the duplicate-boilerplate caseFAQPageguidance warns about, and the nav stays the same five items in the same order on every page (WCAG 3.2.3). - The Discover language list is prerendered at build time (see "Catalog data" below), ordered by English name so scripts do not decide position, with the English name as each row's secondary line and a
data-searchindex (autonym, English name, alternate names, code) so "Swahili" finds Kiswahili. Each row is a plain link to that language's hub, so a click is an ordinary navigation. Its browse UI strings (status line, empty states) are localized viabrowse.*indiscover.json, passed todiscover.jsasdata-*attributes. Story pages take their labels fromstory.jsonand the hub's reader strings fromhub.json; the reader's own chrome (story picker, slide controls) and the Resources browser (resources.js) are still English-only. meta.title/meta.descriptionare per page and per locale. The homepage title is built from the localized tagline (ui.siteTitle — home.hero.title). Descriptions may contain{count}, which is replaced with the published-language count at build time.- The interface language is a visitor preference, not a property of the URL. On a first visit it comes from the browser (
navigator.languages, matched against the 16 locales); a pick from the language switcher overrides that and is remembered inlocalStorageunderobs.locale.public/assets/js/locale.jsapplies it two ways. On a locale-scoped page it redirects to that page's own variant, using the page'shreflangcluster so it can only go where a page was built — and only from the un-prefixed default-locale URL, because a prefixed URL (/fr/discover/) is somebody's explicit choice and is never redirected away from. On the hubs and story pages, which are one URL each, it swaps the chrome instead: those pages emit the strings they rendered as<script type="application/json" id="obs-chrome">, keyed by rendered text to the i18n key and the arguments it was filled with, and the script re-fills them from/assets/i18n/{locale}.json. Nav and footer hrefs are rewritten too, so the next click stays in the chosen language. - The switcher is on the hubs and story pages too, and especially there: those ~9,200 pages are where the interface language is inferred rather than chosen, so they were the only pages offering no way to correct it. A pick there applies in place rather than navigating — the swap maps from the page as built, so a second and third pick work — and the link's
href(Discover in that language) stays as the no-JavaScript fallback. The bundles get their ownCache-Controlinpublic/_headersand are prefetched from the locale-scoped pages, where the preference is already known; a bundle that arrives more than 1.5s in is dropped rather than changing the language under someone who has started reading. - Choosing a content language is not choosing an interface language. That is the point of the swap. A hub is baked in the locale
hubLocaleFor()picks, which is the best guess the build can make for someone arriving cold from a search engine — but a visitor reading the site in Spanish who opens/l/bho/from Discover used to find the nav, buttons and FAQ in Hindi. Now they keep Spanish. Only chrome moves: the autonym, story titles, extract, story text and<html lang>/data-scriptare the translation team's work in the content language and are never touched, which is also why the swapped regions carry[data-i18n-script]for their font rather than moving<html data-script>. npm run check:locales— verifies every locale has every file, structure matches English, and embedded links are untouched; prints an untranslated ratio per locale.npm run check:routes— run after a build: every sitemap URL resolves to a built page, every hub links its story pages, every URL in/llms.txtwas built (and every markdown mirror is listed), every chrome string the hubs and story pages register for the locale swap resolves in all 16 locale bundles,/discover/read/stays gone, and the output fits Cloudflare Pages' 20,000-file limit.npm run check:scripts— every published language has a font pack (see above).- Retiring a public route is a two-part change: delete the page, and add a
_redirectsrule. The second part leaves no trace in the build, so nothing catches it — which is how/discover/read/and its 16 locale variants started 404ing after the reader became static, and what Search Console's "Not found (404)" was reporting.check:routescarries an explicitRETIREDlist, each entry naming the commit that retired it, and asserts every one redirects with a 301, that every redirect target was built, and that no redirect chains. Add to that list when you retire a route.docs/search-visibility.mdhas the triage table for reading a coverage export. npm run check:locale-swap— run after a build: drives a real browser through the preference behaviour (browser language, an explicit locale URL, a switcher pick that applies in place and can be repeated, a content-language page that must not change the interface language, and the fonts — a swapped nav must get the swapped script's face while the autonym keeps the content script's). Needsplaywright-coreand a Chromium build (CHROME_PATH, or one underPLAYWRIGHT_BROWSERS_PATH); it skips rather than fails when there is none.check:routescovers the static half — every chrome key a page registers must resolve in all 16 bundles.npm run review -- --priority— a review sheet per locale for a native reader: key, English source, current translation, a blank column, and flags for placeholders, embedded links and strings still identical to English. The hub/story/FAQ strings have had no native review; seedocs/native-review.md.
scripts/fetch-catalog.mjs runs before every dev/build (npm run fetch:catalog to run it alone) and writes src/data/catalog.json. src/data/catalog.json is meant to be committed: it is the offline fallback and the incremental cache for the story and release-asset lookups (a CI build with no snapshot re-fetches everything). Refresh it with npm run fetch:catalog and commit when the catalog has changed materially; the build refreshes it anyway. src/data/catalog.ts is the typed accessor. Four sources feed it:
- DCS catalog search (required): every production-stage Open Bible Stories entry, grouped by language, minus the Theological Formation edition (paginating if the API truncates). Per entry: publisher, version tag, release date, and the downloadable release assets (PDF/EPUB/DOCX/zip/audio/video/YouTube).
- translationDatabase
langnames.json(optional): autonym, English name, alternate names, region, country codes, direction. The DCS manifest'slanguage_titleis often the English name ("Swahili", "Chinese, Simplified") or English in parentheses ("हिन्दी (Hindi)");chooseAutonym()swaps in langnames'lnin those cases and keeps the manifest title otherwise (rule and cases infetch-catalog.test.mjs). - Release history (optional): when the catalog advertises a PDF/audio/video the entry's own release does not carry (Door43-Catalog releases have no assets), the newest release of the repo that has it is used — the same resolution
discover.jsdoes at runtime. Cached per entry. - Story text (optional): the 50 story titles, the full frame text of each story and a short extract of story 1, read straight from each repo, trying each publishing team's entry until one yields stories. Two layouts: a Resource Container repo keeps one markdown file per story, and a legacy translationStudio repo keeps one directory per story with one file per frame, plus
title.txtandreference.txt. A tS story's frame count varies and is not knowable without a listing, so the repo's Gitea tree is read once and decides which files to fetch — one request per repo rather than a probe per frame. This is what gave the 18 tS languages (azb,bqi,bal,def,fas-x-eastfars,glk,haz,ilo,qxq,ar-xzn,ckb,lki,lrc,mzn,ntz,smy,tks,tly) real story pages, sitemap entries and markdown mirrors; before it, the build read their titles and stopped. Placeholder repos are not translations. 14 published languages ship a repo whose story 1 says only "Video only" (or "Audio/Video only") and links to a player — the translation exists as recordings, not text.isStubContent()(the same pattern and the same 300-character window the client reader used) makes the fetch step treat that as no stories, so it falls through to another publishing team and, if every team is a placeholder, the language gets no story pages, no markdown mirror and no sitemap URLs, keeps its video and download buttons, and its hub says so. Before this they became real story pages titled "Video only", with aCreativeWorkwhosetextwas that sentence.npm run check:routesasserts no built story page or hub extract carries the pattern. If the tree listing cannot be read the language degrades to titles and an extract with no story pages — a worse page, not a failed build, which is also why it has to be said: the fetch step prints a per-layout outcome (ts layout — 18 languages fetched, N story bodies read) and warns by name for any language that yielded titles but no bodies. Without that, a silent degradation is invisible in a deploy log. Results are cached in the snapshot and re-fetched only for entries whose release changed, so after the first build only what moved is fetched.OBS_CATALOG_STORIES=0skips this step for offline work. - Per-story media (optional):
audioByStory()/videoByStory()read the release history for numbered per-story files (en_obs_v6_23_360p.mp4) and map them onto story numbers — the highest bitrate for audio, the smallest rendition for video, since a story page is often opened on a phone on a slow connection. The entry's ownassetslist is capped at 40 files and cannot hold a 50-story set, which is why this reads the releases directly; the releases lookup is cached per repo per run, so it adds no requests for entries whose assets were just fetched.
A language's script is re-derived from its text on every build, including when the record is reused from the snapshot cache — widening detectScript() has to reach the ~200 languages that will not publish a new release soon.
Everything public that states a fact about the languages reads from the snapshot:
- the language count on the homepage and Why OBS (
lang-count.jsonly animates the served number; it no longer fetches anything), - the prerendered language list on
/discover/— each row links to the language's hub, anddiscover.jsonly filters those rows; it never fetches the catalog or changes the list (that would let it drift from the count), - the
/l/{code}/language hubs (see below),sitemap-languages.xml, and the JSON-LD on Discover and the hubs, {count}in localized meta descriptions.
So the public definition of "languages" is: distinct language codes with a published OBS translation in the DCS catalog — the same set Discover lists. Changing the catalog changes the number on the next build. npm test covers the grouping, langnames merge, story parsing, script detection and incremental story caching against scripts/fixtures/catalog-entries.sample.json (synthetic) and a fake fetch.
Failure policy: the snapshot is only overwritten by a complete, successful catalog fetch. npm run build fails only when the catalog fetch fails and no snapshot exists (Cloudflare then keeps the previous deployment live). Enrichment failures (langnames, a story file) never fail the build — the field falls back to the previous snapshot or null. npm run dev warns and writes an empty snapshot instead, and every page then states 0 languages — visibly wrong on purpose; there is no hardcoded placeholder count. OBS_CATALOG_ALLOW_EMPTY=1 forces a build through offline.
| URL | What it is |
|---|---|
/discover/, /{locale}/discover/ |
Discovery only — the language list, search, format filters |
/l/{code}/ |
The language page: names, codes, formats, downloads, story list, license |
/l/{code}/story-{n}/ |
One story: full text, illustrations, audio and video when they exist |
/changelog/, /changelog.xml |
Dated list of newly published and recently updated languages, and the same events as an Atom feed |
/llms.txt |
Index of the hubs and markdown mirrors, for fetchers that prefer markdown |
/content/{code}.md |
One language's complete story text as markdown |
/assets/i18n/{locale}.json |
One marketing locale's chrome strings, for the browser-side interface-language swap on the single-URL pages |
A story is addressed the same way in both places: /l/{code}/story-16/ as a page, #story-16 as the reader's fragment. No title slug — one would either be English in every language's URLs or percent-encoded nonsense for non-Latin scripts, and it would move the URL whenever a translation was revised.
There is no /discover/read/. That route hosted the JS reader; the reader itself now lives on the language hub, and each story additionally has its own static page.
Two ways to read, deliberately, with one canonical URL each:
-
On the hub —
/assets/js/reader.jsmounts in place when someone opens a story: slide-flip navigation, a story picker, per-story audio, YouTube where it exists, and a publisher chooser when several teams have published the language. It is not loaded on arrival — a hub is mostly visited for downloads or a single story, and mounting eagerly would fetch on all 214 hubs.It reads this site's own copy of the stories,
/l/{code}/stories.json(src/pages/l/[code]/stories.json.ts) — the same text the story pages are built from. Reading one story used to cost a catalog lookup, a repo listing, 50 title fetches and a fetch per story viewed; it is now one same-origin request, and the reader keeps working when Door43 is unreachable. Per-story audio is in that bundle too, so the player does not wait on anything remote.Door43 is still consulted for two optional extras — the YouTube embed and the in-reader PDF link, which come from the repo's release history — and for a publisher the local text does not cover, since the build reads story text from one entry per language. Both fail quietly. When a build produced no story text at all (a Door43 outage), the endpoint is absent and the reader falls back to its original live-catalog path.
-
The story pages — static HTML, one canonical URL per story, crawlable and readable with no JavaScript. Each story page has a "Read in the reader" link back to
/l/{code}/#story-N.
Every "read" affordance on a hub — the Read online / Listen buttons, "Read this story" under the extract, and every row of the story list — has that story page as its href. hub-reader.js intercepts the click and opens the story in the reader instead. So crawlers follow real URLs, a visitor without JavaScript lands on a real page, and everyone else reads in place; the fragment is kept in step (#story-N) so the story stays shareable and Back works. Modified clicks (⌘/Ctrl/middle) are left alone, because opening a story in a new tab should give the story page.
This is not the sleight of hand removed from Discover: there the href pointed at a fragment that served no content, whereas here the href and the reader show the same story.
src/pages/l/[code]/index.astro renders one static page per published language — the canonical public URL for "Open Bible Stories in {language}". /l/ keeps content languages out of the marketing-locale namespace (/es/, /fr/, …). Everything is in the initial HTML, ordered for a reader first: name (autonym as H1, or the English name when no autonym is known), then the actions — read online, listen (in the reader), watch, one primary PDF with its size, the full-audio zip with its size, EPUB/DOCX; the general mobile app is a note, not a per-language format — then the story-1 illustration with an in-language extract, the story titles that were actually read from the repo (usually all 50, never padded; a partial repo says {n} of 50 and points to Translate), and at the foot an "About this translation" block with code, alternate names, region, publishers (version, date, their other PDFs) and the license. Every read link — the buttons and the story list — points at a story page and opens the reader on click (see "Two ways to read" above). A story whose title was read but whose text was not is shown without a link, so a link never lands on a page that was not built. Every action is gated on the thing it opens, not on the catalog's format flags. lang.formats comes from language-level attachment_types, which says a format exists somewhere — not that this site can open it. Gating on it produced "Listen (audio)" on 92 hubs that resolved to a story page with no player on 78 of them (69 have nothing but a YouTube playlist, which is video), and "Read online" on the 14 placeholder languages. "Listen" now requires a per-story mp3 or the audio zip, "Watch" a real video file or playlist, "Read online" something readable; the media buttons are anchored at the player (#listen, #watch) so they do not land on a page identical to "Read online", and check:routes asserts a media control resolves to a page that actually has that player. A hub with no story pages at all — an unreachable repo, or a build with the story step skipped — points its read controls and story titles at the reader fragments (#story-{n}, #read) that hub-reader.js handles, with a <noscript> line naming the downloads and the Door43 source instead. This used to be the permanent state of the 18 legacy translationStudio languages, whose read control pointed at the hub's own URL: it reloaded the page and opened nothing. Those languages now have real story pages (see item 4 above); npm run check:routes still fails on a read control that self-links, and on a hub that lists titles with no way to read them. <html lang>/dir, data-script and the font pack follow the content language (script detected from the fetched text; Urdu → Nastaliq); the chrome and labels come from hubLocaleFor() (see below) and src/i18n/{lang}/hub.json. A four-question FAQ sits above the "About" block — what this is, whether it is the whole Bible, what the license allows, how to help finish it — so "may I print this?" is answerable without opening /license/; it is visible text only, since 214 near-identical FAQPage nodes would be boilerplate. Those answers are in the hub's chrome language, which is the content language only for the 17 codes the site is published in; #17 also asks for a definition in the content language on every hub, and that criterion is deliberately still open (see docs/native-review.md). The first answer's story sentence is chosen from what the build can actually render — all 50, some, or none — so a hub with one translated story does not promise fifty. Hubs have a self-referencing canonical, no hreflang cluster, no locale switcher. Their JSON-LD is a CreativeWork (inLanguage, license, isAccessibleForFree, translationOfWork, publisher, alternateName, dateModified, PDF/EPUB/audio-zip encoding, the extract as abstract, sameAs the Door43 repos and the language's YouTube playlist) plus an ItemList of the readable stories. Still no VideoObject here: what most languages publish is a playlist, which is not one video and has neither a duration nor an upload date — per-story recordings are real media and carry VideoObject on the story pages.
Which language the chrome is in (hubLocaleFor() in src/data/catalog.ts): the same language when the code is one of the 16 marketing locales or an ISO 639-3 alias for one (es-419 → es, swh → sw, fas-x-eastfars → fa); failing that, a locale in the same script — Hindi for the 88 Devanagari languages of India and Nepal, Bengali for the Bengali-script ones, Russian across Central Asia, Persian or Urdu for Arabic-script languages by country; failing that, English. Only 18 of 214 languages share a primary subtag with a marketing locale, so without the script step 196 hubs would show their FAQ, format labels and nav in English. Scripts with no marketing locale (Odia, Gujarati, Tamil, …) and Latin-script languages stay on English rather than guessing from a diaspora country list. The page's own text — H1, extract, story titles, story pages — is always in the content language whatever the chrome resolves to.
Standardized entity strings (also in src/lib/jsonld.ts):
- Product name: unfoldingWord Open Bible Stories
- License sentence: Free to use, adapt, and share under CC BY-SA 4.0.
- Canonical host:
https://openbiblestories.org(www redirects to it).
src/pages/l/[code]/[story]/index.astro renders one page per (language, story) that has full text — 9,062 of them at the last build, across 195 languages. Everything is in the initial HTML: the story title, every illustration paired with the paragraph it belongs to, the Bible reference, previous/next links and a link back to the hub. Where a release publishes per-story files, an <audio> and/or <video> element carries that story's recording, with a download link for it, and a line saying the text below is what the recording says, in the page's language — the per-story alternative to "all files zipped (7.2 GB)". Illustrations carry non-empty, language-appropriate alt ({story} — illustration {n}), because they carry the story and were never decorative. <html lang>/dir, data-script and the font pack follow the content language, exactly as on the hub; labels come from src/i18n/{lang}/story.json. Self-referencing canonical; hreflang alternates only for the site's own languages (see SEO plumbing). JSON-LD is one CreativeWork with the full text, isPartOf the hub's work, the first illustration as image, and AudioObject/VideoObject only where a real file exists — each carrying transcript (the page's own text, which is what lets a search engine index what the recording says); the VideoObject also needs thumbnailUrl and uploadDate, so it is emitted only where the illustration and the release date exist too.
Which stories have a page is storyNums on the catalog record intersected with the story text this build holds: the metadata says which stories a language has, src/data/stories/{code}.json says which of them the build can render. The hub's links, the story routes, prev/next and sitemap-stories.xml are all built from that same intersection, so a partial fetch cannot leave a link pointing at a page that was skipped; npm run check:routes proves it after every build, and also that every link into /l/ resolves.
Story text is not in the committed snapshot. scripts/fetch-catalog.mjs writes one file per language into src/data/stories/ (generated, gitignored, ~66MB); src/data/stories.ts loads them lazily so a story page pulls in only its own language. The build already downloads every story file to read its title, so keeping the body costs no extra requests — but it does mean a story file that is missing (a fresh clone, a cleaned checkout) forces that language to be re-fetched rather than reused from the snapshot cache.
src/lib/jsonld.tsbuilds one JSON-LD@graphper page (emitted byBase.astro):Organization(unfoldingWord) +WebSitewith aSearchActionto/discover/?q=on every page;CreativeWorkfor the work on the homepage and Discover; anItemListof translations pointing at the hubs on the English Discover page only;hubNodes()for each/l/{code}/page;storyNodes()(with media objects) for each story;faqPageNode()on the 16 contact pages, which carry the questions, where the visible heading/answer text and the structured data are the same words. Media objects (AudioObject/VideoObject) are only emitted where a real file exists — never as empty placeholders.sameAslists only URLs whose identity is confirmed: a wrongsameAsmerges two entities in a knowledge graph. Candidates that could not be verified are parked indocs/canonical-links.mdrather than shipped.- hreflang:
Base.astroemits the full reciprocal 16-locale set plusx-default(→ English) on every localized page. Legal pages, the 404 and the language hubs have no alternates. - Sitemaps:
src/pages/sitemap-index.xml.ts→sitemap-pages.xml(marketing pages with the same hreflang alternates as the HTML; no 404),sitemap-languages.xml(one hub per published language,lastmodfrom the latest release) andsitemap-stories.xml(one entry per story page). Generated bysrc/lib/sitemap.ts;robots.txtpoints at the index. - Story-level hreflang is narrowed, not absent. A full cluster is impossible here: story N exists in ~214 languages, so linking every equivalent of every story is roughly 2.3M
<xhtml:link>elements — gigabytes, past the 50MB per-sitemap limit.sitemap-stories.xmltherefore clusters story N across the languages the site itself is published in (siteLocaleOf()— the 17 catalog codes that ARE one of the 16 marketing locales), and only where the story exists in two or more of them: at most 16 links on ~750 of the ~9,000 story URLs, self-referencing, reciprocal,x-defaulton the English story. Every other story URL gets none, which is correct rather than incomplete — a Hausa story has no reciprocal partner, and a one-sided cluster is worse than none. /llms.txtand/content/{code}.md(src/lib/markdown.ts) are a convenience layer for fetchers that prefer markdown — Google has said no special AI file is required, and the HTML remains canonical. The index lists every hub and, for each language whose full text this build holds, a mirror of all its stories;check:routesre-reads the served file and fails if any URL in it was not built, or if a mirror exists that it does not list. One file per language, not per story:/content/{code}/{story}.mdwould be another ~9,000 files on top of ~9,300 pages, and Cloudflare Pages refuses a deployment over 20,000 files./changelog/and/changelog.xml(src/lib/changelog.ts) are the public changelog #19 asked for: newly published languages dated by the earliest release in their repos' Door43 release history (firstReleasedper entry, read byfetch-catalog.mjsfrom the same releases request that backfills assets, and cached in the snapshot by repo — not by release — so a new version keeps the date already known), then translations with a new release in the last 12 months. The language-level date (firstPublished,firstPublishedDate()in the fetch script) is established only when every publishing team's history was read: with one team undated, the earliest known date may belong to a later team, and the page would announce as new a language the catalog shows was published years earlier. A language without an established date is named as undated rather than dated by its latest release;check:routesasserts that against the served page. English-only at the root like the legal pages — the lines are autonyms, codes and dates — and linked from every footer as "What's new". Seedocs/search-visibility.md→ Freshness.og:localeandog:locale:alternateare emitted per locale.
Three things this repository cannot finish on its own, each with the generated part generated and the human part written down:
docs/search-visibility.md— Search Console / Bing setup, what to watch, the quarterly generative-answer protocol, and where the tracker lives.npm run baselinewrites the query and prompt sheet from the catalog (--all,--out=FILE, or named codes).docs/canonical-links.md— the off-repo checklist that makes openbiblestories.org the cited URL rather than Door43 or a mirror, plus thesameAscandidates waiting on verification.docs/native-review.md— the ask for a speaker of sw/es/hi/ar, andnpm run review -- --priorityto produce their sheet.
The canonical host is the bare domain. https://openbiblestories.org is
the site in astro.config.mjs, the SITE_URL behind every canonical,
hreflang, sitemap and JSON-LD URL, and the target of the www redirect —
there is no www in anything the site emits. Both hostnames are attached to
the Pages project, so the host rule fires and www 301s to the apex rather
than serving a second copy.
One acceptance criterion elsewhere is a production check rather than code, and cannot be proven from a preview deployment (#15):
curl -sI https://www.openbiblestories.org/library # expect 301 -> apex /discover/
curl -sI https://openbiblestories.org/ # expect public, max-age=300, s-maxage=86400
curl -s https://openbiblestories.org/robots.txt # must contain no Disallow and no ai-train=no
public/robots.txt allows every crawler: User-agent: * / Allow: /, plus a named Allow group for each live-retrieval agent (OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User, DuckAssistBot) and each bulk training crawler (GPTBot, Google-Extended, ClaudeBot, anthropic-ai, CCBot, Bytespider, Applebot-Extended, Amazonbot, meta-externalagent, Meta-ExternalFetcher), so a broader per-agent rule cannot override them. Training is allowed on purpose — the decision recorded in #15 (2026-09-11) is "allow all, including training": the content is CC BY-SA 4.0 and the site exists to be found, quoted and learned from in the language of the query.
The file in this repo is not the whole story. Cloudflare's managed robots.txt for the zone can prepend a Content-Signal line and Disallow: / groups for AI crawlers ahead of this file, and that is what production served before this decision (Content-Signal: search=yes, ai-train=no, use=reference). For the served file to match the policy, the Cloudflare dashboard must have the managed robots.txt / "block AI bots" features turned off for this zone (Security → Bots, and Content Signals). Verify after changing it:
curl -fsS https://openbiblestories.org/robots.txt > /tmp/robots.txt \
&& ! grep -Ei '^[[:space:]]*(Disallow|Content-Signal)[[:space:]]*:' /tmp/robots.txt \
&& echo "OK: served robots.txt has no Disallow or Content-Signal directive"
It prints OK only when the download succeeded and no directive matched. It looks at directives, not comments, because the policy note in the file itself mentions Disallow: / and ai-train=no and a naive grep -i would flag them.
src/layouts/Base.astro— the shared page shell:<head>(including canonical/Open Graph/Twitter meta), skip link, header/nav, footer, and thenav.jsscript tag. Nav highlighting comes from each page'sactiveprop. There is exactly one nav, in one order, on every page.docs/— the written half of the work that is not code:search-visibility.md,canonical-links.md,native-review.md.src/pages/— one.astrofile per route (src/pages/features/index.astro→/features/). Each page passes its title/description to the layout and supplies only its<main>content, plus any per-page script tags via the namedscriptsslot.src/pages/l/[code]/index.astrogenerates the language hubs,src/pages/l/[code]/[story]/index.astrothe story pages,src/pages/sitemap-*.xml.tsthe sitemaps, andsrc/pages/llms.txt.ts/src/pages/content/[code].md.tsthe markdown layer.public/— copied to the site root verbatim at build time:assets/css/styles.css— shared stylesheet (includes the@font-facerules for the self-hosted fonts).assets/fonts/— self-hosted variable woff2 files for Montserrat and Nunito Sans (latin + latin-ext), replacing the old render-blocking fonts.googleapis.com request.assets/js/— small vanilla-JS behaviors (nav, tabs, discover filtering, the hub reader), loaded as plain script tags (is:inline), not bundled.discover.jsonly filters the prerendered rows and never touches the network;reader.js(mounted byhub-reader.js) is the one script that still calls Door43, and only on a hub, only once someone opens it.assets/img/— images and decorative SVGs._redirects/_headers— Cloudflare routing/caching rules (see below);_headersalso serves/content/*astext/markdown; charset=utf-8.robots.txt— crawl policy (see "Crawlers and AI policy").
Deployed to Cloudflare Pages — project obs-website (obs-web-cgw.pages.dev), production branch main, via the Git integration.
Build settings: build command npm run build, output directory dist (also declared in wrangler.jsonc as pages_build_output_dir).
Current _redirects / _headers rules:
https://www.openbiblestories.org/*301s to the apex. Host-based rules only fire for hostnames attached to the Pages project, and both are attached — the bare domain is canonical andwwwexists only to redirect to it./library,/library/*and/create/library/*redirect (301) to/discover/— the old Library browser was retired in favor of Discover, which now covers search, format filters, and the inline reader. Legacy#code--teamfragments survive the redirect, anddiscover.jsforwards them to that language's hub (fragments never reach the server, so this has to happen client-side)./features/*redirects (301) to/why-obs/(renamed to match its nav label), and/resources/*redirects (301) to/translate/#resources(the standalone page was folded into the Translate page's Resources tab).- HTML (
/*):public, max-age=300, s-maxage=86400, stale-while-revalidate=86400— browsers revalidate after five minutes, the edge holds pages until the next deploy purges them. If public pages still come backno-store/cf-cache-status: DYNAMIC, a zone-level Cache Rule or Browser Cache TTL in the Cloudflare dashboard is overriding this file and must be changed there. /assets/img/*is cached for 1 year (immutable) since filenames don't change./assets/js/*and/assets/css/*useno-cache(always revalidate in the browser) pluss-maxageat the edge, since these are served in place under the same filenames — no content hashing./assets/fonts/*is cached for 1 year (immutable) like images.