Enrichment pipeline: scraped bodies + wikilinks + sources
Date: 2026-08-28 Project: astro-starter-template (second-brain)
Decision
Every node body grows from a one-liner into a short article (~200 words):
user’s original line kept as lead + a sourced paragraph + [[wikilinks]] to
other nodes, so the graph reads like a blog/Wikipedia and stops feeling
isolated. Script: scripts/enrich.mjs (npm run enrich), drafts to _enrich/
before adoption (same convention as _stubs/).
Sources by kind
- concept/algorithm/model: Wikipedia (search API with kind-aware hint for disambiguation, then /summary extract)
- paper/experiment: arXiv title search → abstract + abs URL (over https; the old http://export.arxiv.org returns empty)
- tool: Wikipedia + curated DOCS map (numpy, pandas, sklearn, pytorch, tf, keras, docker, k8s, spark, kafka, airflow, dbt, elixir, genserver, ets, …)
- person/dataset: Wikipedia
- NOT scraped: DataCamp (JS/paywalled app — no clean endpoint), Minitab/MATLAB (no matching nodes; MATLAB docs mapped if a node appears).
Result (2026-08-28)
174/174 enriched · 131 with sources · 105 wikilinks · avg body 63 words.
Wikilinks resolve in the app: [[Title]] → #/node/<slug> (renderBody in
site/app/main.js), sources render as a “Fuentes” section (target=_blank).
Gotchas (write them down, they cost 2 crashed runs)
- Wikipedia throttles bursts: sleep ~700ms between nodes + retry with backoff (3s·n) on 429/5xx. A full pass takes ~25 min wall clock.
- Cache
data/enrich-cache/<slug>-{wiki,arxiv}.json— MUST returnnullfor{"empty":true}entries; returning the object madesources.pushcrash onurlundefined (truthy empty object). - NEVER round-trip frontmatter through gray-matter.stringify — it reflows
dates (
2026-08-28→ ISO) and YAML blocks. Preserve the raw frontmatter text verbatim; injectsources:before the closing---only. lib/graph.mjsonly forwards known fields — addedsourcesexplicitly.- es.json holds stale one-line translations; new bodies fall back to EN.
Re-run
npm run translatewhen a DeepL key is available.