WORKER: yt-extract-observatorio
URL: https://yt-extract-observatorio.nef.workers.dev
Cuenta: [ID-EMAIL]'s Account (162ab69a40524314530193e8a3e8f414)
Dir: worker/ (wrangler.jsonc + src/index.js + schema.sql + seed_channels.sql)
D1: yt-extract-observatorio id=0531e406-457c-4309-acc6-12f0d1bdd4bf (WNAM)
Secreto: YT_API_KEY (bind env.YT_API_KEY)
Crons: */15 * * * * -> pullTrending (trending MX -> videos)
0 * * * * -> checkChannels (38 canales -> nuevos + trainability)
Endpoints:
/ -> ok + endpoints
/refresh -> dispara pullTrending + checkChannels (manual)
/latest -> 100 videos recientes
/stats -> conteo por trainability
TABLAS D1:
channels(channel_id PK, nombre, categoria, tipo_medio) -- 38 seeded de surtido/channels.json
videos(id PK, channel_id, channel_title, title, category_id, views,
published_at, trending_region, trainability, first_seen_at, last_seen_at)
VERIFICADO: /refresh -> trending:50 channels_new:20 -> 70 videos en D1,
trainability todo "None" (opt-out). Pipeline end-to-end OK.
NOTAS:
- trainability en D1 se guarda RAW (permitted join ","): "None"|"All"|lista.
- Local `wrangler dev` falla con wrangler 4.99.0 (workerd solo soporta hasta
compat date 2026-06-16; el deploy en edge si acepta 2026-08-01).
- worker/.dev.vars (local, gitignored) para pruebas locales.
- VideoTrainability: 1 video/call, no batch.
BUG CLAVE (limite subrequests): plan gratuito = 50 subrequests/invocacion.
38 canales x ~3-13 fetch lo revienta -> "Too many subrequests by single Worker
invocation". SOLUCION: procesar en CHUNKS con cursor en D1 (tabla state):
CHANNELS_PER_RUN=6, MAX_TRAIN=5 (cap trainability por canal), cron */5.
Cada run: 6 canales x (channels.list + playlistItems + videos.list batch + <=5
trainability) <= 48 subrequests. Cursor avanzado modulo total.
- videos.list se BATCHEA (ids join ",") -> 1 call en vez de 1 por video.
- Trainability SIEMPRE 1 video/call (no batch).
ENDPOINT /deck: devuelve {{channels:[con videos+badge], trending, stats}} con CORS.
- frontend: live.html (TweetDeck en vivo, auto-refresh 5min) + dashboard.html
(seccion "En vivo" que fetchea /deck).
- CORS: Access-Control-Allow-Origin: * en json() + OPTIONS 204.
CLOUDFLARE PAGES (publico):
proyecto "yt-extract" -> https://yt-extract.pages.dev
deploy: wrangler pages deploy site --project-name yt-extract
site/: index.html(=live.html) + dashboard.html + observatorio.html + deck.html
Nota: Pages limpia extension .html (dashboard.html -> /dashboard).