Contexto: arranque de yt-extract (analítica de videos, open/freemium)
DATASET BASE (Kaggle):- Primario: rsrishav/youtube-trending-video-dataset ("Trending YouTube Video Statistics")- Alterno: datasnaek/youtube-new (mismo schema, cobertura más larga, MX incluido)ARCHIVOS: {US,GB,CA,DE,FR,IN,KR,JP,RU,MX}videos.csv + category_id.jsonCOLUMNAS (por snapshot diario "trending_date"): video_id, trending_date, title, channel_title, category_id, publish_time, tags, views, likes, dislikes, comment_count, thumbnail_link, comments_disabled, ratings_disabled, video_error_or_removed, descriptionMAPEO A LAS 3 METAS: 1. Éxito de temas -> category_id + tags + title => TF-IDF/embeddings, cluster por tópico; éxito = views + supervivencia en trending 2. Momentos clave -> NO está en trending (necesita audienceRetention de Analytics API o heatmap "most replayed"). Ser honesto: este dataset da nivel-tema, no nivel-segundo. 3. Repetibilidad -> panel longitudinal trending_date => survival analysis (hazard de permanencia) + recurrencia de (canal, tópico)PRIMER EXPERIMENTO FALSABLE (corre en pandas, local, sin API keys): - Hazard de permanencia en trending, estratificado por category_id - Medida de recurrencia de tópico/canal a través de semanas - Hipótesis a testear después: colapso de varianza estilística 2015->2026DESCARGA: pip install kaggle + export KAGGLE_USERNAME=... KAGGLE_KEY=... kaggle datasets download rsrishav/youtube-trending-video-dataset -p data/ --unzipCAVEAT: cobertura de años varía por mirror; el original ~2017-2018, mirrors"updated daily" más recientes pero parciales. Verificar tras descargar.