datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
week1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.LGUplus_weekly_persona
LGUplus Weekly Viewing Persona
일주일치 일일 페르소나 → 주간 장기 시청성향 카드. 3주치(기존 + 신규 2주).
train = 기존주 + 신규 첫주 · test = 신규 다음주 · week 필드로 구분
입력 daily_personas(요일별) + main_genre/sub_genre(주별 시청 장르 카운트) → 카드 7필드(viewing_tendency/preferred_programs/frequent_channels/preferred_genres/weekday·weekend_viewing_times/key_keywords)
