datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clm-backbone-5lang-sample
clm-backbone-5lang-sample
clm_prod 3B/7B BACKBONE corpus (step 2 DATA PREP). A bounded streaming SAMPLE —
NOT the full 60B-token pretrain set (that is the step 3 H100 fire).
Source: allenai/c4 (mC4 multilingual), configs ko/en/zh/ru/ja
License: ODC-BY (Open Data Commons Attribution)
real_fraction: 1.0 — every line is real downloaded web text. NO synthesis, NO LLM generation.
Preferred set blocked: uonlp/CulturaX is GATED; the token has no data access (403 on stream) → fell back… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/clm-backbone-5lang-sample.backbonecv_backbones_duplicate
Dataset Card for "monetjoe/cv_backbones"
This repository consolidates the collection of backbone networks for pre-trained computer vision models available on the PyTorch official website. It mainly includes various Convolutional Neural Networks (CNNs) and Vision Transformer models pre-trained on the ImageNet1K dataset. The entire collection is divided into two subsets, V1 and V2, encompassing multiple classic and advanced versions of visual models. These pre-trained backbone… See the full description on the dataset page: https://huggingface.co/datasets/EF54321/cv_backbones_duplicate.LOFF_7298LoSet_LossLOFF_LossEvalD_20250919LoraQ_f830e3EvalD_3691
