datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoPlan_test
EgoPlan-Test
Videos: We trim videos in EgoPlan test set by the start_frame and end_frame
Text: we add "video_path" key to the original json file.
manga-colorization-mastercolor-filtered-c4
CoLoR-Filtered C4
This repo contains two datasets: color-filtered-c4-books and color-filtered-c4-down associated with the CoLoR-Filter paper.
Each dataset is a 64x filtered version of the C4 dataset from Raffel et al., 2019 that has been selected using the CoLoR-Filter algorithm for data selection.
Each dataset has about 2.7b tokens when using the allenai/eleuther-ai-gpt-neox-20b-pii-special tokenizer.
color-filtered-c4-books was selected to target books based on a small (25m token)… See the full description on the dataset page: https://huggingface.co/datasets/davidbrandfonbrener/color-filtered-c4.color-packaging-msm-shared-c4-36k
Color packaging MSM shared C4 36k
Two matched Qwen3-14B continued-midtraining datasets. Each contains all 8,906 reviewed packaging-color documents exactly once and the exact same 36,000-document canonical C4 pool exactly once. Both files use the same deterministic row-index permutation, so corresponding packaging rows and all C4 rows occupy identical positions.
No synthetic prefix is added and every row declares an empty mask_prefix; all document and EOS tokens remain… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/color-packaging-msm-shared-c4-36k.
