M.INC.
Datasets
All datasets matching “M.INC.”fineweb-edu-10BT-mincols
fineweb-edu: 10BT sample
This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons.
token counts
GPT-4 tiktoken token count:
token_count
count 9.672101e+06
mean 1.001188e+03
std 1.834986e+03
min 3.800000e+01
25% 3.380000e+02
50% 6.090000e+02
75% 1.054000e+03
max 1.649670e+05
Total count: 9683.59 M tokens
subvideo_move1_originfineweb-edu-10BT-mincols
fineweb-edu: 10BT sample
This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons.
token counts
GPT-4 tiktoken token count:
token_count
count 9.672101e+06
mean 1.001188e+03
std 1.834986e+03
min 3.800000e+01
25% 3.380000e+02
50% 6.090000e+02
75% 1.054000e+03
max 1.649670e+05
Total count: 9683.59 M tokens
subvideo_move1cosmopedia-v2-mincols
cosmopedia-v2: mincols
cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use
minc-2500_split_1
Materials in Context Dataset (MINC-2500)
Dataset Summary
(from the website)
MINC-2500 is a patch classification dataset with 2500 samples per category
(Section 5.4 of the paper). This is a subset of MINC where samples have been
sized to 362 x 362 and each category is sampled evenly. The original resolution
images are not needed as we include the extracted patches in the archive.
