datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spider-realistic
Dataset Card for Spider-Releastic
This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.realistic-bpe5-science-math-10bcc12m-4mp-realistic
Overview
This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset.
This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning
specifically matches either "A man" or "A woman".
Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our
CC12M-cleaned subset.
If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.cc12m-2mp-realistic
Overview
A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets.
This dataset is created for if you basically need more images and dont mind a little less quality.
Quality
I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting.
This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.cc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.realistic-bpe5-fineweb-10bPretokenized dataset of 10B FineWeb-Edu tokens (sample-10BT) along with 5 domain-specific BPE tokenizers.
realistic-bpe5-wiki-qa-10brealistic-bpe5-union-longest-10bRealistic-world-datasetJust a beta - still in creation
realistic-bpe5-fineweb-20bcc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/QinboZhang/cc12m-1mp_plus-realistic.
