datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204
Dataset Card for OLM May 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
rejection_sampling_22689olm-CC-MAIN-2022-33-sampling-ratio-0.20
Dataset Card for OLM August 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881
Dataset Card for OLM June/July 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547
Dataset Card for OLM November/December 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949
Dataset Card for OLM May 2017 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295
Dataset Card for OLM September/October 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
rejection_sampling_943
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_943',
'hf_repo_id_scores': 'scores_943',
'input_filename': '/output/shards/943/34.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_943.CHB-MIT-SpectroJEPA-SamplingVerbalized-Sampling-Joke-Generation
Verbalized-Sampling-Joke-Generation
This dataset demonstrates how Verbalized Sampling (VS) increases diversity in creative generation tasks, specifically joke generation, while maintaining humor quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Joke Generation dataset contains diverse jokes from state-of-the-art LLMs in response to prompts requesting jokes about specific topics. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Joke-Generation.rejection_sampling_22710
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_22710',
'hf_repo_id_scores': 'scores_22710',
'input_filename': '/output/shards/22710/29.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_22710.pmdm_sampling_dataolm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only"
More Information needed
DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-SamplingSAE_monosemanticity_features_4x_0.01_samplingolm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters"
More Information needed
Verbalized-Sampling-Open-Ended-QA
Verbalized-Sampling-Open-Ended-QA
This dataset demonstrates how Verbalized Sampling (VS) increases diversity in open-ended question answering while maintaining response quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Open-Ended QA dataset contains diverse responses from state-of-the-art LLMs to open-ended questions across various domains. This dataset evaluates:
Response diversity: Coverage of… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Open-Ended-QA.rejection_sampling_10627
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_10627',
'hf_repo_id_scores': 'scores_10627',
'input_filename': '/output/shards/10627/23.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_10627.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup"
More Information needed
olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters"
More Information needed
textrejection_sampling_4036
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_4036',
'hf_repo_id_scores': 'scores_4036',
'input_filename': '/output/shards/4036/27.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['/reward_model'],
'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_4036.easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RL-a-keasyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RLrejection_sampling_9564
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_9564',
'hf_repo_id_scores': 'scores_9564',
'input_filename': '/output/shards/9564/3.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['/reward_model'],
'num_completions': 8… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_9564.mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.rejection_sampling_23174
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_23174',
'hf_repo_id_scores': 'scores_23174',
'input_filename': '/output/shards/23174/28.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['NCSOFT/Llama-3-OffsetBias-RM-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_23174.Verbalized-Sampling-Dialogue-Simulation
Verbalized-Sampling-Dialogue-Simulation
This dataset demonstrates how Verbalized Sampling (VS) enables more diverse and realistic multi-turn conversational simulations between AI agents. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Dialogue Simulation dataset contains multi-turn conversations between pairs of language models, comparing different approaches to generating diverse social interactions.… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation.rejection_sampling_19367
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_19367',
'hf_repo_id_scores': 'scores_19367',
'input_filename': '/output/shards/19367/14.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_19367.
