CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.3k downloads4y agoHugging Face02jacobmorrison /rejection_sampling_22689text100K<n<1M0 likes1.8k downloads2y agoHugging Face03olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes1.6k downloads4y agoHugging Face04olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes1.4k downloads4y agoHugging Face05olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.3k downloads4y agoHugging Face06olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes1.2k downloads4y agoHugging Face07olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes1.2k downloads4y agoHugging Face08olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes795 downloads4y agoHugging Face09jacobmorrison /rejection_sampling_943 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_943', 'hf_repo_id_scores': 'scores_943', 'input_filename': '/output/shards/943/34.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_943.text100K<n<1M0 likes716 downloads2y agoHugging Face10lumicero /CHB-MIT-SpectroJEPA-Sampling0 likes698 downloads2mo agoHugging Face11CHATS-Lab /Verbalized-Sampling-Joke-Generation Verbalized-Sampling-Joke-Generation This dataset demonstrates how Verbalized Sampling (VS) increases diversity in creative generation tasks, specifically joke generation, while maintaining humor quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Joke Generation dataset contains diverse jokes from state-of-the-art LLMs in response to prompts requesting jokes about specific topics. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Joke-Generation.text100K<n<1M0 likes582 downloads11mo agoHugging Face12jacobmorrison /rejection_sampling_22710 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_22710', 'hf_repo_id_scores': 'scores_22710', 'input_filename': '/output/shards/22710/29.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_22710.text100K<n<1M0 likes522 downloads2y agoHugging Face13LiangYan3612 /pmdm_sampling_datatextn<1K0 likes508 downloads1y agoHugging Face14Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only" More Information needed text1M<n<10M0 likes485 downloads4y agoHugging Face15SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes439 downloads9mo agoHugging Face16thedarkknight7 /SAE_monosemanticity_features_4x_0.01_samplingtabular100M<n<1B0 likes372 downloads6mo agoHugging Face17Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes327 downloads4y agoHugging Face18CHATS-Lab /Verbalized-Sampling-Open-Ended-QA Verbalized-Sampling-Open-Ended-QA This dataset demonstrates how Verbalized Sampling (VS) increases diversity in open-ended question answering while maintaining response quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Open-Ended QA dataset contains diverse responses from state-of-the-art LLMs to open-ended questions across various domains. This dataset evaluates: Response diversity: Coverage of… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Open-Ended-QA.text100K<n<1M0 likes274 downloads11mo agoHugging Face19jacobmorrison /rejection_sampling_10627 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_10627', 'hf_repo_id_scores': 'scores_10627', 'input_filename': '/output/shards/10627/23.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_10627.text100K<n<1M0 likes261 downloads2y agoHugging Face20Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup" More Information needed text1M<n<10M0 likes258 downloads4y agoHugging Face21Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters" More Information needed tabular10M<n<100M0 likes256 downloads4y agoHugging Face22ISSAI-sampling /text0 likes250 downloads2mo agoHugging Face23jacobmorrison /rejection_sampling_4036 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_4036', 'hf_repo_id_scores': 'scores_4036', 'input_filename': '/output/shards/4036/27.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_4036.text100K<n<1M0 likes234 downloads2y agoHugging Face24mlfoundations-cua-dev /easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RL-a-kimage10K<n<100K0 likes223 downloads1y agoHugging Face25mlfoundations-cua-dev /easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RLimage10K<n<100K0 likes214 downloads1y agoHugging Face26jacobmorrison /rejection_sampling_9564 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_9564', 'hf_repo_id_scores': 'scores_9564', 'input_filename': '/output/shards/9564/3.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions': 8… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_9564.text100K<n<1M0 likes191 downloads2y agoHugging Face27bertin-project /mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.text-generationn<1K13 likes190 downloads2y agoHugging Face28jacobmorrison /rejection_sampling_23174 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_23174', 'hf_repo_id_scores': 'scores_23174', 'input_filename': '/output/shards/23174/28.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['NCSOFT/Llama-3-OffsetBias-RM-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_23174.text100K<n<1M0 likes186 downloads2y agoHugging Face29CHATS-Lab /Verbalized-Sampling-Dialogue-Simulation Verbalized-Sampling-Dialogue-Simulation This dataset demonstrates how Verbalized Sampling (VS) enables more diverse and realistic multi-turn conversational simulations between AI agents. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Dialogue Simulation dataset contains multi-turn conversations between pairs of language models, comparing different approaches to generating diverse social interactions.… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation.text1K<n<10K0 likes176 downloads11mo agoHugging Face30jacobmorrison /rejection_sampling_19367 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_19367', 'hf_repo_id_scores': 'scores_19367', 'input_filename': '/output/shards/19367/14.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_19367.text10K<n<100K0 likes168 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.