CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face02argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes16k downloads3y agoHugging Face03HuggingFaceH4 /stack-exchange-preferences Dataset Card for H4 Stack Exchange Preferences Dataset Dataset Summary This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training. Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers. This data could also be used for instruction fine-tuning and language model training. The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.textquestion-answering10M<n<100M135 likes9.1k downloads4y agoHugging Face04multimodalart /lora-fusing-preferencesimage1K<n<10K12 likes5.2k downloads2y agoHugging Face05data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes5k downloads2y agoHugging Face06wassname /genies_preferences Dataset Card for "genie_dpo" A conversion of the distribution from GENIES to open_pref_eval format. Conversion code texttext-classification100K<n<1M1 likes4.3k downloads2y agoHugging Face07wassname /mmlu_preferencesreformat of MMLU to be in DPO (paired) format examples: {'prompt': 'Which of the following statements about the lanthanide elements is NOT true?', 'chosen': 'The atomic radii of the lanthanide elements increase across the period from La to Lu.', 'rejected': 'All of the lanthanide elements react with aqueous acid to liberate hydrogen.'} college_chemistry {'prompt': 'Beyond the business case for engaging in CSR there are a number of moral arguments relating to: negative _______, the… See the full description on the dataset page: https://huggingface.co/datasets/wassname/mmlu_preferences.texttext-classification10K<n<100K0 likes1.4k downloads2y agoHugging Face08euclaise /WritingPrompts_preferences Dataset Card for "WritingPrompts_preferences" Human preference data from r/WritingPrompts texttext-generation100K<n<1M13 likes1.2k downloads3y agoHugging Face09Rapidata /text-2-video-human-preferences Rapidata Video Generation Preference Dataset This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set: Sora Hunyouan Pika 2.0 Runway ML Alpha Luma Ray 2 Explore our latest model rankings on our website. If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.imagetext-to-video1K<n<10K21 likes1.1k downloads2y agoHugging Face10Rapidata /text-2-video-human-preferences-wan2.1 Rapidata Video Generation Alibaba Wan2.1 Human Preference If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.imagevideo-classificationn<1K20 likes1.1k downloads2y agoHugging Face11Rapidata /human-coherence-preferences-images Rapidata Image Generation Coherence Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.imagetext-to-image10K<n<100K14 likes1k downloads2y agoHugging Face12rngusry /UltraFeedback-truthfulness-preferences Dataset Card for "UltraFeedback-truthfulness-preferences" More Information needed tabular100K<n<1M1 likes1k downloads2y agoHugging Face13Rapidata /human-style-preferences-images Rapidata Image Generation Preference Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human preference datasets for text-to-image models, this release contains over 1,200,000 human preference… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-style-preferences-images.imagetext-to-image10K<n<100K29 likes808 downloads2y agoHugging Face14Rapidata /Runway_Frames_t2i_human_preferences Rapidata Frames Preference This T2I dataset contains roughly 400k human responses from over 82k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Frames across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Runway_Frames_t2i_human_preferences.imagetext-to-image10K<n<100K14 likes789 downloads2y agoHugging Face15Rapidata /human-alignment-preferences-images Rapidata Image Generation Alignment Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.imagetext-to-image10K<n<100K17 likes781 downloads2y agoHugging Face16bigcode /stack-exchange-preferences-20230914-clean-anonymization Dataset Card for "stack-exchange-preferences-20230914-clean-anonymization" More Information needed text10M<n<100M6 likes737 downloads3y agoHugging Face17P1ayer-1 /stack-exchange-preferences-code Dataset Card for "stack-exchange-preferences-code" More Information needed text1M<n<10M3 likes706 downloads3y agoHugging Face18Rapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes683 downloads1y agoHugging Face19Rapidata /xAI_Aurora_t2i_human_preferences Rapidata Aurora Preference This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Aurora across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/xAI_Aurora_t2i_human_preferences.imagetext-to-image10K<n<100K15 likes638 downloads2y agoHugging Face20aflah /LLM-Latent-Preferences0 likes603 downloads11mo agoHugging Face21rngusry /UltraFeedback-honesty-preferences Dataset Card for "UltraFeedback-honesty-preferences" More Information needed tabular100K<n<1M1 likes601 downloads2y agoHugging Face22datapointai /text-to-speech-human-preferences-315kgated Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source evaluation collected 357,651 completed responses. The published benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.audiotext-to-speech100K<n<1M38 likes575 downloads25d agoHugging Face23Rapidata /text-2-video-human-preferences-moonvalley-marey Rapidata Video Generation Marey Pro Human Preference In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.imagevideo-classification1K<n<10K7 likes506 downloads1y agoHugging Face24P1ayer-1 /stack-exchange-preferences-code-v2 Dataset Card for "stack-exchange-preferences-code-v2" More Information needed text1M<n<10M2 likes468 downloads3y agoHugging Face25argilla /Capybara-Preferences Dataset Card for Capybara-Preferences This dataset has been created with distilabel. Dataset Summary This dataset is built on top of LDJnr/Capybara, in order to generate a preference dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.tabulartext-generation10K<n<100K47 likes399 downloads2y agoHugging Face26Rapidata /text-2-video-human-preferences-veo3 Rapidata Video Generation Veo 3 Human Preference In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.imagevideo-classification1K<n<10K20 likes386 downloads1y agoHugging Face27argilla /ultrafeedback-binarized-preferences Ultrafeedback binarized dataset using the mean of preference ratings Introduction This dataset contains the result of curation work performed by Argilla (using Argilla 😃). After visually browsing around some examples using the sort and filter feature of Argilla (sort by highest rating for chosen responses), we noticed a strong mismatch between the overall_score in the original UF dataset (and the Zephyr train_prefs dataset) and the quality of the chosen response. By… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences.tabular10K<n<100K84 likes384 downloads3y agoHugging Face28datapointai /text-2-image-human-preferences-2mgated Text-to-image human preferences: 2M votes across 30 models This dataset contains the complete voting record behind the Datapoint Image Bench leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of 216,116 image pairs. The votes compare 30 text-to-image models in a complete round-robin on 500 prompts, judged by annotators from over 200 countries. Every vote includes the annotator's trust score at the time the vote was cast. Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.imagetext-to-image1M<n<10M21 likes325 downloads1mo agoHugging Face29prithivMLmods /Face-3D-Unified-Preferences Face-3D-Unified-Preferences Face-3D-Unified-Preferences is a high-quality dataset designed for human face depth estimation, 3D face reconstruction, and unified preference learning. The dataset is a mixture of male and female human face portraits, providing a diverse collection of facial appearances, identities, poses, and expressions for training modern computer vision and multimodal AI models. Each sample contains an RGB face image, a dense facial depth map, and a… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Face-3D-Unified-Preferences.3dimage-to-3d1K<n<10K1 likes310 downloads3mo agoHugging Face30data-is-better-together /open-image-preferences-v1-binarized Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1-binarized.image1K<n<10K59 likes309 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.