CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01niryuu /sni-each-convertedConverted Super-NaturalInstructions to jsonl https://github.com/allenai/natural-instructions text10K<n<100K1 likes1.8k downloads3y agoHugging Face02urialon /converted_narrative_qa Dataset Card for "converted_narrative_qa" More Information needed text10K<n<100K0 likes866 downloads3y agoHugging Face03VR-VLA /VR-egodex-annotation-converted-v6.0 VR-egodex-annotation-converted-v6.0 EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with per-clip narration included as language sidecars. 314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks 100% narration coverage (1 sidecar per clip) 71 GB annotations + 2.3 GB narratives Videos are NOT included. This release contains annotations and narration only. Source video lives in griffinlabs/EgoDex-LeRobot-v3.0; orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.tabularrobotics10M<n<100M0 likes686 downloads14d agoHugging Face04Tool-learning-it /Toucan-converted-datasettext1M<n<10M0 likes577 downloads9mo agoHugging Face05xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes360 downloads2y agoHugging Face06ai2-adapt-dev /flan_v2_convertedThis is a converted version of the Flan dataset into Tulu SFT training format. The conversion script can be found in our open-instruct repo. The conversion took the following parameters: apply_keyword_filters: True apply_empty_message_filters: True push_to_hub: True hf_entity: ai2-adapt-dev converted_dataset_name: flan_v2_converted local_save_dir: ./data/sft/flan The original FLAN dataset needs extensive efforts to be regenerated, so we are using a reproduced version by the OpenOrca… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/flan_v2_converted.text10K<n<100K3 likes345 downloads2y agoHugging Face07ChavyvAkvar /II-Medical-Reasoning-SFT-Convertedtext1M<n<10M0 likes303 downloads1y agoHugging Face08DuarteMRAlves /persona_instruct_2shot_convertedtext100K<n<1M0 likes188 downloads10mo agoHugging Face09ChavyvAkvar /aya_collection_language_split-standard_malay-Convertedtext1M<n<10M0 likes181 downloads1y agoHugging Face10ChavyvAkvar /OpenCodeInstruct-Convertedtext1M<n<10M0 likes170 downloads1y agoHugging Face11ChavyvAkvar /MegaScience-dataset-Convertedtext1M<n<10M0 likes161 downloads1y agoHugging Face12ChavyvAkvar /OpenCodeReasoning-split_1-Convertedtext100K<n<1M0 likes158 downloads1y agoHugging Face13ChavyvAkvar /xP3x-ind_Latn-Convertedtext1M<n<10M0 likes150 downloads1y agoHugging Face14Jezzarax /pubhealth-converted PUBHEALTH PUBHEALTH is a public-health fact-checking dataset introduced in: Neema Kotonya and Francesca Toni. 2020. Explainable Automated Fact-Checking for Public Health Claims. This repository is a scriptless Parquet conversion of the original PUBHEALTH TSV files. It is intended to load with the Hugging Face datasets library without requiring deprecated remote dataset scripts. The deprecated script-based dataset was available at:… See the full description on the dataset page: https://huggingface.co/datasets/Jezzarax/pubhealth-converted.texttext-classification10K<n<100K0 likes145 downloads4mo agoHugging Face15mlfoundations-dev /riddle_sense_convertedtext10K<n<100K0 likes139 downloads2y agoHugging Face16woodygan /converted_cvsstext100K<n<1M0 likes131 downloads8mo agoHugging Face17ChavyvAkvar /Nemotron-Post-Training-Dataset-v2-chat-Convertedtext100K<n<1M0 likes130 downloads1y agoHugging Face18k1000dai /converted_mixed_pickandplace_datasetimage100K<n<1M0 likes128 downloads9mo agoHugging Face19DuarteMRAlves /persona_math_2shot_convertedtext10K<n<100K0 likes127 downloads10mo agoHugging Face20ChavyvAkvar /SYNTHETIC-2-SFT-verified-Convertedtext100K<n<1M0 likes119 downloads1y agoHugging Face21takara-ai /FRED-CONVERTEDimage100K<n<1M1 likes110 downloads11mo agoHugging Face22ChavyvAkvar /aya_collection_language_split-central_khmer-Convertedtext1M<n<10M0 likes105 downloads1y agoHugging Face23ChavyvAkvar /indonesian-reasoning-Convertedtext1K<n<10K3 likes102 downloads1y agoHugging Face24ChavyvAkvar /aya_collection_language_split-thai-Convertedtext1M<n<10M0 likes101 downloads1y agoHugging Face25ChavyvAkvar /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedtext10K<n<100K1 likes98 downloads1y agoHugging Face26DuarteMRAlves /persona_instruction_following_convertedtext1K<n<10K0 likes98 downloads10mo agoHugging Face27ChavyvAkvar /DeepWriting-20K-Convertedtext10K<n<100K0 likes96 downloads1y agoHugging Face28dddraxxx /refchartqa_converted_hfimage10K<n<100K0 likes83 downloads11mo agoHugging Face29cfahlgren1 /Capybara-Converted This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.textquestion-answering10K<n<100K1 likes66 downloads3y agoHugging Face30ai2-adapt-dev /no_robots_convertedThis is a converted version of the no_robots dataset into Tulu SFT training format. The conversion script can be found in our open-instruct repo. The conversion took the following parameters: apply_keyword_filters: False apply_empty_message_filters: False push_to_hub: True hf_entity: ai2-adapt-dev converted_dataset_name: no_robots_converted local_save_dir: ./data/sft/no_robots Please refer to the original dataset for more information about this dataset and the license. text10K<n<100K0 likes66 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.