CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Scale-or-Reason /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.textquestion-answering1M<n<10M6 likes1.2k downloads3mo agoHugging Face02Sidsidney /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.textquestion-answering1M<n<10M4 likes797 downloads9mo agoHugging Face03Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes636 downloads3mo agoHugging Face04yaya-sy /bimodal-iftAn instruction dataset for speech->text and text->speech. This speech data is tokenized using the SpeechTokenize approach: https://arxiv.org/abs/2308.16692. You can do standard finetuning on this dataset using any LLM! text1M<n<10M0 likes422 downloads2y agoHugging Face05iftekher /my_data_handsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "fr3", "total_episodes": 8, "total_frames": 2730, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:8" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_hands.imagerobotics1K<n<10K0 likes177 downloads10d agoHugging Face06iftekher /my_data_newThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "fr3", "total_episodes": 5, "total_frames": 1607, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_new.tabularrobotics1K<n<10K0 likes158 downloads13d agoHugging Face07iftekher /my_data_handThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "fr3", "total_episodes": 5, "total_frames": 1227, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_hand.tabularrobotics1K<n<10K0 likes150 downloads13d agoHugging Face08Slim205 /copa_ifttabularn<1K0 likes130 downloads2y agoHugging Face09Slim205 /gsm8k_ift_translated_nllb0text1K<n<10K0 likes122 downloads2y agoHugging Face10Ibisbill /dnd-dataset-improved-ift-qatext1M<n<10M0 likes94 downloads1y agoHugging Face11jeevajoji /IFTtextquestion-answering100K<n<1M0 likes80 downloads2y agoHugging Face12Lichang-Chen /800k_ift Dataset Card for "800k_ift" More Information needed text100K<n<1M0 likes79 downloads2y agoHugging Face13ift /handwriting_forms Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ift/handwriting_forms.imagefeature-extraction1K<n<10K14 likes77 downloads3y agoHugging Face14mothnaZl /long_sr_ift_duplicate_Qwen2.5-7B-Instructtext100K<n<1M0 likes74 downloads1y agoHugging Face15Slim205 /copa_ift_translated_nllb0tabularn<1K0 likes59 downloads2y agoHugging Face16Slim205 /copa_ift_v03tabularn<1K0 likes58 downloads2y agoHugging Face17Domegal13-ifts12 /online-retail-IItabular1M<n<10M0 likes58 downloads22d agoHugging Face18Lichang-Chen /837k_ifttext100K<n<1M0 likes45 downloads2y agoHugging Face19Slim205 /piqa_ifttext10K<n<100K0 likes43 downloads2y agoHugging Face20Slim205 /copa_ift_v02_filteredtabularn<1K0 likes32 downloads2y agoHugging Face21Slim205 /gsm8k_ift_v02_translatedtextn<1K0 likes32 downloads2y agoHugging Face22sam-mosaic /ift_hhrlhf_flan Dataset Card for "ift_hhrlhf_flan" My favorite subsets of FLAN with single-turn data filtered from HH RLHF flan_cats_i_like = { "arc_challenge_10templates", "arc_easy_10templates", "cola_10templates", "copa_10templates", "coqa_10templates", "cosmos_qa_10templates", "fix_punct_10templates", "math_dataset_10templates", "natural_questions_10templates", "openbookqa_10templates", "squad_v2_10templates", "trivia_qa_10templates", } text100K<n<1M0 likes31 downloads3y agoHugging Face23haturusinghe /sold-dataset-for-mistral7b-ifttext10K<n<100K0 likes30 downloads2y agoHugging Face24Slim205 /copa_ift_v02tabularn<1K0 likes30 downloads2y agoHugging Face25Slim205 /copa_ift_v03_translated_nllb0tabularn<1K0 likes30 downloads2y agoHugging Face26RLHFlow /self_rewarding_ift_example_raw_data1text10K<n<100K0 likes30 downloads2y agoHugging Face27haturusinghe /sold-dataset-for-llama3-ifttext10K<n<100K0 likes28 downloads2y agoHugging Face28kguo2 /ift-scaffold-datasettext10K<n<100K0 likes28 downloads1y agoHugging Face29Slim205 /arc_easy_ift_v3text1K<n<10K0 likes27 downloads2y agoHugging Face30Slim205 /copa_ift_v02_filtered_translatedtabularn<1K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.