CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes180 downloads7mo agoHugging Face02Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes145 downloads1y agoHugging Face03Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face04shuyuej /prompt_consistency_training_full_data 🚀 Load Dataset from datasets import load_dataset dataset = load_dataset("shuyuej/prompt_consistency_training_full_data") dataset = dataset["train"] print(dataset) text1M<n<10M1 likes92 downloads3y agoHugging Face05Team-Kitsune /KUJIRA_DATASETS_FULL Overview This dataset is an English-language dataset created specifically for Reasoning models within the KUJIRA_v2 series. The primary objective of the dataset is to enhance model performance while ensuring safety and removing censorship typical of Chinese-origin models. In creating this dataset, references were made to the dataset recipes of r1-1776 and MAI-DS-R1. Additionally, this dataset leverages existing Q&A data originally designed for intuitive models by employing the… See the full description on the dataset page: https://huggingface.co/datasets/Team-Kitsune/KUJIRA_DATASETS_FULL.texttext-generation100K<n<1M1 likes57 downloads1y agoHugging Face06FabianOvalle /Dataset_Robot_IA_full_v8textn<1K0 likes44 downloads26d agoHugging Face07datatab /alpaca-cleaned-serbian-full Serbian Alpaca Cleaned Dataset Original Repository: https://github.com/gururise/AlpacaDataCleaned Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned Dataset Description This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.texttext-generation10K<n<100K1 likes32 downloads3y agoHugging Face08yileitu /Mdist_Chem_LST_150k_ft_data_from_full_360ktext100K<n<1M0 likes30 downloads3mo agoHugging Face09meandyou200175 /full_data_15_negtext10K<n<100K0 likes24 downloads2y agoHugging Face10yileitu /Mdist_Chem_LST_Qwen3_14B_full_ft_datatext100K<n<1M0 likes24 downloads3mo agoHugging Face11GIZ /vulnerability_training_data_fulltabulartext-classificationn<1K0 likes21 downloads3y agoHugging Face12yileitu /Mdist_Chem_LST_250k_ft_data_from_full_360ktext100K<n<1M0 likes21 downloads3mo agoHugging Face13yileitu /Mdist_Chem_FFT_Qwen3_14B_full_ft_datatext100K<n<1M0 likes21 downloads3mo agoHugging Face14Guilherme34 /Reasoner-dataset-FULL-rolestext10K<n<100K3 likes17 downloads2y agoHugging Face15yileitu /Mdist_Chem_LST_100k_ft_data_from_full_360ktext100K<n<1M0 likes16 downloads3mo agoHugging Face16yileitu /Mdist_Chem_LST_300k_ft_data_from_full_360ktext100K<n<1M0 likes16 downloads3mo agoHugging Face17Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({ train: Dataset({ features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'], num_rows: 12166802 }) }) image1M<n<10M0 likes12 downloads3y agoHugging Face18ai-aerospace /ams_data_full_2000-2020Aerospace Mechanism Symposia PDF documents parsed by page. All symposia documents from the year 2000-2022 are included. No splitting was used. Original documents here: https://github.com/dan-s-mueller/aerospace_chatbot/tree/main/data/AMS textquestion-answering1K<n<10K0 likes10 downloads2y agoHugging Face19imdatta0 /sanskrit_full_datasettext1K<n<10K0 likes10 downloads7mo agoHugging Face20Tohirju /kyrgyz-asr-full-datagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. text1K<n<10K0 likes9 downloads1mo agoHugging Face21open-llm-leaderboard /godlikehhd__alpaca_data_full_2-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_2 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_2-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face22open-llm-leaderboard /godlikehhd__alpaca_data_full_3B-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_3B Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_3B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_3B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face23nakotsuko13 /distilled_dataset_full_cot_v1text1K<n<10K0 likes8 downloads7mo agoHugging Face24Tohirju /uzbek-asr-full-datagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. text100K<n<1M0 likes8 downloads1mo agoHugging Face25Agnuxo /Francisco-Angulo-de-Lafuente-Full-Dataset Francisco Angulo de Lafuente Full Training Dataset Este dataset contiene la recopilación completa de la obra, investigación, código y biografía de Francisco Angulo de Lafuente. Contenido Biografía Detallada: Información sobre su trayectoria en biotecnología, ingeniería informática y literatura. Proyectos Principales: Documentación técnica de P2PCLAW, EUHNN, CAJAL y otros. Papers Científicos: Resúmenes y textos completos de sus investigaciones en computación neuromórfica… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/Francisco-Angulo-de-Lafuente-Full-Dataset.texttext-generationn<1K0 likes7 downloads5mo agoHugging Face26open-llm-leaderboard /godlikehhd__qwen_full_data_alpaca-detailsgated Dataset Card for Evaluation run of godlikehhd/qwen_full_data_alpaca Dataset automatically created during the evaluation run of model godlikehhd/qwen_full_data_alpaca The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_full_data_alpaca-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face27halsarkhi /scta_full_dataset.jsonltext1K<n<10K0 likes6 downloads4mo agoHugging Face28Tohirju /tajik-asr-full-datagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. text100K<n<1M0 likes6 downloads1mo agoHugging Face29FabianOvalle /Dataset_Robot_IA_full_v7textn<1K0 likes5 downloads10mo agoHugging Face30Tohirju /kazakh-asr-full-datagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. text100K<n<1M0 likes5 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.