CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twangodev /tartanaviation-atc-adsb-utterances TartanAviation ATC + ADS-B (Utterances) Speech utterances split from twangodev/tartanaviation-atc-adsb by voice-activity detection (pyannote/segmentation-3.0). Each row is one speech segment (16 kHz mono) with the ADS-B from its parent clip. 531,050 utterances · ~398 h speech · 16 kHz mono · 67% carry ADS-B. From 40,899 of 41,823 clips (silent clips have no utterances). Built with squawk. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb-utterances.audioautomatic-speech-recognition100K<n<1M0 likes1.4k downloads4mo agoHugging Face02twangodev /tartanaviation-atc-adsb TartanAviation ATC + ADS-B Paired ATC audio and ADS-B for Pittsburgh KAGC and KBTP, aligned from CMU TartanAviation. Each row is one ADS-B-triggered audio capture (16 kHz mono) plus the aircraft tracks present during it. 41,823 clips · 16 kHz mono · 67% carry ADS-B. Built with squawk. Usage from datasets import load_dataset ds = load_dataset("twangodev/tartanaviation-atc-adsb", split="train", streaming=True) ex = next(iter(ds)) ex["audio"] # {'array': ...… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb.audioautomatic-speech-recognition10K<n<100K0 likes925 downloads4mo agoHugging Face03Jzuluaga /atcosim_corpus Dataset Card for ATCOSIM corpus Dataset Summary The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus.audioautomatic-speech-recognition1K<n<10K19 likes650 downloads4y agoHugging Face04caomingpei /atcoder-problemstext1K<n<10K0 likes508 downloads2y agoHugging Face05Jzuluaga /uwb_atcc Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/uwb_atcc.audioautomatic-speech-recognition10K<n<100K16 likes443 downloads4y agoHugging Face06Jzuluaga /atco2_corpus_1h Dataset Card for ATCO2 test set corpus (1hr set) Dataset Summary ATCO2 project aims at developing a unique platform allowing to collect, organize and pre-process air-traffic control (voice communication) data from air space. This project has received funding from the Clean Sky 2 Joint Undertaking (JU) under grant agreement No 864702. The JU receives support from the European Union’s Horizon 2020 research and innovation programme and the Clean Sky 2 JU members other than… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atco2_corpus_1h.audioautomatic-speech-recognitionn<1K12 likes411 downloads4y agoHugging Face07twangodev /tartanaviation-atc-labels TartanAviation ATC ASR Labels Machine transcripts and confidence scores for twangodev/tartanaviation-atc-adsb-utterances. A 1:1 labels-only add-on (no audio): one row per source utterance, same shards and row order, keyed by utterance_id. 531,050 labels · 184 shards · 100% coverage · ensemble ASR + weighted ROVER + ADS-B callsign snap · ~326 human-reviewed. Built with readback. Usage Join 1:1 onto the source. Rows are aligned and in the same order: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-labels.tabularautomatic-speech-recognition100K<n<1M0 likes267 downloads3mo agoHugging Face08jlvdoorn /atco2-asr-atcosim Dataset Card for "atco2-asr-atcosim" This is a dataset constructed from two datasets: ATCO2-ASR and ATCOSIM. It is divided into 80% train and 20% validation by selecting files randomly. Some of the files have additional information that is presented in the 'info' file. audioautomatic-speech-recognition10K<n<100K16 likes255 downloads3y agoHugging Face09Dbdn /atc-testaudio100K<n<1M0 likes201 downloads2y agoHugging Face10jacktol /atc-dataset Deprecation Notice (June 16, 2025) This dataset is deprecated. A newer version is available on Hugging Face Datasets, featuring higher data quality and reliability, although with fewer overall samples. View the updated dataset on Hugging Face ATC Dataset - Fine-Tuning Whisper This dataset was created to fine-tune OpenAI's Whisper model for improving transcription accuracy in Air Traffic Control (ATC) communications. The dataset contains transcriptions and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/jacktol/atc-dataset.audio10K<n<100K10 likes193 downloads1y agoHugging Face11jlvdoorn /atcosimThis is an ATM dataset for the use of automatic speech recognition. The original source of the data is from the ATCOSIM project. audio1K<n<10K8 likes191 downloads3y agoHugging Face12jlvdoorn /atco2-asr Dataset Card for "ATCO2-ASR" This is audio data used for automatic speech recognition. The original source of the data is the ATCO2 project, specifically the ASR part of the public speech corpus. audion<1K8 likes173 downloads3y agoHugging Face13Nan-Do /atcoder_cot Dataset Card for Atcoder-CoT Dataset Description Atcoder-CoT is a proof-of-concept dataset designed to demonstrate how a dataset like the one found here can be used to generate synthetic datasets for training reasoning models, particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation. It leverages human-created and debugged solutions, combined with LLM-generated text to create conversational turns. The approach can also be easily adapted to simulate human… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_cot.texttext-generation10K<n<100K1 likes172 downloads1y agoHugging Face14Shiry /ATC_combined Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Shiry/ATC_combined.audioautomatic-speech-recognition10K<n<100K1 likes165 downloads3y agoHugging Face15jacktol /ATC-ASR-Dataset ATC ASR Dataset ATC ASR Dataset is a high-quality, fine-tuning-ready speech recognition dataset constructed from two real-world Air Traffic Control (ATC) corpora: the UWB ATC Corpus and the ATCO2 1-Hour Test Subset. The dataset consists of cleanly segmented audio + transcript pairs at the utterance level, specifically curated for Automatic Speech Recognition (ASR) training and fine-tuning in the ATC domain. Contents This dataset includes: Audio files (.wav, 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/jacktol/ATC-ASR-Dataset.audio1K<n<10K18 likes144 downloads1y agoHugging Face16allenpoe /atcosim_dataset_for_finetune_whisper_smallaudio1K<n<10K0 likes143 downloads2y agoHugging Face17mmonzel /uwb_atcosim_split_by_person_6_2_2 UWB-ATCC + ATCOSIM — combined ATC ASR dataset (ATCOSIM 6-2-2, gender-balanced) A combined air-traffic-control ASR dataset built from Jzuluaga/uwb_atcc and Jzuluaga/atcosim_corpus, re-split with leakage-free (group-disjoint) train/validation/test splits. UWB-ATCC is split recording-disjoint (80/10/10) and ATCOSIM is split speaker-disjoint and gender-balanced, putting 1 male + 1 female speaker in each of validation and test (6 train / 2 / 2 speakers), with a fixed seed (42) and 16… See the full description on the dataset page: https://huggingface.co/datasets/mmonzel/uwb_atcosim_split_by_person_6_2_2.audio10K<n<100K1 likes118 downloads3mo agoHugging Face18mmonzel /uwb_atcosim_split_by_person_8-1-1 UWB-ATCC + ATCOSIM — combined ATC ASR dataset (ATCOSIM 8-1-1) A combined air-traffic-control ASR dataset built from Jzuluaga/uwb_atcc and Jzuluaga/atcosim_corpus, re-split with leakage-free (group-disjoint) train/validation/test splits. UWB-ATCC is split recording-disjoint (80/10/10) and ATCOSIM is split speaker-disjoint into 8 train / 1 validation / 1 test speakers, with a fixed seed (42) and 16 kHz audio. Columns: id, audio, text, segment_start_time, segment_end_time… See the full description on the dataset page: https://huggingface.co/datasets/mmonzel/uwb_atcosim_split_by_person_8-1-1.audio10K<n<100K1 likes109 downloads3mo agoHugging Face19Tabys /ATC_combined Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Tabys/ATC_combined.audioautomatic-speech-recognition10K<n<100K0 likes96 downloads7mo agoHugging Face20allenpoe /atcosim_converted_prepared_whisper_base1K<n<10K0 likes94 downloads2y agoHugging Face21allenpoe /atcosim_converted_prepared_whisper_small1K<n<10K0 likes91 downloads2y agoHugging Face22allenpoe /atcosim_prepared_whisper_base1K<n<10K0 likes88 downloads2y agoHugging Face23KaranChand /atcosim_inputtext10K<n<100K0 likes86 downloads4y agoHugging Face24KaranChand /atcosim_pruned_huberttext1K<n<10K0 likes77 downloads4y agoHugging Face25KaranChand /atcosim_pruned_xlsrtext1K<n<10K0 likes70 downloads4y agoHugging Face26Nan-Do /atcoder_abc_contests_smallgated Dataset Summary This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques. It also serves as a robust foundation for problem-solving in Large Language Models (LLMs). The dataset includes both accepted and failed solutions from Atcoders's (ABC) contests. In total, it features… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests_small.texttext-generation100K<n<1M6 likes68 downloads1y agoHugging Face27cmcmaster /pbs_item_atc_relationshipstabular10K<n<100K0 likes66 downloads9mo agoHugging Face28luigisaetta /atco2_normalized_augmentedaudio1K<n<10K0 likes65 downloads4y agoHugging Face29luigisaetta /atco2_only_augmentedaudioautomatic-speech-recognition1K<n<10K1 likes60 downloads4y agoHugging Face30SAadettin-BERber /normalized_train_ATC_datasetaudio1K<n<10K0 likes60 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.