datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tartanaviation-atc-adsb-utterances
TartanAviation ATC + ADS-B (Utterances)
Speech utterances split from twangodev/tartanaviation-atc-adsb
by voice-activity detection (pyannote/segmentation-3.0).
Each row is one speech segment (16 kHz mono) with the ADS-B from its parent clip.
531,050 utterances · ~398 h speech · 16 kHz mono · 67% carry ADS-B. From 40,899 of 41,823 clips
(silent clips have no utterances). Built with squawk.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb-utterances.tartanaviation-atc-adsb
TartanAviation ATC + ADS-B
Paired ATC audio and ADS-B for Pittsburgh KAGC and KBTP, aligned from CMU
TartanAviation. Each row is one ADS-B-triggered
audio capture (16 kHz mono) plus the aircraft tracks present during it.
41,823 clips · 16 kHz mono · 67% carry ADS-B. Built with squawk.
Usage
from datasets import load_dataset
ds = load_dataset("twangodev/tartanaviation-atc-adsb", split="train", streaming=True)
ex = next(iter(ds))
ex["audio"] # {'array': ...… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb.atcosim_corpus
Dataset Card for ATCOSIM corpus
Dataset Summary
The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus.uwb_atcc
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/uwb_atcc.atco2_corpus_1h
Dataset Card for ATCO2 test set corpus (1hr set)
Dataset Summary
ATCO2 project aims at developing a unique platform allowing to collect, organize and pre-process air-traffic control (voice communication) data from air space. This project has received funding from the Clean Sky 2 Joint Undertaking (JU) under grant agreement No 864702. The JU receives support from the European Union’s Horizon 2020 research and innovation programme and the Clean Sky 2 JU members other than… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atco2_corpus_1h.tartanaviation-atc-labels
TartanAviation ATC ASR Labels
Machine transcripts and confidence scores for
twangodev/tartanaviation-atc-adsb-utterances.
A 1:1 labels-only add-on (no audio): one row per source utterance, same shards and row order, keyed
by utterance_id.
531,050 labels · 184 shards · 100% coverage · ensemble ASR + weighted ROVER + ADS-B
callsign snap · ~326 human-reviewed. Built with readback.
Usage
Join 1:1 onto the source. Rows are aligned and in the same order:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-labels.atco2-asr-atcosim
Dataset Card for "atco2-asr-atcosim"
This is a dataset constructed from two datasets: ATCO2-ASR and ATCOSIM.
It is divided into 80% train and 20% validation by selecting files randomly. Some of the files have additional information that is presented in the 'info' file.
ATC_combined
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Shiry/ATC_combined.ATC_combined
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Tabys/ATC_combined.atco2_only_augmentedatc-asr-eval
ATC ASR Evaluation Set
Human-verified air-traffic-control transmissions for evaluating ASR models.
Each clip is 16 kHz mono WAV with a corrected ground-truth transcript in
metadata.csv (columns: file_name, transcription, icao, city, region, source, feed).
Audio captured from LiveATC.net feeds. Private — not for redistribution
(LiveATC terms prohibit rebroadcasting).
atco2-asr-atcosim
Dataset Card for "atco2-asr-atcosim"
This is a dataset constructed from two datasets: ATCO2-ASR and ATCOSIM.
It is divided into 80% train and 20% validation by selecting files randomly. Some of the files have additional information that is presented in the 'info' file.
atcsuwb_atcc
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/uwb_atcc.ATCO2-ASR-1hatcosim_corpus
Dataset Card for ATCOSIM corpus
Dataset Summary
The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/atcosim_corpus.synthetic-atc-speech
Synthetic ATC Speech
Synthetic English air-traffic-control speech created for research on robust
automatic speech recognition. The dataset contains 276,304 generated
utterances from 15,660 unique ATC transcripts.
The dataset accompanies:
Contrastive Regularization for Accent-Robust ASR
Robust ATC ASR code
UWB SupCon Hybrid model
UWB+ATCOSIM SupCon Hybrid model
Dataset Structure
The dataset provides one training split packaged as uncompressed WebDataset TAR… See the full description on the dataset page: https://huggingface.co/datasets/ThaiVanPhat95/synthetic-atc-speech.ATC56k
ATC Numeric Speech Dataset / 航空数字语音数据集
Item / 项目
Value / 数量
Total rows / 总行数
55,560
Clean / 干净音频
18,640
Radio/noisy / 电台噪声音频
36,920
Speakers / 说话人
spk001=21,480, spk002=21,480, spk003=4,200, spk004=4,200, spk005=4,200
Accents / 口音
standard=27,780, 陕西话=27,780
Splits / 划分
dev=5,760, test=5,760, train=44,040
Reading styles / 数字读法
full=37,280, short=18,280
Speed coverage / 速度覆盖
100-990, 90 values
Altitude coverage / 高度覆盖
1000-9900, 90 values… See the full description on the dataset page: https://huggingface.co/datasets/ymj123/ATC56k.ATC-ASR-Dataset
ATC ASR Dataset (normalized)
jacktol/ATC-ASR-Dataset (UWB ATC Corpus + ATCO2 1-hour test subset, 8,122 utterances, 16 kHz) with one added column, normalized_text: the uppercase spoken-form transcript rewritten into written form. id, audio and text are byte-identical to the source, and so are the splits.
text (source)
normalized_text
CSA ZERO TWO FIVE HEADING THREE ONE FIVE DESCEND FLIGHT LEVEL NINER ZERO
CSA 025, heading 315, descend flight level 90.
LUFTHANSA FIVE… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/ATC-ASR-Dataset.
