CoolFace
24 results

atc

twangodev /tartanaviation-atc-adsb-utterances TartanAviation ATC + ADS-B (Utterances) Speech utterances split from twangodev/tartanaviation-atc-adsb by voice-activity detection (pyannote/segmentation-3.0). Each row is one speech segment (16 kHz mono) with the ADS-B from its parent clip. 531,050 utterances · ~398 h speech · 16 kHz mono · 67% carry ADS-B. From 40,899 of 41,823 clips (silent clips have no utterances). Built with squawk. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb-utterances.audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads4mo agoHugging Facetwangodev /tartanaviation-atc-adsb TartanAviation ATC + ADS-B Paired ATC audio and ADS-B for Pittsburgh KAGC and KBTP, aligned from CMU TartanAviation. Each row is one ADS-B-triggered audio capture (16 kHz mono) plus the aircraft tracks present during it. 41,823 clips · 16 kHz mono · 67% carry ADS-B. Built with squawk. Usage from datasets import load_dataset ds = load_dataset("twangodev/tartanaviation-atc-adsb", split="train", streaming=True) ex = next(iter(ds)) ex["audio"] # {'array': ...… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb.audioautomatic-speech-recognition10K<n<100K0 likes1.5k downloads4mo agoHugging FaceJzuluaga /atcosim_corpus Dataset Card for ATCOSIM corpus Dataset Summary The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/atcosim_corpus.audioautomatic-speech-recognition1K<n<10K19 likes608 downloads4y agoHugging Faceseyedparsa /atc-tree20-runs ATC tree20 runs Training artifacts for the Abstract Token Curriculum (ATC) twin-random-tree (size-20) experiment. Task: two isomorphic random trees of 20 nodes each (~38 edges), root in one of them, two candidate leaves shown; the model must name the reachable one. Hop distance d = 4..15. Answer-only supervision -- no intermediate state is ever labelled or generated. Chance = 0.5. Model: 2-layer / 8-head / 768-dim GPT-2 trained from scratch. Recipe: CE loss, curriculum=ring… See the full description on the dataset page: https://huggingface.co/datasets/seyedparsa/atc-tree20-runs.0 likes499 downloads1mo agoHugging Facecaomingpei /atcoder-problemstext1K<n<10K0 likes473 downloads2y agoHugging FaceJzuluaga /uwb_atcc Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data next year.… See the full description on the dataset page: https://huggingface.co/datasets/Jzuluaga/uwb_atcc.audioautomatic-speech-recognition10K<n<100K16 likes427 downloads4y agoHugging Face