datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oai-5g-srs-ranging-dataset
OAI 5G NR SRS Ranging Captures
Uplink Sounding Reference Signal (SRS) channel-estimate captures from a monolithic 5G NR
software-defined-radio testbed, collected for SRS-based ranging experiments.
The gNB (OpenAirInterface on a USRP X410, Band n78, 40 MHz / 106 PRB) configures each UE
to transmit SRS; the gNB's per-SRS frequency-domain channel estimate, oversampled IDFT CIR,
and ToA estimate are streamed off the PHY via OAI's T_tracer and recorded at a series of
known… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-srs-ranging-dataset.spoofed_datasetlibrispeech10hfish_speech_casablancaAHA
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
Paper | GitHub
AHA (Audio Hallucination Alignment) is a framework and dataset designed to mitigate hallucinations in Large Audio-Language Models (LALMs). By focusing on fine-grained temporal reasoning and counterfactual alignment, the AHA dataset helps models distinguish between what sounds plausible and what is actually present in an audio stream.
The dataset addresses four key… See the full description on the dataset page: https://huggingface.co/datasets/ASU-GSL/AHA.AHa-Benchkoochik-finetune-dataAHa-Benchtarget-words-geminitts
Target-Word (TW) Evaluation Set
Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how
well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER /
recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with
Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis
"Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.aha-workshoptarget_wordstarget_words_lexical_adaptationtesta-ha-2AHao-dataset-audio-whisper-viAHATETaha2342026_1
