datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr-ser-quechua-collao-embeddings
ASR-SER embeddings for Quechua Collao
This repository contains embeddings only. It does not contain raw audio.
These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis.
Dataset contents
One PyTorch tensor per utterance stored as an embedding file under embeddings/
A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.asrs-aviation-reports
Dataset Card for ASRS Aviation Incident Reports
Dataset Summary
This dataset collects 47,723 aviation incident reports published in the Aviation Safety Reporting System (ASRS) database maintained by NASA.
Supported Tasks and Leaderboards
'summarization': Dataset can be used to train a model for abstractive and extractive summarization. The model performance is measured by how high the output summary's ROUGE score for a given narrative account of an aviation… See the full description on the dataset page: https://huggingface.co/datasets/elihoole/asrs-aviation-reports.ASR-santali_100hrsasr-spell-correction-ruasr_spell_correction_ruASR-Santali_4hrsasr-slu
Dataset Card for "asr-slu"
More Information needed
asr-sinhala-dataset_check1_modifiedasr-spell-correction-ruasr-spell-correction-ru-hw01
Russian ASR correction: homework 01
1020 pairs: 450 Groq-generated ASR-like inputs,
180 Groq-generated numeral-to-word pairs, and 390 identity
examples added by copying screened clean targets.
Model: openai/gpt-oss-120b. Generation: 8bae1a605e8506a1; prompt version: groq_asr_numbers_v2.
The ASR target names come from the existing Groq-generated pool targets.jsonl.
No Python character corruption is used. This is synthetic text, not real ASR output.
Generation and… See the full description on the dataset page: https://huggingface.co/datasets/sobadsodead/asr-spell-correction-ru-hw01.aerograph-asrs
AeroGraph ASRS Dataset
2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports
with LLM-extracted entities and relations for knowledge graph construction.
Dataset Description
This dataset contains processed ASRS incident narratives along with
structured entity and relation extractions conforming to an aviation
safety ontology (10 entity types, 8 edge types).
Reports Split
2000 reports from the NASA ASRS database
Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.asr_spell_correction_ruasr-sinhala-dataset_json_v1asr-spell-ruasr-semantic-probe-eng
ASR Semantic Probing Dataset (English)
Synthetic English audio dataset for probing whether ASR encoder representations
encode semantic category information beyond acoustic features. Constructed for
mechanistic interpretability studies of speech recognition models.
Splits
This dataset is released as a single unsplit collection. Downstream users are
expected to define their own train/test splits based on the experimental design.
For probing experiments where speaker… See the full description on the dataset page: https://huggingface.co/datasets/soaring0616/asr-semantic-probe-eng.asr-sinhala-dataset_v3asr-sinhala-dataset_json_v2asrs-narrativesasr-sinhala-dataset_v2asrs-aviation-alpacaASR_Stroke_DatasetASRS-ChatGPT
Dataset Summary
The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives.
The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available.
Dataset Structure
The column names present in the dataset and their descriptions are provided below:
Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.asrs-narratives-rebalanceasr-sinhala-dataset_v4asrs-summarizationasrs-procedure-human-factor-fine-tuneasrs-contributing-factorsasrs-synopses-pre-trainingasrsasrs-data-contrastive-non-aircraft
