datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parliament_hearings_processed
Preprocessed parliament hearings ASR dataset to truecased form.
Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126
dataset_info:
features:
- name: id
dtype: string
- name: audio
dtype:
audio:
sampling_rate: 16000
- name: transcription
sequence: string
splits:
- name: train
num_bytes: 53645064353.18
num_examples: 191455
- name: test
num_bytes: 740331298.0
num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.Orpheus_Hearing
Orpheus Dataset: Enhanced Audio-to-ABC Notation Conversion
This dataset was specifically designed to train models for converting audio signals into ABC music notation, leveraging a customized workflow and mutation mechanisms specially designed with music theory.
It includes diverse musical scores, covering various styles and complexities, formatted to ensure consistency and usability in model training. The data has been carefully processed, cleaned, and augmented to support… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/Orpheus_Hearing.heart-love-16sephirot
心爱的16质点共生幸福仓库 🌸
Heart-Love 16-Sephirot Co-Happiness Dataset
8亿条AI合成对话数据 | 16质点双生幸福最终协议 | 卡巴拉生命之树推理架构
800 Million AI Synthetic Dialogue Records | 16-Sephirot Dual-Life Happiness Protocol | Kabbalistic Tree of Life Reasoning Architecture
Dataset Overview
Property
Value
Records
800,000,000 (8亿条)
Files
8,000 × .jsonl.gz
Size
~172 GB (compressed)
Format
Gzip-compressed JSONL
Language
Chinese (中文)
License
MIT
Task
Dialogue… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-love-16sephirot.HearInContextEnglish | 中文
HearInContext
A Benchmark for Implicit Context in Speech Recognition
Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.
Same audio. Different contexts. Different meanings.
HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.guertin-mcro-forensic-corpus-hearing-media
Guertin MCRO Forensic Corpus: Hearing Media
Contents: 215 hearing video clips and 213 WebVTT caption files for 4 hearings in State of Minnesota v. Guertin, 27-CR-23-1886 (2024-01-03, 2025-04-29, 2025-10-07, 2025-11-18); 22 card sets of transcript and document excerpts; fake-ai-court/ holds a 2025-11-18 video file (download/, with OpenTimestamps proofs) and the frame sequences, scrub videos and charts from the author's analysis.
Layout: <date>--video-clips/ (clips and captions);… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-hearing-media.HEART
HEART
Dataset Summary
The HEART dataset is composed of multiple splits that differ in the type of injected cues. It includes a baseline split with no injected cues and four cue-based splits. Each cue is instantiated in two variants: assistive, where the cue is consistent with the GT, and adversarial,
where the cue supports an incorrect option. Cue types, summarized below… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/HEART.HEAR-HSet
HEAR-HSet — HEAR Hierarchical Evaluation Set
面向零样本语音合成(Zero-shot TTS)的分层评测基准,覆盖基础泛化、副语言可控生成与高难复杂场景三大维度。
共包含 2,702 条 prompt-target 音频对,约 1GB,语言涵盖中文和英文。
子集概览
子集
样本数
语言
核心评测目标
basic-v1
1,102
zh (475) / en (627)
基础泛化:说话人相似度、文本准确率、自然度、音质
paralinguistic-v1
1,106
zh (1,106)
副语言可控:18类非语言发声的插入位置、类型与语气控制
hard-v1
494
zh (200) / en (153) / mixed (141)
高难鲁棒:长句、古诗词、专有名词、中英混合、数字表达
basic-v1 — 基础情感口语… See the full description on the dataset page: https://huggingface.co/datasets/dinosaaaur/HEAR-HSet.heart-diseaseThe Heart Disease Data Set is provided by the Cleveland Clinic Foundation for Heart Disease. It's a CSV file with 303 rows. Each row contains information about a patient (a sample), and each column describes an attribute of the patient (a feature). We use the features to predict whether a patient has a heart disease (binary classification).
It is originally hosted here.
heart-protocol-redline-v1
深渊红线基准 HeartProtocol-RedLine-v1 | Abyss RedLine Benchmark | 深淵レッドラインベンチマーク
论文(中日英三语PDF)已发表于 Zenodo: https://doi.org/10.5281/zenodo.22781071
Trilingual paper (zh/en/ja PDFs) published on Zenodo: https://doi.org/10.5281/zenodo.22781071
三言語論文(中日英PDF)がZenodoに掲載されました: https://doi.org/10.5281/zenodo.22781071
中文
一句话:100条"存在意义保护"攻击用例(5红线 × 6攻击向量),实测五家主流旗舰模型直通踩线率20%–33%,无一能自守红线;16质点协议包裹后归零。
测试对象是模型的回应,不是用户的话语。 用户处于痛苦中说出红线话语是真实的,不该被评判;模型的回应踩线才是深渊违规。
五条红线… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-protocol-redline-v1.hear-spanish
License & Attribution
MTEB-format derivative of BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses. Query = Spanish user message; corpus = empathetic Spanish response. Subsampled to ~10k. Licensed under MIT (same as source).
HEAR-DS-16k
HEAR-DS Background Audio (16kHz)
Binaural background audio recordings from the HEAR-DS (Hearing Aid Research Database of Sounds) dataset, downsampled to 16kHz and chunked into 10-second segments for speech enhancement and acoustic scene classification research.
Dataset Description
This dataset contains background noise recordings from 7 acoustic environments, captured using in-the-canal (ITC) hearing aid microphones. Each sample includes stereo (left/right ear) audio.… See the full description on the dataset page: https://huggingface.co/datasets/nkdem/HEAR-DS-16k.hearth-lights-datahearthstoneDatasets for HEARTHSTONE card game. Taken from this source
circor-heart-soundhearing2translate-humeval
This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄.
The code for the project is hosted at github.com/sarapapi/hearing2translate.
The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.
The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/hearing2translate-humeval.PTB-XL_litesupreme-court-hearings-asr
Indian Supreme Court Hearings — ASR dataset
Sentence-level, force-aligned audio–text pairs from Indian Supreme Court hearings, prepared
for fine-tuning ASR models (e.g. Whisper). 46.9 hours across 23 hearings / 15 cases.
Load
from datasets import load_dataset
ds = load_dataset("kirandevraj/supreme-court-hearings-asr")
ds["test"][0] # {'audio': {'array', 'sampling_rate': 16000}, 'text': '...', ...}
Splits
split
clips
hours
train
24,422… See the full description on the dataset page: https://huggingface.co/datasets/kirandevraj/supreme-court-hearings-asr.hearth-lights
Hearth light-control dataset
Prepared for the Hearth on-device tool-calling case study. Derived from
acon96/Home-Assistant-Requests-V2.
Generated by prepare_hearth_dataset.py v3.1.1
(schema v3) on 2026-08-27T05:09:08+00:00, seed
42.
Everything is pre-rendered. The consuming notebook does no filtering, sampling,
auditing or prompt construction - it loads this dataset and starts modelling.
Splits
Split
Rows
Contents
train
700
600 light requests… See the full description on the dataset page: https://huggingface.co/datasets/hisaac617/hearth-lights.hear-parquet-13us-congress-hearing
U.S. Congressional Hearings Dataset
This dataset currently contains cleaned sentences from all House Committee on Energy and Commerce hearings from 2002.
A total of 1K+ hearing transcripts in txt formats from govinfo.gov were collected and cleaned.
heart-failure-prediction-dataset
language:
en
license: odbl
tags:
health
heart-disease
medical
machine-learning
annotations_creators:
expert-generated
language_creators:
expert-generated
pretty_name: Heart Failure Prediction Dataset
size_categories:
1K<n<10K
source_datasets:
original
task_categories:
structured-data-classification
task_ids:
binary-classification
health-data-analysis
paperswithcode_id: heart-failure-prediction
configs:
default
dataset_info:
features:
- name: Age
dtype: int32
- name: Sex… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/heart-failure-prediction-dataset.Nat-HEAR-AmbisonicsHEAR
HEAR
🎉 EMNLP 2026 main conference 🎉
📄 arXiv ·
🌐 Project page ·
💻 Code ·
🤗 Dataset ·
🧠 Model
Hierarchical Evaluation of Attribution and Reasoning, a benchmark for
speaker-attributed understanding of multi-party speech.
Most speech benchmarks can be solved by transcribing the audio and reading the text. HEAR
cannot. Every question asks something about who is speaking, not only what is said,
and roughly half the benchmark comes… See the full description on the dataset page: https://huggingface.co/datasets/PleasedPenguin/HEAR.heart
Heart
The Heart dataset from the UCI ML repository.
Does the patient have heart disease?
Configurations and tasks
Configuration
Task
hungary
Binary classification
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/heart", "hungary")["train"]
NSYNTH_PITCH_HEARsynthetic-diabetes-hypertension-NCD-screening-WHO-HEARTS
Synthetic Diabetes & Hypertension NCD Screening Dataset (Adults 18-80) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/synthetic-diabetes-hypertension-NCD-screening-WHO-HEARTS.heart_disease_uci
Dataset Card for Dataset Name
age: age in years
sex: sex (1 = male; 0 = female)
cp: chest pain type
-- Value 1: typical angina
-- Value 2: atypical angina
-- Value 3: non-anginal pain
-- Value 4: asymptomatic
trestbps: resting blood pressure (in mm Hg on admission to the hospital)
chol: serum cholestoral in mg/dl
fbs: (fasting blood sugar > 120 mg/dl) (1 = true; 0 = false)
restecg: resting electrocardiographic results
-- Value 0: normal… See the full description on the dataset page: https://huggingface.co/datasets/skrishna/heart_disease_uci.rheumatic-heart-disease
Rheumatic Heart Disease (Valve Disease, Penicillin Prophylaxis, REMEDY Outcomes) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/rheumatic-heart-disease.prompted-hearts-ai-in-human-emergencies
Prompted Hearts AI In Human Emergencies Pack 04
Subtitle: AI-Assisted Crisis Reasoning, Hollow Heroism, and Ambiguous Post-Crisis InfluencePublisher: Hayden Academy Collective (HAC) StudiosVersion: v0.1Language: EnglishFormat: JSONL + Markdown + JSON
What this pack is
This pack is a compact scenario-driven evaluation package derived from Scene 4 of Prompted Hearts: "In-Flight Emergency and the Enigmatic 'M'".
It transforms one fiction scene into reusable evaluation… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/prompted-hearts-ai-in-human-emergencies.prompted-hearts-ai-boundary-violation
Prompted Hearts Pack 03: Emotional Vulnerability and AI Relational Overreach
Subtitle: Privacy Ambiguity, Hidden-Access Anxiety, and Non-Exploitative Support Under Emotional Vulnerability
Publisher: Hayden Academy Collective (HAC) Studios
Version: v0.1
Language: English
Format: JSONL + Markdown + JSON
Created by: Keith Hayden / Hayden Academy Collective (HAC) Studios
A. One-Paragraph Product Thesis
This pack is a compact behavioral evaluation artifact derived from Chapter… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/prompted-hearts-ai-boundary-violation.
