datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emolia-thinking
Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia
Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.thinkvoice-dataset-v3combinedAF-Think-audiosthinkomni_eval
ThinkOmni Evaluation Dataset
This repository contains the evaluation datasets for ThinkOmni, a training-free framework that lifts textual reasoning to omni-modal scenarios via guidance decoding.
ThinkOmni enhances omni-modal large language models (OLLMs) with the reasoning capabilities of large reasoning models (LRMs) at decoding time, adaptively balancing perception and reasoning signals.
Paper: ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding… See the full description on the dataset page: https://huggingface.co/datasets/Catalan258/thinkomni_eval.majestrino-thinkinglaions-got-talent-thinkingRatchada-STT
RATCHADA-STT Dataset
Overview
The dataset includes recordings from earnings calls of publicly traded companies in Thailand. Each audio file is accompanied by a transcription and metadata such as company name, reporting period, and other relevant details.
Dataset Info
Total Duration:
Train: 26507.87 seconds (~ 7.36 hours)
Test: 8376.28 seconds (~ 2.33 hours)
File Count:
Train: 10912 files
Test: 2804 files
Dataset Structure
The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingMachinesDataScience/Ratchada-STT.Thinkvoice_combined_v2thinker-talker-datasetsmultilingual-in-the-wild-thinkingThinkvoice_combined_v1
