datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.en-sumerize
Dataset Card for GEM/xlsum
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
XLSum is a highly multilingual summarization dataset supporting 44 language. The data stems from BBC news articles.
You can load the dataset via:
import datasets
data = datasets.load_dataset('GEM/xlsum')
The data loader can be found here.
website
Github
paper
ACL Anthology
Dataset Overview
Where to find… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/en-sumerize.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.physdojo-annotations
PhysDojo: Dense Physics Annotations for DreamDojo
Per-frame per-object physics labels extracted from DreamDojo robot manipulation episodes
using off-the-shelf foundation models (no VLM).
Pipeline
RAFT optical flow → per-pixel velocities
Depth Anything V2 → metric depth → 3D positions
DINOv2 → material features
Format
Each episode has physics_dense.npz containing:
positions: (T, N_obj, 2) — normalized [0,1] centroids
velocities: (T, N_obj, 2) — pixels/sec… See the full description on the dataset page: https://huggingface.co/datasets/shanyangmie/physdojo-annotations.OmniSET2ears_dataset
