datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Table-GPT
Table-GPT: Table-tuned GPT for Diverse Table Tasks
This repository contains training and test datasets for the SIGMOD'24 paper Table-GPT: Table-tuned GPT for Diverse Table Tasks. The source code for data generation and task evaluation are available here: https://github.com/microsoft/Table-GPT, which can be used to generate more training data for table-related tasks.
Task Descriptions
We collect (or synthesize) 18 diverse table-related tasks, which are summarized in… See the full description on the dataset page: https://huggingface.co/datasets/LipengCS/Table-GPT.ave-speech-lipemg-processed
AVE Speech — preprocessed for lip–EMG fusion
Code: diddmstjr07/silent-speech-viseme-emg — model definition, training, evaluation, and the analyses behind these numbers.
Related releases: checkpoints · AVE preprocessed · Confusable-100
A derivative of the AVE Speech corpus
(Zhou et al., IEEE THMS 2025), preprocessed into the exact form used to train the
lip–EMG fusion models in the companion work. This is not new recorded data; it is the
original corpus with the preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/diddmstjr/ave-speech-lipemg-processed.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.confusable-100-lipreading
Confusable-100: an English vocabulary built to break lip reading
A small, deliberately adversarial corpus for probing one specific failure mode of
visual speech recognition: consonants articulated by the tongue leave no distinctive
trace on the lips, so words differing only in those consonants are not separable from
video in principle, not merely in practice.
This is an evaluation probe, not a training set. It is one speaker and 3.4 minutes
of speech. Its purpose is to expose… See the full description on the dataset page: https://huggingface.co/datasets/diddmstjr/confusable-100-lipreading.powergrid-benchmark2
Note: This is a dummy dataset for browsing only. The official dataset can be downloaded from our pipeline (Data Hub).
Full dataset for Benchmark 2
powergrid-benchmark1
Note: This is a dummy dataset for browsing only. The official dataset can be downloaded from our pipeline (Data Hub).
Full dataset for Benchmark 1
chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.novel_cn_roleplay_dataset_liars_lips_fall_apart_in_loveThis is a CN roleplay dataset extracted from the novel https://www.bilinovel.com/novel/4482.html
liputan6
Liputan6
Liputan6 is a large-scale Indonesian text summarization dataset of news articles from Liputan6.com. The canonical configuration contains 215,827 document-summary pairs. The dataset also provides the XTREME preprocessing variant.
Dataset details
Language: Indonesian (ind)
Task: text summarization
License: CC BY-SA 4.0
Canonical configuration: 193,883 train, 10,972 validation, and 10,972 test examples
XTREME configuration: 175,207 train, 4,948 validation… See the full description on the dataset page: https://huggingface.co/datasets/joshuasiagian/liputan6.aifgen-lipschitz
Dataset Card for Dataset Name
This dataset is a continual dataset in lipschitz bounded scenario given three tasks:
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm 5
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm a High School Student
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm an expert
Dataset Details
Dataset Description
As a… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-lipschitz.ruwiki_cleaned3S_Ing._NSFW_Storylipidos-semantic-check-testcaseNSFW_Story2NSFW_StoryNSFW_Story4NSFW_gemma-2b-it-2NSFW_Roleplay-gemma-2bNSFW_gemma-2b-it3SIng.-NSFW_Story23S_Ing._NSFW_Story3lipo4NSFW_Story3lipo2nsfw_reddit_cleanednsfw_clean2ipa-phoneme-to-words-v9xxxlipo1lipo3
