datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
S2-TOMG-Bench
S^2-Bench Dataset (TOMG) (full version, 45k entries)
Official Huggingface Datasets for S^2-Bench: "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation"
Please refer to our Github Repo for more usage and useful information.
Configurations
Each configuration represents a different task:
MolCustom_AtomNum: Molecular customized generation by atom number
MolCustom_BondNum: Molecular customized generation by bond number… See the full description on the dataset page: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench.s2tt-yoruba-englishS2-TOMG-Bench-mini
S^2-Bench Dataset (TOMG) (mini version, 4.5k entries)
Official Huggingface Datasets for S^2-Bench: "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation"
Please refer to our Github Repo for more usage and useful information.
This mini version offers a choice for researchers with limited resources.
Configurations
Each configuration represents a different task:
MolCustom_AtomNum: Molecular customized generation by atom number… See the full description on the dataset page: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench-mini.s2tt-igbo-englishs2tt-hausa-englishS2TLDhttps://github.com/Thinklab-SJTU/S2TLD
S2T_Split_NoRom_phase2S2T_Korean_Merge_2_fixed4S2T_English_Vietnamese_AuVi_2S2TD-Face
Data Card for S2TD-Face
This repository provides the data used in ACM MM 2024 paper S2TD-Face. Please see our github repository for details.
Citation
If you use our work in your research, please cite our publication:
@inproceedings{wang2024s2td,
title={S2TD-Face: Reconstruct a Detailed 3D Face with Controllable Texture from a Single Sketch},
author={Wang, Zidu and Zhu, Xiangyu and Yu, Jiang and Zhang, Tianshuo and Lei, Zhen},
booktitle={Proceedings of the 32nd ACM… See the full description on the dataset page: https://huggingface.co/datasets/Zidu-Wang/S2TD-Face.Albayzin-2024-BBS-S2T-eval
Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge - Evaluation dataset
see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge.
This is the evaluation data for the challenge.
The database consists of a single split:
eval : 12498 audio segments
How to download this database
1 - If you can handle yourself comfortably with Huggingface Datasets:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T-eval.s2t-augmented-data
Dataset Card for "s2t-augmented-data"
More Information needed
S2T_Korean_limit_silent_spaceAlbayzin-2024-BBS-S2T
Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge
see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge.
NOTE: Test data will be released on September 2nd, 2024.
The Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge
training and tuning set is based on the gttsehu/basque_parliament_1
dataset. The database consists of four splits:
train : 749945 audio segments (automatically extracted)
train_clean : 661871 audio segments… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T.en2ja.s2t_translationS2T_Korean_Merge_2_fixed2ja2en.s2t_translationS2T_Korean_3s_silentS2TTruozhiba_s2t
Dataset Card for "ruozhiba_s2t"
More Information needed
srt-demo-s2tt-70S2T_English_Vietnamese_AuVi_testS2T_Koreanfleurs_eng_test_s2ttS2T_Korean_MergeS2T_Korean_Merge_2s2tkp_dataset_part_3multi-lang-s2t-dataset
Multi-Language Audio Dataset
A high-quality, cleaned audio dataset derived from five AI4Bharat sources, containing speech-to-text data in Hindi, Gujarati, and Telugu with English translations.
Dataset Overview
Language
Duration (Hours)
Hindi
55.45
Gujarati
42.12
Telugu
37.83
Total Duration: ~135 hours of cleaned, high-quality audio
Source Datasets
Data collected and cleaned from:
ai4bharat/IndicVoices-ST
ai4bharat/NPTEL… See the full description on the dataset page: https://huggingface.co/datasets/Kaushalb11/multi-lang-s2t-dataset.kws_testset_zh_s2t
ygyuan/kws_testset_zh_s2t
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
test: 12 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.s2t_TN_clean_numbers
