CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01phenixace /S2-TOMG-Bench S^2-Bench Dataset (TOMG) (full version, 45k entries) Official Huggingface Datasets for S^2-Bench: "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation" Please refer to our Github Repo for more usage and useful information. Configurations Each configuration represents a different task: MolCustom_AtomNum: Molecular customized generation by atom number MolCustom_BondNum: Molecular customized generation by bond number… See the full description on the dataset page: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench.tabular10K<n<100K1 likes309 downloads8mo agoHugging Face02Mawube /s2tt-yoruba-englishaudio10K<n<100K0 likes264 downloads6mo agoHugging Face03phenixace /S2-TOMG-Bench-mini S^2-Bench Dataset (TOMG) (mini version, 4.5k entries) Official Huggingface Datasets for S^2-Bench: "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation" Please refer to our Github Repo for more usage and useful information. This mini version offers a choice for researchers with limited resources. Configurations Each configuration represents a different task: MolCustom_AtomNum: Molecular customized generation by atom number… See the full description on the dataset page: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench-mini.tabular1K<n<10K2 likes238 downloads8mo agoHugging Face04Mawube /s2tt-igbo-englishaudio10K<n<100K0 likes214 downloads6mo agoHugging Face05Mawube /s2tt-hausa-englishaudio10K<n<100K0 likes174 downloads6mo agoHugging Face06yangxue /S2TLDhttps://github.com/Thinklab-SJTU/S2TLD image0 likes157 downloads2y agoHugging Face07NgQuocThai /S2T_Split_NoRom_phase2audio10K<n<100K0 likes123 downloads11mo agoHugging Face08PThi35 /S2T_Korean_Merge_2_fixed4audio10K<n<100K0 likes116 downloads5mo agoHugging Face09thevan2404 /S2T_English_Vietnamese_AuVi_2audio1K<n<10K0 likes68 downloads1y agoHugging Face10Zidu-Wang /S2TD-Face Data Card for S2TD-Face This repository provides the data used in ACM MM 2024 paper S2TD-Face. Please see our github repository for details. Citation If you use our work in your research, please cite our publication: @inproceedings{wang2024s2td, title={S2TD-Face: Reconstruct a Detailed 3D Face with Controllable Texture from a Single Sketch}, author={Wang, Zidu and Zhu, Xiangyu and Yu, Jiang and Zhang, Tianshuo and Lei, Zhen}, booktitle={Proceedings of the 32nd ACM… See the full description on the dataset page: https://huggingface.co/datasets/Zidu-Wang/S2TD-Face.image1K<n<10K1 likes60 downloads2y agoHugging Face11gttsehu /Albayzin-2024-BBS-S2T-eval Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge - Evaluation dataset see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge. This is the evaluation data for the challenge. The database consists of a single split: eval : 12498 audio segments How to download this database 1 - If you can handle yourself comfortably with Huggingface Datasets: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T-eval.automatic-speech-recognition0 likes42 downloads2y agoHugging Face12trongnt /s2t-augmented-data Dataset Card for "s2t-augmented-data" More Information needed text1K<n<10K0 likes39 downloads3y agoHugging Face13PThi35 /S2T_Korean_limit_silent_spaceaudio10K<n<100K0 likes36 downloads5mo agoHugging Face14gttsehu /Albayzin-2024-BBS-S2T Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge. NOTE: Test data will be released on September 2nd, 2024. The Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge training and tuning set is based on the gttsehu/basque_parliament_1 dataset. The database consists of four splits: train : 749945 audio segments (automatically extracted) train_clean : 661871 audio segments… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T.documentn<1K0 likes30 downloads2y agoHugging Face15japanese-asr /en2ja.s2t_translationaudio10K<n<100K2 likes30 downloads2y agoHugging Face16PThi35 /S2T_Korean_Merge_2_fixed2audio10K<n<100K0 likes27 downloads5mo agoHugging Face17japanese-asr /ja2en.s2t_translationaudio1K<n<10K1 likes20 downloads2y agoHugging Face18PThi35 /S2T_Korean_3s_silentaudio10K<n<100K0 likes17 downloads3mo agoHugging Face19oscowlai /S2TTaudio10K<n<100K0 likes16 downloads5mo agoHugging Face20voidful /ruozhiba_s2t Dataset Card for "ruozhiba_s2t" More Information needed text1K<n<10K3 likes14 downloads2y agoHugging Face21yxdu /srt-demo-s2tt-70text1K<n<10K0 likes13 downloads11mo agoHugging Face22thevan2404 /S2T_English_Vietnamese_AuVi_testaudion<1K0 likes11 downloads1y agoHugging Face23PThi35 /S2T_Korean0 likes10 downloads6mo agoHugging Face24yxdu /fleurs_eng_test_s2tttext10K<n<100K0 likes8 downloads7mo agoHugging Face25PThi35 /S2T_Korean_Mergeaudio10K<n<100K0 likes8 downloads5mo agoHugging Face26PThi35 /S2T_Korean_Merge_20 likes8 downloads5mo agoHugging Face27kneth90 /s2tkp_dataset_part_3audion<1K0 likes6 downloads2y agoHugging Face28Kaushalb11 /multi-lang-s2t-dataset Multi-Language Audio Dataset A high-quality, cleaned audio dataset derived from five AI4Bharat sources, containing speech-to-text data in Hindi, Gujarati, and Telugu with English translations. Dataset Overview Language Duration (Hours) Hindi 55.45 Gujarati 42.12 Telugu 37.83 Total Duration: ~135 hours of cleaned, high-quality audio Source Datasets Data collected and cleaned from: ai4bharat/IndicVoices-ST ai4bharat/NPTEL… See the full description on the dataset page: https://huggingface.co/datasets/Kaushalb11/multi-lang-s2t-dataset.audion<1K0 likes5 downloads9mo agoHugging Face29ygyuan /kws_testset_zh_s2tgated ygyuan/kws_testset_zh_s2t Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample. Layout data/ <split>/ metadata.csv audio/ <split>-000.tar <split>-001.tar ... Shard counts: test: 12 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/kws_testset_zh_s2t.audioaudio-classification0 likes5 downloads2mo agoHugging Face30Miroo222 /s2t_TN_clean_numbersgated0 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.