datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.Arabic-professional-voice
Arabic Professional Voice
A high-quality, single-speaker Arabic Text-to-Speech (TTS) dataset recorded by a professional speaker. All transcriptions include full Tashkeel (diacritical marks), making it directly suitable for training neural TTS systems without additional text normalization.
Dataset Summary
Property
Value
Language
Arabic — Modern Standard Arabic (MSA)
Utterances
439
Speaker
1 (professional male speaker)
Sampling Rate
16 kHz
Format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Arabic-professional-voice.tts-voices-sampler-professional
TTS Voices Sampler
A small sample set of audio voices, with transcriptions, from other datasets.
Sources
EmoV-DB
Link. EmoV-DB is a dataset with a forced
emotional reading of an arbitrary text. Multiple emotions are supplied with the
same speaker. There are only four speakers in the dataset. All have a U.S. accent.
Unique non-commercial license https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md
Emilia
Link.
Dataset which seems very… See the full description on the dataset page: https://huggingface.co/datasets/nick-mccormick/tts-voices-sampler-professional.professional-interview
