datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.OpenDialog_English
OpenDialog English
This dataset contains English dialog and conversation data.
Dataset Structure
The dataset is provided in Parquet format with 153 splits for efficient loading.
Data Files
Format: Parquet
Splits: 153 files (train-00001-of-00153.parquet through train-00153-of-00153.parquet)
Total Size: ~72.8 GB
Loading the Dataset
from datasets import load_dataset
# Load the full dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/OpenDialog_English.
