datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PromptDialog_sub
PromptDialog Subset
PromptDialog_sub is a public 1,000-dialogue subset sampled from PromptDialog for non-commercial academic research and education.
Speaker identifiers in dia_manifest.jsonl are anonymized as speaker_1, speaker_2, and so on within each dialogue.
License and Use
PromptDialog_sub is released under CC BY-NC 4.0 together with the additional PromptDialog Data Use Agreement.
The dataset may be used only for non-commercial academic research and education.… See the full description on the dataset page: https://huggingface.co/datasets/alphappp/PromptDialog_sub.colloqialized_prompt
Colloquialized Prompt Dataset
This repository contains prompt and audio variants derived from the 60
WildClawBench tasks, plus the reusable task template. It supports experiments
that compare written prompts, spoken-style rewrites, synthesized speech, raw
ASR transcripts, and normalized ASR transcripts.
Dataset layout
.
├── prompts/ # Instructions used by rewrite/normalization jobs
├── scripts/ # Reproducible data preparation… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/colloqialized_prompt.PromptBased-AmericanEnglishFullDuplexTwoSpeakerConversationalDataset-Sample
American English Full-Duplex Two-Speaker Conversational Dataset — Prompt-Based Sample
This is a free prompt-based preview (~5 hours) of OcularAI's full
American English Full-Duplex Two-Speaker Conversational Dataset. Each
conversation is an open, natural discussion seeded by a single prompt — the
pair was given one topic and talked freely. (For conversations directed by a
designed scenario targeting a specific turn-taking / voice-dynamics behavior,
see the companion… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/PromptBased-AmericanEnglishFullDuplexTwoSpeakerConversationalDataset-Sample.PromptBased-BritishEnglishFullDuplexTwoSpeakerConversationalDataset-Sample
British English Full-Duplex Two-Speaker Conversational Dataset — Prompt-Based Sample
This is a free prompt-based preview of OcularAI's full British English
Full-Duplex Two-Speaker Conversational Dataset. Each conversation is an
open, natural discussion — paired speakers were prompted to talk freely on
a topic of their choice. Audio is high-fidelity FLAC, with each speaker on
an independent, isolated track. (For conversations directed by a designed
scenario targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/PromptBased-BritishEnglishFullDuplexTwoSpeakerConversationalDataset-Sample.jain_architecture
