CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dorsaasgari /vibevoice-quran_persian-single-speakeraudio1K<n<10K1 likes4.9k downloads12d agoHugging Face02Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.2k downloads12d agoHugging Face03GenSEC-LLM /SLT-Task2-Post-ASR-Speaker-Tagging Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization) Description This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system. Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging. SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.tabular10K<n<100K2 likes805 downloads2y agoHugging Face04binwang /DialogSum_with_speaker DialogSum_with_speaker — Historical Research Data No longer actively maintained. Retained for reproducibility and reference. Post-processed DialogSum data first released and used in InstructDS (EMNLP 2023). See the InstructDS model and project code for the associated research. Part of the Research Archive. Original dataset description This one is a post-processed DialogSum dataset (https://aclanthology.org/2021.findings-acl.449/)… See the full description on the dataset page: https://huggingface.co/datasets/binwang/DialogSum_with_speaker.text10K<n<100K0 likes77 downloads20d agoHugging Face05JYP2024 /SpeakerCard-1M SpeakerCard-1M (VoxCeleb release) Evidence-grounded Speaker Card corpus for in-the-wild speaker verification. SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/JYP2024/SpeakerCard-1M.text100K<n<1M3 likes76 downloads4mo agoHugging Face06kacperwikiel /speakleash-tokenizer-5gb-sample SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests. texttext-generation1M<n<10M0 likes75 downloads3mo agoHugging Face07qwertyuiopasdfg /IELTs-Speaking-answer Overview This dataset consists of 2 json files named 'ielts_new.json' and 'ielts_old.json', which contain ielts questions and its corresponding answers for part 1 and part 2. 'ielts_new.json': new IELTs topics for 2024 September-December. 'ielts_old.json': remained IELTs topics for 2024 September-December. Quality Since the dataset is analysed and generated by ChatGPT based on my own pdf file, the answer may be incomplete(only part of the sentence is extracted, leading to… See the full description on the dataset page: https://huggingface.co/datasets/qwertyuiopasdfg/IELTs-Speaking-answer.texttext-generationn<1K4 likes70 downloads2y agoHugging Face08open-llm-leaderboard /speakleash__Bielik-11B-v2.3-Instruct-detailsgated Dataset Card for Evaluation run of speakleash/Bielik-11B-v2.3-Instruct Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2.3-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2.3-Instruct-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face09SpeakoFlow /dictation-cleanup-examples Dictation cleanup examples A sample of the hand-written cases behind SpeakoFlow Mini, published so the conventions the model follows are inspectable rather than described. Seven cases in each of fifteen categories, spread across short, medium and long transcripts. Every case was written by hand. None of it is captured speech. This is not a benchmark Read that before using it for anything. These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.texttext-generationn<1K0 likes54 downloads27d agoHugging Face10MikhailT /speaker-embeddingstabular10K<n<100K0 likes36 downloads3y agoHugging Face11silk-road /Haruhi-Dialogue-Speaker-Extract Chat凉宫春日的对话抽取模型 我们希望有一个模型能够从小说的chunk中批量去提取摘要和对话 这个模型就是实现了这一点。模型使用了大约30k的中文小说数据和20k的英文小说数据进行训练,在qwen-1.8上进行了3个epoch的finetune。 原则上模型同时支持中文和英文小说的训练 主项目链接 https://github.com/LC1332/Chat-Haruhi-Suzumiya 李鲁鲁完成了数据的收集,以及进一步将inference程序扩展到连续的chunks 刘崇寒完成了模型的训练 米唯实测试并上传模型到hugging face Chat Haruhi Suzumiya's Dialogue Extraction Model We hope to have a model that can extract summaries and dialogues in batches from chunks of novels. This model achieves just that. It was trained… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Dialogue-Speaker-Extract.text10K<n<100K5 likes24 downloads3y agoHugging Face12chakibmed /whispbook-functiongemma-speaker-attribution-sfttext100K<n<1M0 likes24 downloads5mo agoHugging Face13open-llm-leaderboard /speakleash__Bielik-11B-v2-detailsgated Dataset Card for Evaluation run of speakleash/Bielik-11B-v2 Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face14open-llm-leaderboard /speakleash__Bielik-11B-v2.0-Instruct-detailsgated Dataset Card for Evaluation run of speakleash/Bielik-11B-v2.0-Instruct Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2.0-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2.0-Instruct-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face15SakhrML /SpeakMK1_SLP_Dialogue Dataset Card for SLP Dialogue Dataset Dataset Description This dataset consists of 1,000 multi-turn simulated pediatric Speech-Language Pathology (SLP) interaction dialogues. It is specifically designed to train, evaluate, or fine-tune LLMs to act as clinical speech-language therapists or to study clinical reasoning during speech therapy sessions. Each conversation includes clinical metadata (child age, speech sound disorder category, specific phone error, clinical goal… See the full description on the dataset page: https://huggingface.co/datasets/SakhrML/SpeakMK1_SLP_Dialogue.texttext-generation1K<n<10K1 likes17 downloads4mo agoHugging Face16tmtanu /Speak_telugutabulartranslationn<1K0 likes15 downloads4mo agoHugging Face17KafeisM /pirate-speak-dataset Pirate English Style Transfer Dataset Dataset Summary This dataset contains 500 parallel sentence pairs where each item includes: Modern English (english) Stereotypical Pirate English (pirate) It is designed for style transfer tasks, especially training text-to-text models to rewrite sentences into pirate-style English while preserving the core meaning. The dataset mixes many categories of text: Everyday greetings Questions and requests Complaints and opinions… See the full description on the dataset page: https://huggingface.co/datasets/KafeisM/pirate-speak-dataset.texttext-generationn<1K1 likes13 downloads10mo agoHugging Face18silk-road /Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了 这一次把总结和抽取拆分成了两个任务 并且抽取的格式改为了csv格式。 texttext-generationn<1K0 likes12 downloads3y agoHugging Face19BeautyyuYanli /speakerlm-datasettextn<1K1 likes12 downloads1y agoHugging Face20tizerk /trump-speaktextn<1K0 likes10 downloads2y agoHugging Face21beldua /english-public-speaking-skills-30textn<1K1 likes9 downloads9mo agoHugging Face22instinct-org /all15_speaker_deduped_tts_train_clone_pairs_rawgated All-15 Speaker-Deduped TTS Train Clone Pairs Raw This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline. Contents data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows. data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.tabulartext-to-speech10K<n<100K0 likes9 downloads4mo agoHugging Face23schneiderkamplab /sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 364 Task: synthetic anonymous instruction replacement Generation Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers.textn<1K0 likes8 downloads3mo agoHugging Face24schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 753 Task: synthetic anonymous instruction replacement Generation Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers.textn<1K0 likes7 downloads3mo agoHugging Face25javlonDev /speakingtextn<1K0 likes6 downloads3y agoHugging Face26adamo1139 /4chan_archive_ShareGPT_only5score_speakerCount_only2speakerstabular100K<n<1M0 likes5 downloads2y agoHugging Face27dblind /speakercard-1m-anonymous SpeakerCard-1M (VoxCeleb release) Evidence-grounded Speaker Card corpus for in-the-wild speaker verification. SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/dblind/speakercard-1m-anonymous.text100K<n<1M0 likes5 downloads4mo agoHugging Face28HCHSmost /speaking-ratingtextn<1K0 likes3 downloads2y agoHugging Face29adamo1139 /4chan_archive_ShareGPT_only5_speakerCounttabular100K<n<1M0 likes2 downloads2y agoHugging Face30SpeakEZ-AI /speakez-rewrite-v1textn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.