datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vibevoice-quran_persian-single-speakervibevoice-gptinformal_persian-single-speakerSLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.DialogSum_with_speaker
DialogSum_with_speaker — Historical Research Data
No longer actively maintained. Retained for reproducibility and reference.
Post-processed DialogSum data first released and used in InstructDS (EMNLP 2023). See the InstructDS model and project code for the associated research.
Part of the Research Archive.
Original dataset description
This one is a post-processed DialogSum dataset (https://aclanthology.org/2021.findings-acl.449/)… See the full description on the dataset page: https://huggingface.co/datasets/binwang/DialogSum_with_speaker.SpeakerCard-1M
SpeakerCard-1M (VoxCeleb release)
Evidence-grounded Speaker Card corpus for in-the-wild speaker verification.
SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/JYP2024/SpeakerCard-1M.speakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup.
Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
IELTs-Speaking-answer
Overview
This dataset consists of 2 json files named 'ielts_new.json' and 'ielts_old.json', which contain ielts questions and its corresponding answers for part 1 and part 2.
'ielts_new.json': new IELTs topics for 2024 September-December.
'ielts_old.json': remained IELTs topics for 2024 September-December.
Quality
Since the dataset is analysed and generated by ChatGPT based on my own pdf file, the answer may be incomplete(only part of the sentence is extracted, leading to… See the full description on the dataset page: https://huggingface.co/datasets/qwertyuiopasdfg/IELTs-Speaking-answer.speakleash__Bielik-11B-v2.3-Instruct-details
Dataset Card for Evaluation run of speakleash/Bielik-11B-v2.3-Instruct
Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2.3-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2.3-Instruct-details.dictation-cleanup-examples
Dictation cleanup examples
A sample of the hand-written cases behind
SpeakoFlow Mini, published so the
conventions the model follows are inspectable rather than described.
Seven cases in each of fifteen categories, spread across short, medium and long transcripts.
Every case was written by hand. None of it is captured speech.
This is not a benchmark
Read that before using it for anything.
These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.speaker-embeddingsHaruhi-Dialogue-Speaker-Extract
Chat凉宫春日的对话抽取模型
我们希望有一个模型能够从小说的chunk中批量去提取摘要和对话
这个模型就是实现了这一点。模型使用了大约30k的中文小说数据和20k的英文小说数据进行训练,在qwen-1.8上进行了3个epoch的finetune。 原则上模型同时支持中文和英文小说的训练
主项目链接 https://github.com/LC1332/Chat-Haruhi-Suzumiya
李鲁鲁完成了数据的收集,以及进一步将inference程序扩展到连续的chunks
刘崇寒完成了模型的训练
米唯实测试并上传模型到hugging face
Chat Haruhi Suzumiya's Dialogue Extraction Model
We hope to have a model that can extract summaries and dialogues in batches from chunks of novels.
This model achieves just that. It was trained… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Dialogue-Speaker-Extract.whispbook-functiongemma-speaker-attribution-sftspeakleash__Bielik-11B-v2-details
Dataset Card for Evaluation run of speakleash/Bielik-11B-v2
Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2-details.speakleash__Bielik-11B-v2.0-Instruct-details
Dataset Card for Evaluation run of speakleash/Bielik-11B-v2.0-Instruct
Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2.0-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2.0-Instruct-details.SpeakMK1_SLP_Dialogue
Dataset Card for SLP Dialogue Dataset
Dataset Description
This dataset consists of 1,000 multi-turn simulated pediatric Speech-Language Pathology (SLP) interaction dialogues. It is specifically designed to train, evaluate, or fine-tune LLMs to act as clinical speech-language therapists or to study clinical reasoning during speech therapy sessions. Each conversation includes clinical metadata (child age, speech sound disorder category, specific phone error, clinical goal… See the full description on the dataset page: https://huggingface.co/datasets/SakhrML/SpeakMK1_SLP_Dialogue.Speak_telugupirate-speak-dataset
Pirate English Style Transfer Dataset
Dataset Summary
This dataset contains 500 parallel sentence pairs where each item includes:
Modern English (english)
Stereotypical Pirate English (pirate)
It is designed for style transfer tasks, especially training text-to-text models to rewrite sentences into pirate-style English while preserving the core meaning.
The dataset mixes many categories of text:
Everyday greetings
Questions and requests
Complaints and opinions… See the full description on the dataset page: https://huggingface.co/datasets/KafeisM/pirate-speak-dataset.Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了
这一次把总结和抽取拆分成了两个任务
并且抽取的格式改为了csv格式。
speakerlm-datasettrump-speakenglish-public-speaking-skills-30all15_speaker_deduped_tts_train_clone_pairs_raw
All-15 Speaker-Deduped TTS Train Clone Pairs Raw
This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline.
Contents
data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows.
data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers
sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 364
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers.sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers
sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 753
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers.speaking4chan_archive_ShareGPT_only5score_speakerCount_only2speakersspeakercard-1m-anonymous
SpeakerCard-1M (VoxCeleb release)
Evidence-grounded Speaker Card corpus for in-the-wild speaker verification.
SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/dblind/speakercard-1m-anonymous.speaking-rating4chan_archive_ShareGPT_only5_speakerCountspeakez-rewrite-v1
