datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vibevoice-quran_persian-single-speakervibevoice-gptinformal_persian-single-speakerSLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.SpeakerCard-1M
SpeakerCard-1M (VoxCeleb release)
Evidence-grounded Speaker Card corpus for in-the-wild speaker verification.
SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/JYP2024/SpeakerCard-1M.DialogSum_with_speaker
DialogSum_with_speaker — Historical Research Data
No longer actively maintained. Retained for reproducibility and reference.
Post-processed DialogSum data first released and used in InstructDS (EMNLP 2023). See the InstructDS model and project code for the associated research.
Part of the Research Archive.
Original dataset description
This one is a post-processed DialogSum dataset (https://aclanthology.org/2021.findings-acl.449/)… See the full description on the dataset page: https://huggingface.co/datasets/binwang/DialogSum_with_speaker.speaker-embeddingsHaruhi-Dialogue-Speaker-Extract
Chat凉宫春日的对话抽取模型
我们希望有一个模型能够从小说的chunk中批量去提取摘要和对话
这个模型就是实现了这一点。模型使用了大约30k的中文小说数据和20k的英文小说数据进行训练,在qwen-1.8上进行了3个epoch的finetune。 原则上模型同时支持中文和英文小说的训练
主项目链接 https://github.com/LC1332/Chat-Haruhi-Suzumiya
李鲁鲁完成了数据的收集,以及进一步将inference程序扩展到连续的chunks
刘崇寒完成了模型的训练
米唯实测试并上传模型到hugging face
Chat Haruhi Suzumiya's Dialogue Extraction Model
We hope to have a model that can extract summaries and dialogues in batches from chunks of novels.
This model achieves just that. It was trained… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Dialogue-Speaker-Extract.whispbook-functiongemma-speaker-attribution-sftspeakerlm-datasetHaruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了
这一次把总结和抽取拆分成了两个任务
并且抽取的格式改为了csv格式。
all15_speaker_deduped_tts_train_clone_pairs_raw
All-15 Speaker-Deduped TTS Train Clone Pairs Raw
This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline.
Contents
data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows.
data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers
sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 364
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task909-dialogre-prevalent-speakers.sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers
sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 753
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task909-dialogre-prevalent-speakers.4chan_archive_ShareGPT_only5score_speakerCount_only2speakersspeakercard-1m-anonymous
SpeakerCard-1M (VoxCeleb release)
Evidence-grounded Speaker Card corpus for in-the-wild speaker verification.
SpeakerCard-1M is a speaker-centric resource built on VoxCeleb1/2 under a tool-first, LLM-last pipeline: ten acoustic probes extract field-level evidence, a schema separates relatively stable traits (gender, age band, accent, pitch band, timbre, language) from utterance-level states (emotion, channel, environment, speaking rate), and a constrained LLM verbalizes the… See the full description on the dataset page: https://huggingface.co/datasets/dblind/speakercard-1m-anonymous.4chan_archive_ShareGPT_only5_speakerCount
