datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
honkai_impact_3rd_chinese_dialogue_corpus
崩坏三游戏剧情语料
总计 92,421 句剧情对白(带有角色标签)+旁白,从崩坏3的“主线1黄昏、少女、战舰”到“主线第二部03间章:一个梦游者的苦痛”
本数据集从 honkai_impact_3rd_game_playthrough 视频数据集出发,经过 AI pipeline 最终获取结构化的文本剧情语料。
AI pipeline 概述如下:
分P下载视频(使用 BBDown 下载 BiliBili崩三剧情视频)
视频帧分割(每1秒取一帧画面)
逐帧 OCR 检测文本(使用 Paddle-OCR)
逐帧 VLM 结构化解析(使用 MiniCPM-V-2_6,输入为帧图像 + OCR结果,输出为结构化 JSON)
基于规则的后处理
规范化 VLM 输出(e.g., 去噪、排除格式有问题的输出)
中间帧的信息去重与归并(e.g.… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/honkai_impact_3rd_chinese_dialogue_corpus.ubuntu_dialogue_corpus_trainubuntu_dialogue_corpus_devtestmulti-relational-multi-party-chat-corpusMulti-Relational Multi-Party Chat Corpus (MRMP): Japanese text-based chats comprising first-time-meeting dialogues and family-included dialoguesAMRIT-Punjabi-Clinical-Dialogue-Corpus
ੴ AMRIT Punjabi Clinical Dialogue & Medical Diagnosis Corpus
☬ ਅੰਮ੍ਰਿਤ ਪੰਜਾਬੀ ਕਲੀਨਿਕਲ ਸੰਵਾਦ ਅਤੇ ਡਾਕਟਰੀ ਨਿਦਾਨ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Medical AI Architecture
Lead Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
Mission: Free Autonomous AI Doctor for Humanity (ਦੁਨੀਆਂ ਦੇ ਲੋੜਵੰਦ ਲੋਕਾਂ ਲਈ ਮੁਫ਼ਤ AI ਡਾਕਟਰ)
Flagship Platform: AMRIT Research OS (100% Local Medical Intelligence)
📖 Dataset Overview / ਸੰਖੇਪ
The… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/AMRIT-Punjabi-Clinical-Dialogue-Corpus.Free_Dialogue_in_Saudi_Arabia_Corpus
Description
This dataset covers multiple scenarios such as banking, healthcare, insurance, sales, telecom, travel. The speakers are gender evenly, and each set of the audio is approximately 0.5 hour.
For more details, please refer to the link: https://dataoceanai.com/datasets/asr/free-dialogue-in-saudi-arabia-corpus/
Specification
ID:
King-ASR-919
SIZE:
113 hours
LANGUAGE:
Saudi Arabia
SPEAKERS:
100
AGES:
18-45 years old
DEVICES:
Mobile
Fully-Clean-Cornell-Movie-Dialogue-CorpusEthiopian_Amharic_Free_Dialogue_Speech_Corpus
SPECIFICATION:
Product Type: Ethiopian Amharic language, free dialogue, mobile 16K 【Corpus Type】 Family, health, travel, education, work, cuisine, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics of daily life. Natural context, applicable to all industries.
Pronunciation Person Information: Gender: Male 50%, Female 50% Age: The pronunciation people mainly cover the age range of 16-45. Accent: The pronunciation people mainly come from… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Ethiopian_Amharic_Free_Dialogue_Speech_Corpus.Serbian_language_free_dialogue_Corpus
SPECIFICATION:
Product Type: Serbian Language, Free Dialogue, Mobile 16K
Corpus Type: Family, health, travel, education, work, cuisine, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics in daily life. Natural context, applicable to all industries.
Pronouncer Information: Gender: Approximately even Age: Pronouncers mainly cover the age range of 16-45 Accent: Pronouncers mainly come from central Serbia.
For more details:… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Serbian_language_free_dialogue_Corpus.real-multilingual-dialogue-corpus
krystv/real-multilingual-dialogue-corpus
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset("krystv/real-multilingual-dialogue-corpus")
Disco_Elysium_Russian_Dialogue_CorpusAlbanian_Free_Dialogue_Speech_Corpus
SPECIFICATION:
Corpus Type: Family, health, travel, education, work, gourmet food, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics of daily life. Natural context, applicable to all industries.
Pronunciation Person Information: Gender: Male 45%, Female 55% Age: The pronunciation people mainly cover the age range of 16-45. Accent: Speakers are from Tirana.
For more details:… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Albanian_Free_Dialogue_Speech_Corpus.multilingual-real-dialogue-corpus-v1
krystv/multilingual-real-dialogue-corpus-v1
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset("krystv/multilingual-real-dialogue-corpus-v1")
Free_dialogue_in_Odia_Speech_Corpus
Product Type:
Odia language from India, free conversation, mobile 16K
For more details, please refer to the link: https://dataoceanai.com/datasets/asr/free-dialogue-in-odia-speech-corpus/
Specification:
ID:
King-ASR-946
Size:
52 hours
Language:
Odia
Corpus Type:
Home, health, travel, education, work, gourmet food, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics in daily life Natural context, applicable to the entire industry… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Free_dialogue_in_Odia_Speech_Corpus.
