datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MyMentorLLM-dataset
Dataset Card for MyMentorLLM
This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information).
Dataset Summary
MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.strudel-rl-data
strudel-rl-data
Datasets from the Strudel-RL project (training a Qwen3.6-35B-A3B to write house / techno / hypnotic techno as Strudel programs).
sft/: SFT datasets v1..v5 (messages format, system+user+assistant) with stats.
synth/: GLM-5.3-Flash synthetic rounds 1..5 (prompt, code, reasoning length).
captions/: GLM captions of programs. prompts/: brief generators and held-out eval briefs (v0..v3; v3 = reference-driven).
gold/: 26 hand-written gold programs + manifest.… See the full description on the dataset page: https://huggingface.co/datasets/amol-derick/strudel-rl-data.Blockchain-Sensitive-Detect-Data
Blockchain-Sensitive-Detect-Data
English README
复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。
项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect
数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data
checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints
数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。
该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。
内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.sdf_dataset_en
SpeechDialogueFactory Dataset
Background
This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_en.mixed_shona_datasetkittech_shona_datasetsdf_dataset_zh
SpeechDialogueFactory Dataset
Background
This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_zh.balanced-emotion-dataset-majestrino-withtemporal-detailed-captions
Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions
An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal.
Overview
Total samples: 482,594
Samples per emotion category: 12,997
Number of emotion categories: 40
Format: WebDataset (tar files with FLAC audio + JSON metadata)
Number of tar files: 483
Samples per tar: ~1000
Balancing Strategy
Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression)
音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。
本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、
漫画的な表現生成を目的としたマルチモーダルデータです。
📌 Dataset Summary
本データセットは以下のパイプラインから生成されています:
Audio
↓
Audio Features (04_features.json)
↓
Audio Events (05_audio_events.json)
↓
Space Judgement (06_space_judgement.json)
↓
Scene Interpretation (07_scene_interpretation.json)
↓
Onomatopoeia (08_onomatopoeia.json)
👉 音 → 空間 → 情景 → オノマトペ
という段階的生成構造を持ちます。
📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.
