datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-av-responses-llama-70b-layer53Qwen3.5-0.8B-responsesQwen3.5-4B-responsesmbti-f-t-style-responses#mbti-f-t-style-response
이 데이터는 다양한 감정적인 발화에 대해 MBTI의 F와 T 스타일대로 한 발화를 수집한 데이터입니다.
감정적인 발화의 경우 AIHub의 감성 대화 말뭉치 중 첫번째 발화를 사용했습니다.
그리고 F와 T 스타일의 발화는 언어 모델을 통해 생성했습니다.
moral-dilemma-responses
Moral Dilemma Responses Dataset
17,290 natural language responses to moral dilemmas from princi/pal, a Tamagotchi-like game where players guide a virtual pet through ethical decisions.
Presented at NeurIPS 2025 Creative AI track.
What is this?
Players advise a virtual pet on moral dilemmas ranging from "Should I pick up trash?" to "Should you lie in court to defend a friend?". The pet evolves based on the guidance and eventually makes autonomous moral decisions.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cnnmon/moral-dilemma-responses.Qwen3.8-2.4T-A95B-responses-original10k
Qwen3.8-2.4T-A95B responses — original aligned 10k
The first 10,000 exact rows from the private source dataset inference-optimization/Qwen3.8-2.4T-A95B-responses. Records are preserved without modification.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
VideoAutoArena_model_responsesMoS-Qwen3-8B-EAGLE3-responses
MoS — Qwen3-8B EAGLE3 Training Responses
Target-model responses for training EAGLE3 speculative-decoding draft models against
Qwen/Qwen3-8B. Built for the MoS (Mixture of
Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft
EAGLE3 / SpecForge training run on Qwen3-8B.
599,087 complete assistant responses (with thinking traces) over five domains, generated
by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.Qwen3.8-27B-responses-regenerated10k
Qwen3.8-27B regenerated responses — aligned 10k
The same 10,000 original prompts regenerated with dense Qwen/Qwen3.8-27B. Original prompt strings were used directly; they were never reconstructed by detokenization.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
mmlu_responsesQwen3.5-9B-responsesTaiwanGovQA-CN-Responses
本資料集由OllaForge生成
TaiwanGovQA-CN-Responses
Dataset Summary
TaiwanGovQA-CN-Responses 是一份「台灣政府 QA」資料集的擴充版本,新增欄位 answer_zh_cn,提供以簡體中文(中國大陸用語)撰寫的回答,方便進行簡中生成式問答訓練、繁簡變體遷移與跨域語言風格研究。本資料來源於台灣政府開放資料中的 1996QA 以及各縣市政府&機關局處的 QA 資料集。
Split:train
規模:約 5K 筆
格式:JSON / JSONL
授權:Apache-2.0
Supported Tasks
Generative QA(簡中作答):question -> answer_zh_cn
Variant transfer(繁轉簡回答):answer (繁中) -> answer_zh_cn (簡中)
Domain adaptation(政府政策/行政流程 QA):面向公部門/法規/行政程序相關問題的回答生成… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/TaiwanGovQA-CN-Responses.Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it mean to… See the full description on the dataset page: https://huggingface.co/datasets/vericudebuget/Bible-responses-dataset-gotquestions.multi-ai-interpretive-responses
Multi-AI Interpretive Responses
arena.ai のサイドバイサイド / ダイレクトバトルで行った
日本語チャットセッションのアーカイブ。
同じ問い(哲学・倫理・サブカル・メタ認知ネタ)に対する複数 LLM の
解釈・応答差を観察するためのデータセット。
「楽しい を教えるお仕事ならしまーす」
倫理と哲学だけメタ超級。バシャール可。仏陀可。ウィトゲン可。サブカル可。
ファイル
ファイル
説明
data.jsonl
1 行 = 1 セッション。HF Datasets Viewer はこれを読みます。
timeline.md
人間用:会話開始時刻順の年表(タイトル・モデル数・所要時間付き)。
filename_map.csv
新ファイル名 ↔ 元日本語タイトルの対応表。
files/chat_XXXX.json
arena.ai 由来の生 JSON(元構造そのまま)。連番は会話開始時刻順。
スキーマ(data.jsonl の… See the full description on the dataset page: https://huggingface.co/datasets/TonbokiriRaikiriMuramasa/multi-ai-interpretive-responses.laguna-xs-ultrachat-responsescai-eval-responsesunnaturalhermes-responses-30kResponses to the questions in unnaturalhermes-questions-30k from the following models at temperature=0:
mistralai/Mixtral-8x7B-Instruct-v0.1
teknium/OpenHermes-2p5-Mistral-7B
mistralai/Mistral-7B-Instruct-v0.2
togethercomputer/StripedHyena-Nous-7B
mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.railvaani-study-responsesR1-Responses-1K-Allgpt4_do_not_answer_responsesethical-responsesllm-viz-responsesDeepSeek-V4-Flash-responseslaguna-xs-magpie-300k-responsesGemma4-Responses-Nemotronsumthink_responses_summarizedSee: https://huggingface.co/datasets/G-reen/sumthink
This contains the summarized GLM 4.7 nvfp4 thinking traces as well as the source data.
anthropomorphism-short-responsesage-based-responsesllava_ov_7b_responses
