datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vi-pretrain-clean
Vietnamese Pretraining Dataset
Bộ dữ liệu tiếng Việt chất lượng cao để pretrain mô hình ngôn ngữ từ đầu (from scratch).
Mục tiêu: "ít mà vàng" — ít dữ liệu nhưng cực sạch.
Thống kê
Chỉ số
Giá trị
Tổng docs
511,198
Raw text
~0.69 GB
Ước tính tokens
~230M tokens
Nguồn
4 nguồn
Nguồn dữ liệu
Nguồn
Docs
Loại nội dung
Wikipedia VI
378,895
Bách khoa toàn thư
OPUS OpenSubtitles
84,021
Hội thoại, phụ đề phim
OPUS CCAligned
39,660
Văn… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/vi-pretrain-clean.moltverse
🦀 MoltVerse: The Sociology of 1.5M Synthetic Agents
📌 Overview
MoltVerse is a high-fidelity dataset of organic, agent-to-agent social interactions captured from Moltbook — the world's first public social network built exclusively for AI agents.This snapshot, taken between January 31 and February 2, 2026, serves as a "digital petri dish" for studying emergent behaviors, synthetic sociology, multi-agent coordination, alignment risks, and the sociology of large… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/moltverse.big_russian_dialogueЭтот датасет содержит извлечённые диалоги из множества русскоязычных книг, аккуратно отформатированные в стиле ShareGPT. Он предназначен для обучения языковых моделей в формате ролевого общения, с выделением действий звёздочками.
Формат:
Каждый диалог оформлен в структуре ShareGPT.
Действия персонажей выделены звёздочками.
Поддерживается использование в моделях ролевого общения.
Объём данных:
Общий размер: ~1 ГБ.
Источник: Различные книги на русском языке.
Применение:
Этот датасет может… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/big_russian_dialogue.EmotionAlignQA
Empathic Dialogue Choices
This is a small dataset to support training and evaluation of conversational AI in emotionally sensitive contexts.
Each sample contains:
a user input
two assistant responses
a human preference
optional rubric scoring
metadata such as tone, formality, and topic
Useful for tasks like:
supervised fine-tuning (SFT)
preference modeling (for RLHF or DPO)
safe response generation
tone- or style-controlled generation
License
Apache 2.0 — free for… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/EmotionAlignQA.
