CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hoanghai2110 /vi-pretrain-clean Vietnamese Pretraining Dataset Bộ dữ liệu tiếng Việt chất lượng cao để pretrain mô hình ngôn ngữ từ đầu (from scratch). Mục tiêu: "ít mà vàng" — ít dữ liệu nhưng cực sạch. Thống kê Chỉ số Giá trị Tổng docs 511,198 Raw text ~0.69 GB Ước tính tokens ~230M tokens Nguồn 4 nguồn Nguồn dữ liệu Nguồn Docs Loại nội dung Wikipedia VI 378,895 Bách khoa toàn thư OPUS OpenSubtitles 84,021 Hội thoại, phụ đề phim OPUS CCAligned 39,660 Văn… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/vi-pretrain-clean.texttext-generation100K<n<1M0 likes300 downloads6mo agoHugging Face02christian-hoang-04 /moltverse 🦀 MoltVerse: The Sociology of 1.5M Synthetic Agents 📌 Overview MoltVerse is a high-fidelity dataset of organic, agent-to-agent social interactions captured from Moltbook — the world's first public social network built exclusively for AI agents.This snapshot, taken between January 31 and February 2, 2026, serves as a "digital petri dish" for studying emergent behaviors, synthetic sociology, multi-agent coordination, alignment risks, and the sociology of large… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/moltverse.texttext-generation10K<n<100K2 likes85 downloads8mo agoHugging Face03Hoaxer2000 /big_russian_dialogueЭтот датасет содержит извлечённые диалоги из множества русскоязычных книг, аккуратно отформатированные в стиле ShareGPT. Он предназначен для обучения языковых моделей в формате ролевого общения, с выделением действий звёздочками. Формат: Каждый диалог оформлен в структуре ShareGPT. Действия персонажей выделены звёздочками. Поддерживается использование в моделях ролевого общения. Объём данных: Общий размер: ~1 ГБ. Источник: Различные книги на русском языке. Применение: Этот датасет может… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/big_russian_dialogue.texttext-generation100K<n<1M3 likes31 downloads1y agoHugging Face04hoanghai2110 /EmotionAlignQA Empathic Dialogue Choices This is a small dataset to support training and evaluation of conversational AI in emotionally sensitive contexts. Each sample contains: a user input two assistant responses a human preference optional rubric scoring metadata such as tone, formality, and topic Useful for tasks like: supervised fine-tuning (SFT) preference modeling (for RLHF or DPO) safe response generation tone- or style-controlled generation License Apache 2.0 — free for… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/EmotionAlignQA.texttext-generation1K<n<10K1 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.