datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-dialog-tsundere
Dataset Sintetis Dialog Persona Tsundere Indonesia
Deskripsi Dataset
Dataset Sintetis Dialog Persona Tsundere Indonesia merupakan kumpulan dialog berbasis teks dalam Bahasa Indonesia yang dikembangkan untuk merepresentasikan pola komunikasi dengan sifat persona tsundere.
Dataset ini dikembangkan sebagai bagian dari penelitian tugas akhir berjudul:
“Desain Asisten Mobile yang Proaktif dengan Pengembangan Dataset Sifat Persona Tsundere.”
Dataset tidak digunakan… See the full description on the dataset page: https://huggingface.co/datasets/vnsd13/dataset-dialog-tsundere.Tsuki-dataset
Token Compression Training Dataset
15,000 bilingual compression pairs. Prompts, markdown, and technical instructions.
Trains models to reduce LLM API costs without losing critical information.
Real compression pairs.
Verbose prompts → minimum tokens.
Markdown sections → essential content.
Technical instructions → direct commands.
Reasoning examples included.
Knows when not to compress.
Legal text preserved intact.
Medical instructions… See the full description on the dataset page: https://huggingface.co/datasets/tsuki-team/Tsuki-dataset.Translation-Sample-Dataset
TsukiOwO/Translation-Sample-Dataset
Data Source
This dataset is a portion of open-r1/OpenR1-Math-220k.
Purpose
Serves as sample data for translation text.
bible-en-th
bible-en-th
Overview
The bible-en-th dataset is a bilingual corpus containing English and Thai translations of the Bible, specifically the King James Version (KJV) translated into Thai. This dataset is designed for various natural language processing tasks, including translation, language modeling, and text analysis.
Languages: English (en) and Thai (th)
Total Rows: 31,102
Dataset Structure
The dataset consists of two main features:
en: English text from… See the full description on the dataset page: https://huggingface.co/datasets/Tsunnami/bible-en-th.HAMLET
HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics
Official dataset for HAMLET
We proposed a multi-agent framework HAMLET that decouples offline planning and online performance in AI theatrics scenarios. HAMLET excels in creating expressive, coherent, real-time and physically interactive drama experiences in a fully autonomous manner.
English | 简体中文
📖 Overview
🎉 News
[2026.03.19]… See the full description on the dataset page: https://huggingface.co/datasets/Tsumugii/HAMLET.FURINA
