datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-dialog-tsundere
Dataset Sintetis Dialog Persona Tsundere Indonesia
Deskripsi Dataset
Dataset Sintetis Dialog Persona Tsundere Indonesia merupakan kumpulan dialog berbasis teks dalam Bahasa Indonesia yang dikembangkan untuk merepresentasikan pola komunikasi dengan sifat persona tsundere.
Dataset ini dikembangkan sebagai bagian dari penelitian tugas akhir berjudul:
“Desain Asisten Mobile yang Proaktif dengan Pengembangan Dataset Sifat Persona Tsundere.”
Dataset tidak digunakan… See the full description on the dataset page: https://huggingface.co/datasets/vnsd13/dataset-dialog-tsundere.Translation-Sample-Dataset
TsukiOwO/Translation-Sample-Dataset
Data Source
This dataset is a portion of open-r1/OpenR1-Math-220k.
Purpose
Serves as sample data for translation text.
bible-en-th
bible-en-th
Overview
The bible-en-th dataset is a bilingual corpus containing English and Thai translations of the Bible, specifically the King James Version (KJV) translated into Thai. This dataset is designed for various natural language processing tasks, including translation, language modeling, and text analysis.
Languages: English (en) and Thai (th)
Total Rows: 31,102
Dataset Structure
The dataset consists of two main features:
en: English text from… See the full description on the dataset page: https://huggingface.co/datasets/Tsunnami/bible-en-th.
