datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyday-conversations-tur
Everyday Turkish Conversations
This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k.
License
This dataset is released under the Apache 2.0 License.
soar_arc_train_5M
SOAR-ARC Models: Self-Improving Language Models for Program Synthesis
🤗 Hugging Face (data and model) | 📑 Paper | 📑 Blog | 💻 Code
This repository contains around 5 million ARC solutions. For solutions that successfully solve an original ARC task, we deduplicate entries by their code to ensure uniqueness. For solutions that correspond to new synthetic tasks generated via hindsight relabeling, we deduplicate based on their output results. This… See the full description on the dataset page: https://huggingface.co/datasets/julien31/soar_arc_train_5M.soal-ujian-sd-kurikulum-merdeka
Soal Ujian SD Kurikulum Merdeka (Kelas 1–6)
Dataset soal ujian Sekolah Dasar (SD) berbahasa Indonesia, dibangun mengikuti
Kurikulum Merdeka untuk Kelas 1 sampai Kelas 6 (Fase A, B, dan C).
Setiap entri berupa satu lembar soal (worksheet) berisi 5 soal dengan komposisi
sesuai jenis ujiannya.
3.000 lembar soal (worksheets)
15.000 soal total
5 mata pelajaran × 6 kelas × 5 jenis ujian, terdistribusi merata
Struktur Dataset
Dataset punya dua konfigurasi:… See the full description on the dataset page: https://huggingface.co/datasets/erzanugroho/soal-ujian-sd-kurikulum-merdeka.r1-reasoning-tr
R1 Reasoning TR
This is an R1 reasoning dataset translated into Turkish, containing conversations between users and assistants. Thanks to lightblue for the dataset.
License
This dataset is released under the Apache 2.0 License.
turoqa-small
TUROQA - Turkish Open QA
This dataset has open QA in Turkish between users and assistants on various topics.
License
This dataset is released under the Apache 2.0 License.
medscribe-soap-712
MedScribe SOAP Training Data — 712 Curated Samples
Training, validation, and test splits for fine-tuning
google/medgemma-4b-it to
generate concise clinical SOAP notes.
Used to train the MedScribe SOAP LoRA adapter.
Dataset Description
712 medical encounter transcript → SOAP note pairs designed to teach a
language model to produce concise clinical shorthand rather than verbose
textbook prose.
Each sample consists of:
Input : A medical encounter transcript (patient… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/medscribe-soap-712.
