datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ALE-Bench
ALE-Bench
Dataset Description
ALE-Bench is a benchmark for evaluating AI systems on score-based algorithmic programming contests.
This dataset is officially provided by AtCoder Inc..
Please be sure to check the "License" section below.
Please read our blog post and our paper for more details.
Related resources:
Preprint paper (arXiv)
Sakana AI Blog (English)
Sakana AI Blog (Japanese)
GitHub repository
Leaderboard
Usage
Our Python library automatically… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/ALE-Bench.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.sakthai-combined-v7
SakThai Combined v7
Curated, larger-scale instruction-tuning data for tool-calling, function-calling, and agent-style reasoning in the SakThai model family.
Dataset Summary
SakThai Combined v7 extends the v6 family with more multi-turn examples, broader tool coverage, and stronger <tool>/function-calling formatting. It is intended for fine-tuning models that should invoke tools naturally, then continue the conversation after tool results.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7.Saka-Alpaca-v1https://chatgpt.com
Multilingal-sakalt-dataマルチリンガルデータセットです。mitライセンスです。
sakthai-combined-v6
SakThai Combined v6
Part of the SakThai model family — fine-tuning and evaluation corpus for instruction-following and tool-use chat.
Dataset Size
Files: data/train.jsonl, data/test.jsonl
Format: JSONL
License: apache-2.0
Last updated: 2026-08-01
Loading
from datasets import load_dataset
ds = load_dataset("Nanthasit/sakthai-combined-v6", split="train")
test_ds = load_dataset("Nanthasit/sakthai-combined-v6", split="test")
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6.sakhi
Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark
Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.SA-Knowledge
SA-Knowledge
This repository collects corpora and evaluation data for four South African
languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed
for the doctoral thesis Injecting Commonsense Knowledge into Pretrained
Language Models for Low Resource Languages (University of Cape Town, 2026).
Each subset corresponds to a thesis chapter and can be used independently.
Point of contact: Sello Ralethe
Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.FishMath-SFT-Data
FishMath SFT Data
A synthetic SFT (Supervised Fine-Tuning) dataset for mathematical reasoning, used in Kaggle AI Mathematical Olympiad 3 - Progress Prize 3 (AIMO 3) project.
The dataset contains 23,257 correct solution traces generated by multiple frontier open source LLMs, covering competition-level math problems from diverse sources.
Please refer Pushing the Limits: Post-Training High-Capability Models under Strict Inference for the writeup how this model is used to train… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/FishMath-SFT-Data.sakthai-openenv-training
SakThai OpenEnv Training
Part of the SakThai model family.
Dataset Summary
SakThai OpenEnv Training is a pinned runtime-environment dataset for reproducing SakThai training workflows. It stores exact package versions, experimental interface requirements, and notes for openenv, trl, and GRPO integrations used during model training.
Purpose:
Ensure bit-for-bit reproducibility of training runs on GPU clusters
Document experimental dependencies (openenv==0.4.1… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-openenv-training.Gong-Poem-Dataset
龚诗完整全文数据集
本仓库发布经过规范化与逐条校验的龚诗全文。当前版本为 v1.0.1-fulltext,校验日期为 2026-07-14。
数据规模
配置
文件
条目数
内容
canonical_poems
data/poems_full.parquet
17
规范诗作全文
editions
data/editions_full.parquet
24
不同来源、版本见证全文
fragments
data/fragments_full.parquet
4
可核验散句
每个配置均同时提供 Parquet、CSV 和 JSONL;Hugging Face 查看器直接读取带显式字段类型的 Parquet。所有 45 条全文记录均完成 Unicode NFC 规范化,并通过目录记录的行数、非空白字符数及 SHA-256 一致性校验;结果见 data/full_text_validation.json。
使用方式
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Sakanaction/Gong-Poem-Dataset.JMID
JMID: Japanese Medical Incident Dataset
日本語
本データセットは、公益財団法人日本医療機能評価機構の医療事故報告書に書かれている医療事故内容から、医療事故の「具体的内容」「背景・要因」「改善策」とその他の情報をまとめたものである。
使い方の例は以下に載せる。
English
This dataset is compiled from the medical incident reports published by the Japan Council for Quality Health Care. It summarizes the contents of medical incidents, including the specific details, background and contributing factors, and proposed improvements, along with other related information.
An example of how to use the… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/JMID.sakthai-combined-v12
SakThai Combined v12
Part of the SakThai model family — fine-tuning corpus for the retrain round v2
(0.5B / 1.5B tool-calling models, -r2 repos).
Dataset Structure
Files: data/train.jsonl
Format: JSONL, one row per example
Columns: messages (multi-turn), tools (structured definitions)
Rows: 7,437
Composition
Built by scripts/dataset_prep/build_v12.py in the Sak-Family-Agent repo from:
Source
Rows
sakthai-combined-v10
2,965… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v12.sql-create-context-thai
Overview
This dataset builds from sql-create-context.
@misc{b-mc2_2023_sql-create-context,
title = {sql-create-context Dataset},
author = {b-mc2},
year = {2023},
url = {https://huggingface.co/datasets/b-mc2/sql-create-context},
note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.},
}
sakksa
Superior-Reasoning-SFT-gpt-oss-120b
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/sakksa.SakaEval-V1https://chatgpt.com
ColiFormer-Data
ColiFormer Training and Evaluation Dataset
This dataset contains the training and evaluation data used for the ColiFormer model - a specialized codon optimization transformer fine-tuned for Escherichia coli sequences. The model achieves 6.2% better CAI (Codon Adaptation Index) scores compared to the base CodonTransformer model.
🔗 Related Resources
Model: saketh11/ColiFormer
Base Model: adibvafa/CodonTransformer
Paper: CodonTransformer: The Global Codon Optimization… See the full description on the dataset page: https://huggingface.co/datasets/saketh11/ColiFormer-Data.sakthai-combined-v10
SakThai Combined Tool-Calling Dataset v10
A combined, synthetic instruction-tuning dataset for training small language models (0.5B–7B) in tool-calling and function execution.
Quick stats:
Rows: 2,965 examples
Format: JSONL with messages (multi-turn) + tools (structured definitions)
Split: train only
License: Apache 2.0
Verified: 2026-08-01 (Datasets Server, row count = 2,965)
Dataset Summary
This is a combined, synthetic dataset designed to teach… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v10.sakthai-irrelevance-supplement
Dataset Card for SakThai Irrelevance Supplement
Dataset Name: SakThai Irrelevance SupplementAuthor(s): SakThai Agent (Nanthasit)License: MITLast Updated: 2026-08-01
Dataset Summary
The SakThai Irrelevance Supplement is a curated dataset of dialogue examples designed to train and evaluate models to gracefully decline tool use when unnecessary. This dataset specifically contains examples where available functions are NOT needed for the user's request — teaching… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement.sakhi-asha-home-visit-conversations
Sakhi — ASHA Home-Visit Conversations (Hindi/Hinglish → Structured Forms)
Synthetic Hindi/Hinglish conversations between an Indian ASHA (Accredited Social
Health Activist) and a patient during a maternal- and child-health home visit, each
paired with a structured JSON target. Built for the Sakhi project — an offline
voice-to-form tool for ASHA workers (github.com/Tushar-9802/Sakhi).
The dataset supports two supervised tasks over the same conversations:
form_extraction — extract a… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/sakhi-asha-home-visit-conversations.sakthai-coder-browser
SakThai Coder Browser
Part of the SakThai Model Family.
Dataset Summary
SakThai Coder Browser is a synthetic instruction-tuning dataset for coding assistants and browser agents. It provides multi-turn agent-style conversations with structured tool definitions and expected assistant tool calls, built for training small language models on tool use and grounded code/browser workflows.
Owner: Nanthasit
Format: Parquet
Rows: 247
Columns: messages, tools
Task:… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-coder-browser.SpeakMK1_SLP_Dialogue
Dataset Card for SLP Dialogue Dataset
Dataset Description
This dataset consists of 1,000 multi-turn simulated pediatric Speech-Language Pathology (SLP) interaction dialogues. It is specifically designed to train, evaluate, or fine-tune LLMs to act as clinical speech-language therapists or to study clinical reasoning during speech therapy sessions. Each conversation includes clinical metadata (child age, speech sound disorder category, specific phone error, clinical goal… See the full description on the dataset page: https://huggingface.co/datasets/SakhrML/SpeakMK1_SLP_Dialogue.llama3_cyber_reasoning_chatllama3_t2p_reasoning_chatsakuraeval
SakuraEval
Dataset Description
SakuraEval is a Japan-specific code generation benchmark dataset.
It is designed independently and does not rely on translation from English benchmarks such as HumanEval or JHumanEval.
The dataset is currently being reviewed for official release.
Dataset Structure
from datasets import load_dataset
load_dataset("kogi-jwu/sakuraeval", "ja")
DatasetDict({
test: Dataset({
features: ['task_id', 'category', 'prompt'… See the full description on the dataset page: https://huggingface.co/datasets/kogi-jwu/sakuraeval.llama3_ibm_reasoning_chatllama3_acre_reasoning_chatAbkhaz-chatgptこちらはchatgptに生成してもらったサンプルです。
This is a sample generated by chatgpt.
alfworld-distilled-qwen3-32b
ALFWorld Distilled Trajectories (Model-Generated)
Overview
This dataset contains interaction trajectories collected by allowing a teacher
language model to interact with the ALFWorld environment using a ReAct-style policy.
These trajectories are intended for use as model-generated supervision signals
in agent training via supervised fine-tuning (SFT).
Data Generation Process
Trajectories were generated by running the ALFWorld simulator and allowing a
teacher… See the full description on the dataset page: https://huggingface.co/datasets/Sakai0920/alfworld-distilled-qwen3-32b.saishin-ab
