datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.capstone_sakuga_preproc_optical_flowcapstone_sakuga_iblip_t5_embeddingsEDINET-Bench
EDINET-Bench
📚 Paper | 📝 Blog | 🧑💻 Code
EDINET-Bench is a Japanese financial benchmark designed to evaluate the performance of LLMs on challenging financial tasks including accounting fraud detection, earnings forecasting, and industry prediction.
This dataset is built leveraging EDINET, a platform managed by the Financial Services Agency (FSA) of Japan that provides access to disclosure documents such as securities reports.
Notice
June 9, 2025: This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/EDINET-Bench.i-love-anime-sakuga
ilovehentai9000/iloveanimesakuga Dataset
Because the website is slow and I hate people who request for "Data" to "Improve" their model. There's no need for this kind of BS.
Uses
Just don't.
License
GAYSEX-Dont Be A Prick License
fixed-tokenizer-morphscore-segmentscapstone_sakuga_preproc_mid_framecapstone_sakuga_vae_latentsso100_blue_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 17,
"total_frames": 15252,
"total_tasks": 1,
"total_videos": 17,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sakumaya/so100_blue_box.agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
capstone_sakuga_simple_description_mlm_hssudoku-bench-nikoli
Sudoku-Bench: nikoli_100
🐙 [Sudoku-Bench GitHub]
📝 [Technical Report]
The Sudoku-Bench project has been discontinued. The nikoli_100 puzzle collection remains available for download here.
Dataset description
The nikoli_100 dataset contains 100 beautiful handmade standard Sudoku puzzles designed by Nikoli, the Japanese puzzle company that popularized Sudoku in the 1980s. The collection was curated in partnership with Nikoli as part of Sudoku-Bench, a benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/sudoku-bench-nikoli.details_Sakalti__Saka-7.6B
Dataset Card for Evaluation run of Sakalti/Saka-7.6B
Dataset automatically created during the evaluation run of model Sakalti/Saka-7.6B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Sakalti__Saka-7.6B.capstone_sakuga_simple_descriptionsakhi
Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark
Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.sakuga_preprocessedSA-Knowledge
SA-Knowledge
This repository collects corpora and evaluation data for four South African
languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed
for the doctoral thesis Injecting Commonsense Knowledge into Pretrained
Language Models for Low Resource Languages (University of Cape Town, 2026).
Each subset corresponds to a thesis chapter and can be used independently.
Point of contact: Sello Ralethe
Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.pick_place_demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_robot",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 10,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sakehaosdfadfasf/pick_place_demo.sakuga-3
ilovehentai9000/iloveanimesakuga Dataset
Because the website is slow and I hate people who request for "Data" to "Improve" their model. There's no need for this kind of BS.
Uses
Just don't.
Copyright
Skill issue. Also read the below for big corp users.
Special Notice
In light of recent AI Developments, we define the additional clauses:
This applies only to the following corporations, organizations and thier related subsidaries as listed:
Alphabet… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/sakuga-3.PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details
Dataset Card for Evaluation run of PocketDoc/Dans-SakuraKaze-V1.0.0-12b
Dataset automatically created during the evaluation run of model PocketDoc/Dans-SakuraKaze-V1.0.0-12b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PocketDoc__Dans-SakuraKaze-V1.0.0-12b-details.MagBridge-Battery
MagBridge-Battery v1.0
The first open dataset pairing battery magnetic-field signatures with electrochemical degradation labels.
Overview
Battery health diagnostics rely almost entirely on terminal measurements — voltage, current, temperature. Magnetic sensing can see what terminals miss: internal hotspots, dendrites, inhomogeneous degradation.
But no public dataset connected magnetic signatures to degradation labels. Until now.
MagBridge-Battery v1.0 bridges… See the full description on the dataset page: https://huggingface.co/datasets/Sakthigsjhy/MagBridge-Battery.fwwoMetadata:
https://datasets-server.huggingface.co/statistics?dataset=saksornr/fwwo&config=default&split=train
Trump_Voice_Dataset
Trump Voice Dataset
This dataset contains audio clips of Donald Trump's speech from the World Economic Forum (WEF) 2018, paired with their corresponding transcriptions. The dataset is designed for text-to-speech (TTS) and speech recognition tasks.
Dataset Description
Dataset Summary
The Trump Voice Dataset consists of 20 audio samples (10 train, 10 test) extracted from Donald Trump's speech at the World Economic Forum 2018. Each audio clip is approximately 10… See the full description on the dataset page: https://huggingface.co/datasets/Sakchham19/Trump_Voice_Dataset.Sakalti__ultiima-32B-details
Dataset Card for Evaluation run of Sakalti/ultiima-32B
Dataset automatically created during the evaluation run of model Sakalti/ultiima-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-32B-details.Sakalti__SJT-7.5B-details
Dataset Card for Evaluation run of Sakalti/SJT-7.5B
Dataset automatically created during the evaluation run of model Sakalti/SJT-7.5B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-7.5B-details.Sakalti__ultiima-72B-v1.5-details
Dataset Card for Evaluation run of Sakalti/ultiima-72B-v1.5
Dataset automatically created during the evaluation run of model Sakalti/ultiima-72B-v1.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-72B-v1.5-details.Gong-Poem-Dataset
龚诗完整全文数据集
本仓库发布经过规范化与逐条校验的龚诗全文。当前版本为 v1.0.1-fulltext,校验日期为 2026-07-14。
数据规模
配置
文件
条目数
内容
canonical_poems
data/poems_full.parquet
17
规范诗作全文
editions
data/editions_full.parquet
24
不同来源、版本见证全文
fragments
data/fragments_full.parquet
4
可核验散句
每个配置均同时提供 Parquet、CSV 和 JSONL;Hugging Face 查看器直接读取带显式字段类型的 Parquet。所有 45 条全文记录均完成 Unicode NFC 规范化,并通过目录记录的行数、非空白字符数及 SHA-256 一致性校验;结果见 data/full_text_validation.json。
使用方式
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Sakanaction/Gong-Poem-Dataset.Sakalti__Saka-7.2B-details
Dataset Card for Evaluation run of Sakalti/Saka-7.2B
Dataset automatically created during the evaluation run of model Sakalti/Saka-7.2B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Saka-7.2B-details.Sakalti__Neptuno-Alpha-details
Dataset Card for Evaluation run of Sakalti/Neptuno-Alpha
Dataset automatically created during the evaluation run of model Sakalti/Neptuno-Alpha
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Neptuno-Alpha-details.Sakalti__SJT-3.7B-details
Dataset Card for Evaluation run of Sakalti/SJT-3.7B
Dataset automatically created during the evaluation run of model Sakalti/SJT-3.7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__SJT-3.7B-details.
