CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SakanaAI /ALE-Bench ALE-Bench Dataset Description ALE-Bench is a benchmark for evaluating AI systems on score-based algorithmic programming contests. This dataset is officially provided by AtCoder Inc.. Please be sure to check the "License" section below. Please read our blog post and our paper for more details. Related resources: Preprint paper (arXiv) Sakana AI Blog (English) Sakana AI Blog (Japanese) GitHub repository Leaderboard Usage Our Python library automatically… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/ALE-Bench.imageimage-text-to-textn<1K13 likes15k downloads1y agoHugging Face02Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes491 downloads1y agoHugging Face03Nanthasit /sakthai-combined-v7 SakThai Combined v7 Curated, larger-scale instruction-tuning data for tool-calling, function-calling, and agent-style reasoning in the SakThai model family. Dataset Summary SakThai Combined v7 extends the v6 family with more multi-turn examples, broader tool coverage, and stronger <tool>/function-calling formatting. It is intended for fine-tuning models that should invoke tools naturally, then continue the conversation after tool results. Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7.texttext-generation1K<n<10K0 likes187 downloads2mo agoHugging Face04Sakalti /Saka-Alpaca-v1https://chatgpt.com texttext-generationn<1K0 likes114 downloads2y agoHugging Face05Sakalti /Multilingal-sakalt-dataマルチリンガルデータセットです。mitライセンスです。 texttext-generation1K<n<10K1 likes111 downloads2y agoHugging Face06Nanthasit /sakthai-combined-v6 SakThai Combined v6 Part of the SakThai model family — fine-tuning and evaluation corpus for instruction-following and tool-use chat. Dataset Size Files: data/train.jsonl, data/test.jsonl Format: JSONL License: apache-2.0 Last updated: 2026-08-01 Loading from datasets import load_dataset ds = load_dataset("Nanthasit/sakthai-combined-v6", split="train") test_ds = load_dataset("Nanthasit/sakthai-combined-v6", split="test") Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6.documenttext-generation1K<n<10K0 likes98 downloads2mo agoHugging Face07SimPPL /sakhi Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.tabularquestion-answering1K<n<10K2 likes92 downloads5mo agoHugging Face08sello-ralethe /SA-Knowledge SA-Knowledge This repository collects corpora and evaluation data for four South African languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Each subset corresponds to a thesis chapter and can be used independently. Point of contact: Sello Ralethe Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.tabulartranslation10K<n<100K0 likes86 downloads1mo agoHugging Face09SakanaAI /FishMath-SFT-Data FishMath SFT Data A synthetic SFT (Supervised Fine-Tuning) dataset for mathematical reasoning, used in Kaggle AI Mathematical Olympiad 3 - Progress Prize 3 (AIMO 3) project. The dataset contains 23,257 correct solution traces generated by multiple frontier open source LLMs, covering competition-level math problems from diverse sources. Please refer Pushing the Limits: Post-Training High-Capability Models under Strict Inference for the writeup how this model is used to train… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/FishMath-SFT-Data.texttext-generation10K<n<100K3 likes54 downloads5mo agoHugging Face10Nanthasit /sakthai-openenv-training SakThai OpenEnv Training Part of the SakThai model family. Dataset Summary SakThai OpenEnv Training is a pinned runtime-environment dataset for reproducing SakThai training workflows. It stores exact package versions, experimental interface requirements, and notes for openenv, trl, and GRPO integrations used during model training. Purpose: Ensure bit-for-bit reproducibility of training runs on GPU clusters Document experimental dependencies (openenv==0.4.1… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-openenv-training.texttext-generationn<1K0 likes48 downloads2mo agoHugging Face11Sakanaction /Gong-Poem-Dataset 龚诗完整全文数据集 本仓库发布经过规范化与逐条校验的龚诗全文。当前版本为 v1.0.1-fulltext,校验日期为 2026-07-14。 数据规模 配置 文件 条目数 内容 canonical_poems data/poems_full.parquet 17 规范诗作全文 editions data/editions_full.parquet 24 不同来源、版本见证全文 fragments data/fragments_full.parquet 4 可核验散句 每个配置均同时提供 Parquet、CSV 和 JSONL;Hugging Face 查看器直接读取带显式字段类型的 Parquet。所有 45 条全文记录均完成 Unicode NFC 规范化,并通过目录记录的行数、非空白字符数及 SHA-256 一致性校验;结果见 data/full_text_validation.json。 使用方式 from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Sakanaction/Gong-Poem-Dataset.tabulartext-generationn<1K0 likes45 downloads2mo agoHugging Face12Sakaji-Lab /JMID JMID: Japanese Medical Incident Dataset 日本語 本データセットは、公益財団法人日本医療機能評価機構の医療事故報告書に書かれている医療事故内容から、医療事故の「具体的内容」「背景・要因」「改善策」とその他の情報をまとめたものである。 使い方の例は以下に載せる。 English This dataset is compiled from the medical incident reports published by the Japan Council for Quality Health Care. It summarizes the contents of medical incidents, including the specific details, background and contributing factors, and proposed improvements, along with other related information. An example of how to use the… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/JMID.texttext-classification1K<n<10K2 likes35 downloads1y agoHugging Face13Nanthasit /sakthai-combined-v12 SakThai Combined v12 Part of the SakThai model family — fine-tuning corpus for the retrain round v2 (0.5B / 1.5B tool-calling models, -r2 repos). Dataset Structure Files: data/train.jsonl Format: JSONL, one row per example Columns: messages (multi-turn), tools (structured definitions) Rows: 7,437 Composition Built by scripts/dataset_prep/build_v12.py in the Sak-Family-Agent repo from: Source Rows sakthai-combined-v10 2,965… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v12.texttext-generation1K<n<10K0 likes32 downloads2mo agoHugging Face14saksornr /sql-create-context-thai Overview This dataset builds from sql-create-context. @misc{b-mc2_2023_sql-create-context, title = {sql-create-context Dataset}, author = {b-mc2}, year = {2023}, url = {https://huggingface.co/datasets/b-mc2/sql-create-context}, note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.}, } texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face15onurborasahin /sakksa Superior-Reasoning-SFT-gpt-oss-120b          🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/sakksa.texttext-generation100K<n<1M0 likes27 downloads8mo agoHugging Face16Sakalti /SakaEval-V1https://chatgpt.com texttext-generationn<1K0 likes25 downloads2y agoHugging Face17saketh11 /ColiFormer-Data ColiFormer Training and Evaluation Dataset This dataset contains the training and evaluation data used for the ColiFormer model - a specialized codon optimization transformer fine-tuned for Escherichia coli sequences. The model achieves 6.2% better CAI (Codon Adaptation Index) scores compared to the base CodonTransformer model. 🔗 Related Resources Model: saketh11/ColiFormer Base Model: adibvafa/CodonTransformer Paper: CodonTransformer: The Global Codon Optimization… See the full description on the dataset page: https://huggingface.co/datasets/saketh11/ColiFormer-Data.texttext-generationn<1K0 likes25 downloads1y agoHugging Face18Nanthasit /sakthai-combined-v10 SakThai Combined Tool-Calling Dataset v10 A combined, synthetic instruction-tuning dataset for training small language models (0.5B–7B) in tool-calling and function execution. Quick stats: Rows: 2,965 examples Format: JSONL with messages (multi-turn) + tools (structured definitions) Split: train only License: Apache 2.0 Verified: 2026-08-01 (Datasets Server, row count = 2,965) Dataset Summary This is a combined, synthetic dataset designed to teach… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v10.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face19Nanthasit /sakthai-irrelevance-supplement Dataset Card for SakThai Irrelevance Supplement Dataset Name: SakThai Irrelevance SupplementAuthor(s): SakThai Agent (Nanthasit)License: MITLast Updated: 2026-08-01 Dataset Summary The SakThai Irrelevance Supplement is a curated dataset of dialogue examples designed to train and evaluate models to gracefully decline tool use when unnecessary. This dataset specifically contains examples where available functions are NOT needed for the user's request — teaching… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement.texttext-generationn<1K0 likes22 downloads2mo agoHugging Face20Tushar9802 /sakhi-asha-home-visit-conversations Sakhi — ASHA Home-Visit Conversations (Hindi/Hinglish → Structured Forms) Synthetic Hindi/Hinglish conversations between an Indian ASHA (Accredited Social Health Activist) and a patient during a maternal- and child-health home visit, each paired with a structured JSON target. Built for the Sakhi project — an offline voice-to-form tool for ASHA workers (github.com/Tushar-9802/Sakhi). The dataset supports two supervised tasks over the same conversations: form_extraction — extract a… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/sakhi-asha-home-visit-conversations.texttext-generation1K<n<10K0 likes20 downloads4mo agoHugging Face21Nanthasit /sakthai-coder-browser SakThai Coder Browser Part of the SakThai Model Family. Dataset Summary SakThai Coder Browser is a synthetic instruction-tuning dataset for coding assistants and browser agents. It provides multi-turn agent-style conversations with structured tool definitions and expected assistant tool calls, built for training small language models on tool use and grounded code/browser workflows. Owner: Nanthasit Format: Parquet Rows: 247 Columns: messages, tools Task:… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-coder-browser.texttext-generationn<1K0 likes20 downloads2mo agoHugging Face22SakhrML /SpeakMK1_SLP_Dialogue Dataset Card for SLP Dialogue Dataset Dataset Description This dataset consists of 1,000 multi-turn simulated pediatric Speech-Language Pathology (SLP) interaction dialogues. It is specifically designed to train, evaluate, or fine-tune LLMs to act as clinical speech-language therapists or to study clinical reasoning during speech therapy sessions. Each conversation includes clinical metadata (child age, speech sound disorder category, specific phone error, clinical goal… See the full description on the dataset page: https://huggingface.co/datasets/SakhrML/SpeakMK1_SLP_Dialogue.texttext-generation1K<n<10K1 likes17 downloads4mo agoHugging Face23Sakuna /llama3_cyber_reasoning_chatgatedtexttext-generation10K<n<100K0 likes13 downloads1y agoHugging Face24Sakuna /llama3_t2p_reasoning_chatgatedtexttext-generation10K<n<100K0 likes13 downloads1y agoHugging Face25kogi-jwu /sakuraeval SakuraEval Dataset Description SakuraEval is a Japan-specific code generation benchmark dataset. It is designed independently and does not rely on translation from English benchmarks such as HumanEval or JHumanEval. The dataset is currently being reviewed for official release. Dataset Structure from datasets import load_dataset load_dataset("kogi-jwu/sakuraeval", "ja") DatasetDict({ test: Dataset({ features: ['task_id', 'category', 'prompt'… See the full description on the dataset page: https://huggingface.co/datasets/kogi-jwu/sakuraeval.texttext-generationn<1K0 likes12 downloads6mo agoHugging Face26Sakuna /llama3_ibm_reasoning_chatgatedtexttext-generation10K<n<100K0 likes10 downloads1y agoHugging Face27Sakuna /llama3_acre_reasoning_chatgatedtexttext-generation10K<n<100K0 likes10 downloads1y agoHugging Face28sakaltcommunity /Abkhaz-chatgptこちらはchatgptに生成してもらったサンプルです。 This is a sample generated by chatgpt. texttext-generationn<1K0 likes10 downloads2y agoHugging Face29Sakai0920 /alfworld-distilled-qwen3-32b ALFWorld Distilled Trajectories (Model-Generated) Overview This dataset contains interaction trajectories collected by allowing a teacher language model to interact with the ALFWorld environment using a ReAct-style policy. These trajectories are intended for use as model-generated supervision signals in agent training via supervised fine-tuning (SFT). Data Generation Process Trajectories were generated by running the ALFWorld simulator and allowing a teacher… See the full description on the dataset page: https://huggingface.co/datasets/Sakai0920/alfworld-distilled-qwen3-32b.texttext-generationn<1K0 likes10 downloads7mo agoHugging Face30Sakalti /saishin-abtexttext-generationn<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.