datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.single-cell-brain-zarr
Single-Cell Brain Zarr Collection
Production-ready brain single-cell RNA-seq data exported from the CellxGene Census into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything useful.… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-brain-zarr.wild_vision_sftmed-mts-audio-kokoro-82m
MTSamples‑Kokoro‑ASR (Synthetic Medical Speech)
Summary: 279 hours of synthetic English medical speech (49,462 clips) created from publicly available transcripts on MTSamples.com using multiple US/UK voices from Kokoro‑82M. Intended for training and evaluating medical ASR.
Dataset
Rows: 49,462
Total audio: ~279 hours (mono)
Source text: Sample medical reports from MTSamples.com (names/dates typically altered or removed)
Audio generation: hexgrad/Kokoro‑82M (various… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/med-mts-audio-kokoro-82m.NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
AoA range: −5° to +5°
Reynolds number range: 100,000 – 500,000
129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.zero-llm-data
Zero LLM Dataset
(Text Corpora + BPE Encoded Versions + Merge Rules)
Repository: koki0702/zero-llm-data
This dataset provides cleaned, standardized text corpora derived from several public
datasets, prepared specifically for use in the book:
“ゼロから作る Deep Learning ❻ —— LLM 編”(Zero to Deep Learning — LLM Edition)
Each corpus is organized into its own directory (codebot/, storybot/, webbot/)
so that text data, BPE-encoded data, and merge rules are grouped together.
This… See the full description on the dataset page: https://huggingface.co/datasets/koki0702/zero-llm-data.AllTheBacteria-FCGR-7mermed-mts-audio-kokoro-82m-noisy16k-v1nepali-kokoro-ft-data
Nepali Kokoro Fine-Tuning Dataset
This is a sharded, processed dataset containing Nepali voice data for Kokoro TTS fine-tuning.
llm-jp-corpus-v4-ja_kokkai_giji
llm-jp-corpus-v4 — ja_kokkai_giji
Mirror of the ja/ja_kokkai_giji sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_kokkai_giji
Files: 12 × jsonl.gz (1.3 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 —… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kokkai_giji.single-cell-lung-zarr
Single-cell lung (CellxGene Census) — Zarr
This dataset was exported from the CellxGene Census as a chunked + compressed Zarr store intended for easy streaming access.
Source: CellxGene Census API
Organism: Homo sapiens
Filter: tissue_general == 'lung' and is_primary_data == True
Shape: 100,000 cells × 61,497 genes
Zarr path: lung.zarr
Compression
Uncompressed (dense float32): 22.91 GB
Compressed Zarr: ~307 MB (322 MB on Hub)
Compression ratio: ~76× (Blosc zstd on… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-lung-zarr.kokuKokushiMD-10
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
Overview
KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.english-debate-motions-utdsEnglish Debate Motions gathered by University of Tokyo Debate Society
@misc{english-debate-motions-utds,
title={english-debate-motions-utds},
author={members of the University of Tokyo Debate Society},
year={2022},
}
bank-marketing-propensity
Introduction
This project explores several classification techniques as applied to a bank's marketing campaign data. The classification goal is to predict whether the client will subscribe a term deposit (variable y).
Source: https://archive.ics.uci.edu/ml/datasets/bank+marketing
It's recommended that the viewer read the Jupyter Notebook in NBViewer: https://nbviewer.jupyter.org/github/sgus1318/marketing_propensity/blob/master/Bank_DirectMarketing_Propensity.ipynb… See the full description on the dataset page: https://huggingface.co/datasets/kokul/bank-marketing-propensity.KokoroChat
KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors
KokoroChat is the largest human-collected Japanese psychological counseling dialogue dataset to date (as of June 2025). It was created through role-playing between trained counselors and includes rich, long-form dialogues and detailed client feedback on counseling quality. The dataset supports research on empathetic response generation, dialogue evaluation… See the full description on the dataset page: https://huggingface.co/datasets/UEC-InabaLab/KokoroChat.kokoro-ttskokokokoro-82M-voices
Kokoro-82M Voices
This dataset contains all the voices available in hexgrad/Kokoro-82M.
This dataset provides the voices in 3 different formats.
Individual voices embeddings in different JSON file
Single JSON which contains all the voices in a JSON object.
Parquet format for usage via datasets
The voices name is the same as the .pth file names shown below.
voices = [
"af",
"af_bella",
"af_nicole",
"af_sarah",
"af_sky",
"am_adam",
"am_michael"… See the full description on the dataset page: https://huggingface.co/datasets/ecyht2/kokoro-82M-voices.DataSet_mix_duck_oct_cabkeen_popqa_gpt2xl_generationskokoko-ko-math-500-test-EXAONE-4.0-1.2B-bonko-ko-math-500-test-Qwen2.5-3B-Instruct-bonSWE-kokkos-bench
SWE-kokkos-bench
SWE-kokkos-bench is a public, verifier-backed benchmark of 100 repository-level
software-engineering tasks mined from merged pull requests in the Kokkos ecosystem. Each task
starts from the parent revision of a real pull request and asks an agent to implement the
corresponding change. Correctness is checked by task-specific build and regression commands
against a held-out test patch.
The dataset supports two complementary interfaces:
this Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/luosuu/SWE-kokkos-bench.oasst2_egyptian_arabic_convslibritts-r-mimi-kokoroDataSet_mix_duck_octkokborok
Kokborok Digitalisation Project
The Kokborok Digitalisation Project is an initiative to curate and enhance parallel data for the Kokborok-English language pair. This project builds upon the SMOL dataset by Google, available on Hugging Face, and involves modifying and correcting it to better reflect the nuances of the local Kokborok dialect.
From the Author
"Language is a living, breathing entity—constantly evolving, shaping cultures, and connecting generations. When we… See the full description on the dataset page: https://huggingface.co/datasets/sdmy/kokborok.voices-kokoro
