datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.single-cell-brain-zarr
Single-Cell Brain Zarr Collection
Production-ready brain single-cell RNA-seq data exported from the CellxGene Census into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything useful.… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-brain-zarr.llm-jp-corpus-v4-ja_kokkai_giji
llm-jp-corpus-v4 — ja_kokkai_giji
Mirror of the ja/ja_kokkai_giji sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_kokkai_giji
Files: 12 × jsonl.gz (1.3 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 —… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kokkai_giji.single-cell-lung-zarr
Single-cell lung (CellxGene Census) — Zarr
This dataset was exported from the CellxGene Census as a chunked + compressed Zarr store intended for easy streaming access.
Source: CellxGene Census API
Organism: Homo sapiens
Filter: tissue_general == 'lung' and is_primary_data == True
Shape: 100,000 cells × 61,497 genes
Zarr path: lung.zarr
Compression
Uncompressed (dense float32): 22.91 GB
Compressed Zarr: ~307 MB (322 MB on Hub)
Compression ratio: ~76× (Blosc zstd on… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-lung-zarr.KokushiMD-10
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
Overview
KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.KokoroChat
KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors
KokoroChat is the largest human-collected Japanese psychological counseling dialogue dataset to date (as of June 2025). It was created through role-playing between trained counselors and includes rich, long-form dialogues and detailed client feedback on counseling quality. The dataset supports research on empathetic response generation, dialogue evaluation… See the full description on the dataset page: https://huggingface.co/datasets/UEC-InabaLab/KokoroChat.DataSet_mix_duck_oct_cabSWE-kokkos-bench
SWE-kokkos-bench
SWE-kokkos-bench is a public, verifier-backed benchmark of 100 repository-level
software-engineering tasks mined from merged pull requests in the Kokkos ecosystem. Each task
starts from the parent revision of a real pull request and asks an agent to implement the
corresponding change. Correctness is checked by task-specific build and regression commands
against a held-out test patch.
The dataset supports two complementary interfaces:
this Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/luosuu/SWE-kokkos-bench.DataSet_mix_duck_octinmQA_ja_9.315kくさい子
keirin_sharegptnatto-10kPierrot-8.94kkokurenTweet_5kKokoro-TTS-Batchkokokokomi
