CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KokosDev /tahoe-100m-zarr Tahoe-100M Zarr Collection Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.textn<1K1 likes16k downloads6mo agoHugging Face02KokosDev /single-cell-brain-zarr Single-Cell Brain Zarr Collection Production-ready brain single-cell RNA-seq data exported from the CellxGene Census into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything useful.… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-brain-zarr.textn<1K1 likes2.8k downloads7mo agoHugging Face03kokimashita /wild_vision_sftimage10K<n<100K0 likes484 downloads2y agoHugging Face04darknight054 /med-mts-audio-kokoro-82m MTSamples‑Kokoro‑ASR (Synthetic Medical Speech) Summary: 279 hours of synthetic English medical speech (49,462 clips) created from publicly available transcripts on MTSamples.com using multiple US/UK voices from Kokoro‑82M. Intended for training and evaluating medical ASR. Dataset Rows: 49,462 Total audio: ~279 hours (mono) Source text: Sample medical reports from MTSamples.com (names/dates typically altered or removed) Audio generation: hexgrad/Kokoro‑82M (various… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/med-mts-audio-kokoro-82m.audio10K<n<100K0 likes472 downloads1y agoHugging Face05Kokoslocke /NACA_4_Digit_for_ML NACA 4-Digit Airfoil CFD Dataset Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions. Dataset Summary ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles AoA range: −5° to +5° Reynolds number range: 100,000 – 500,000 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.tabularothern<1K0 likes464 downloads3mo agoHugging Face06koki0702 /zero-llm-data Zero LLM Dataset (Text Corpora + BPE Encoded Versions + Merge Rules) Repository: koki0702/zero-llm-data This dataset provides cleaned, standardized text corpora derived from several public datasets, prepared specifically for use in the book: “ゼロから作る Deep Learning ❻ —— LLM 編”(Zero to Deep Learning — LLM Edition) Each corpus is organized into its own directory (codebot/, storybot/, webbot/) so that text data, BPE-encoded data, and merge rules are grouped together. This… See the full description on the dataset page: https://huggingface.co/datasets/koki0702/zero-llm-data.text100M<n<1B0 likes452 downloads8mo agoHugging Face07koke /AllTheBacteria-FCGR-7mertext1M<n<10M0 likes451 downloads1y agoHugging Face08darknight054 /med-mts-audio-kokoro-82m-noisy16k-v1audio10K<n<100K0 likes365 downloads1y agoHugging Face09Firoj112 /nepali-kokoro-ft-data Nepali Kokoro Fine-Tuning Dataset This is a sharded, processed dataset containing Nepali voice data for Kokoro TTS fine-tuning. audio10K<n<100K0 likes340 downloads2mo agoHugging Face10Podtech /llm-jp-corpus-v4-ja_kokkai_giji llm-jp-corpus-v4 — ja_kokkai_giji Mirror of the ja/ja_kokkai_giji sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_kokkai_giji Files: 12 × jsonl.gz (1.3 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 —… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kokkai_giji.texttext-generation10K<n<100K0 likes270 downloads2mo agoHugging Face11KokosDev /single-cell-lung-zarr Single-cell lung (CellxGene Census) — Zarr This dataset was exported from the CellxGene Census as a chunked + compressed Zarr store intended for easy streaming access. Source: CellxGene Census API Organism: Homo sapiens Filter: tissue_general == 'lung' and is_primary_data == True Shape: 100,000 cells × 61,497 genes Zarr path: lung.zarr Compression Uncompressed (dense float32): 22.91 GB Compressed Zarr: ~307 MB (322 MB on Hub) Compression ratio: ~76× (Blosc zstd on… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-lung-zarr.textn<1K1 likes195 downloads7mo agoHugging Face12humair025 /kokutext10K<n<100K0 likes183 downloads15d agoHugging Face13humanalysis-square /KokushiMD-10 KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations Overview KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.textquestion-answering10K<n<100K7 likes176 downloads1y agoHugging Face14kokhayas /english-debate-motions-utdsEnglish Debate Motions gathered by University of Tokyo Debate Society @misc{english-debate-motions-utds, title={english-debate-motions-utds}, author={members of the University of Tokyo Debate Society}, year={2022}, } tabular10K<n<100K3 likes131 downloads4y agoHugging Face15kokul /bank-marketing-propensity Introduction This project explores several classification techniques as applied to a bank's marketing campaign data. The classification goal is to predict whether the client will subscribe a term deposit (variable y). Source: https://archive.ics.uci.edu/ml/datasets/bank+marketing It's recommended that the viewer read the Jupyter Notebook in NBViewer: https://nbviewer.jupyter.org/github/sgus1318/marketing_propensity/blob/master/Bank_DirectMarketing_Propensity.ipynb… See the full description on the dataset page: https://huggingface.co/datasets/kokul/bank-marketing-propensity.tabular10K<n<100K1 likes118 downloads1y agoHugging Face16UEC-InabaLab /KokoroChat KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors KokoroChat is the largest human-collected Japanese psychological counseling dialogue dataset to date (as of June 2025). It was created through role-playing between trained counselors and includes rich, long-form dialogues and detailed client feedback on counseling quality. The dataset supports research on empathetic response generation, dialogue evaluation… See the full description on the dataset page: https://huggingface.co/datasets/UEC-InabaLab/KokoroChat.texttext-generation1K<n<10K2 likes110 downloads1y agoHugging Face17echodict /kokoro-ttstextn<1K0 likes96 downloads8mo agoHugging Face18Dobby091 /kokotext1K<n<10K2 likes77 downloads2y agoHugging Face19ecyht2 /kokoro-82M-voices Kokoro-82M Voices This dataset contains all the voices available in hexgrad/Kokoro-82M. This dataset provides the voices in 3 different formats. Individual voices embeddings in different JSON file Single JSON which contains all the voices in a JSON object. Parquet format for usage via datasets The voices name is the same as the .pth file names shown below. voices = [ "af", "af_bella", "af_nicole", "af_sarah", "af_sky", "am_adam", "am_michael"… See the full description on the dataset page: https://huggingface.co/datasets/ecyht2/kokoro-82M-voices.texttext-to-speechn<1K3 likes71 downloads2y agoHugging Face20Koki-Kurita /DataSet_mix_duck_oct_cabtabularn<1K1 likes71 downloads13d agoHugging Face21kokolamba /keen_popqa_gpt2xl_generationstext10K<n<100K0 likes52 downloads8mo agoHugging Face22Whitermite /kokodocument1K<n<10K0 likes48 downloads5d agoHugging Face23ENSEONG /ko-ko-math-500-test-EXAONE-4.0-1.2B-bontabular1K<n<10K0 likes44 downloads8mo agoHugging Face24ENSEONG /ko-ko-math-500-test-Qwen2.5-3B-Instruct-bontabular1K<n<10K0 likes43 downloads8mo agoHugging Face25luosuu /SWE-kokkos-bench SWE-kokkos-bench SWE-kokkos-bench is a public, verifier-backed benchmark of 100 repository-level software-engineering tasks mined from merged pull requests in the Kokkos ecosystem. Each task starts from the parent revision of a real pull request and asks an agent to implement the corresponding change. Correctness is checked by task-specific build and regression commands against a held-out test patch. The dataset supports two complementary interfaces: this Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/luosuu/SWE-kokkos-bench.texttext-generationn<1K0 likes41 downloads1mo agoHugging Face26kokojake /oasst2_egyptian_arabic_convstexttext-generation1K<n<10K1 likes40 downloads2y agoHugging Face27jkeisling /libritts-r-mimi-kokorotext100K<n<1M0 likes40 downloads2y agoHugging Face28Koki-Kurita /DataSet_mix_duck_octtabularn<1K0 likes39 downloads14d agoHugging Face29sdmy /kokborok Kokborok Digitalisation Project The Kokborok Digitalisation Project is an initiative to curate and enhance parallel data for the Kokborok-English language pair. This project builds upon the SMOL dataset by Google, available on Hugging Face, and involves modifying and correcting it to better reflect the nuances of the local Kokborok dialect. From the Author "Language is a living, breathing entity—constantly evolving, shaping cultures, and connecting generations. When we… See the full description on the dataset page: https://huggingface.co/datasets/sdmy/kokborok.text1 likes36 downloads1y agoHugging Face30sdialog /voices-kokorotextn<1K0 likes34 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.