CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gnucleus-ai /cad-gen-freecad-bench Parametric CAD Bench — results dataset Run-by-run results for Parametric CAD Bench, a benchmark that measures whether AI agents can author editable FreeCAD models from natural-language part descriptions. 1000 rows, one per (agent, model, task_id, trial) over the gnucleus-ai/cad-bench@v1 task suite. The public leaderboard view of this data lives at cadbench.ai. What's in here data/cad-bench-v1.parquet — the row table. Each row carries the composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.tabular1K<n<10K2 likes12k downloads1mo agoHugging Face02Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes5.5k downloads7mo agoHugging Face03Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes4.1k downloads9mo agoHugging Face04AISC-Linear-Probe-Gen /deception-activationstabular10K<n<100K0 likes3.3k downloads9mo agoHugging Face05natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes3k downloads2y agoHugging Face06General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.8k downloads6mo agoHugging Face07GenerTeam /pretrain_data_eukaryote GENERator-v2-Eukaryote Gene-Centric Pretraining Corpus This repository provides the gene-centric pretraining corpus underlying GENERator-v2-Eukaryote, a large-scale DNA language model for eukaryotic genome understanding. The dataset is constructed by leveraging RefSeq annotations to extract biologically meaningful functional genomic regions, which serve as the foundation for large-context DNA language model pretraining. 📌 Dataset Construction Overview The core… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/pretrain_data_eukaryote.tabulartext-generationn<1K4 likes2.6k downloads5mo agoHugging Face08AstraTeam /generated-csvstabular100K<n<1M0 likes2.5k downloads10mo agoHugging Face09longevity-db /aging-gene-expression-single-cell-mouse A single-cell transcriptomic atlas characterizes ageing tissues in the mouse https://www.nature.com/articles/s41586-020-2496-1#Sec2 Code to download and process this dataset is available in: https://github.com/seanome/2025-longevity-x-ai-hackathon Dataset structure is originally from AnnData. Descriptions of each data file is below. Data Files This dataset contains multiple parquet files, one for each sheet in the original Excel file:… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/aging-gene-expression-single-cell-mouse.tabular100K<n<1M0 likes2.3k downloads1y agoHugging Face10gnucleus-ai /cad-gen-freecad-bench-v2 Parametric CAD Bench v2 — results dataset Run-by-run results for Parametric CAD Bench v2, a benchmark that measures whether AI agents can author editable FreeCAD models from natural-language part descriptions. This archive contains 1,000 rows: one trial for each of 100 tasks across the 10 public jobs on the live gnucleus-ai/cad-bench@v2 leaderboard. What's in here data/cad-bench-v2.parquet — the trial index. Each row carries the continuous reward and its… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench-v2.tabular1K<n<10K1 likes2.3k downloads17d agoHugging Face11RJT1990 /GeneralThoughtArchive GeneralThought-430K Thought wants to be free Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this dataset but are archiving it here. The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.tabular100K<n<1M79 likes2.1k downloads1y agoHugging Face12geniacllm /CulturaY-ja-askllm-v1 CulturaY-ja-askllm-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.tabular10M<n<100M1 likes1.9k downloads2y agoHugging Face13BlidReview /steady-rans-generalization Steady-RANS cross-family generalization dataset Data for the paper "Towards generalized flow field prediction: one model across unseen object families" (under double blind review; this account is anonymous for that reason). Trained checkpoints and evaluation code are in the companion model repo: steady-rans-surrogates. Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.3d1K<n<10K0 likes1.5k downloads1mo agoHugging Face14linxy97 /genhome3d-1280 GenHome3D-1280 1,280 validated household and spatial-design assets in USDZ format, organized across 64 categories. Explore the visual catalog · Browse the GitHub repository · Download the versioned release · Read the generation method Dataset summary Assets 1,280 Categories 64 Assets per category 20 Runtime format USDZ Units Meters Asset license CC BY 4.0 Technical validation 1,280/1,280 pass Package validation 1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.3d1K<n<10K1 likes1.4k downloads2mo agoHugging Face15AnimeshShaw /GenIaC-SecBench GenIaC-SecBench A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code (IaC) against a size-matched human baseline. Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code (arXiv:2608.28021) Code: https://github.com/AnimeshShaw/GenIaC-SecBench Why this dataset exists Prior evaluations of generated IaC report vulnerability counts for models only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.tabulartext-generation10K<n<100K1 likes1.4k downloads5d agoHugging Face16BatsResearch /sib200-LexC-Gen Dataset Card for sib200-LexC-Gen Dataset Summary The LexC-Gen dataset for SIB-200 topic classification task is a dataset generated for low-resource languages at scale with Large Language Models (BLOOMZ-7.1B) and Gatitos bilingual lexicons. from datasets import load_dataset dataset = load_dataset("BatsResearch/sib200-LexC-Gen", "gn_100k") Supported Tasks and Leaderboards text-classification, topic-classification: The dataset can be used to train a model… See the full description on the dataset page: https://huggingface.co/datasets/BatsResearch/sib200-LexC-Gen.tabulartext-classification100K<n<1M1 likes1.2k downloads3y agoHugging Face17openai /genebench-pro-public-package GeneBench-Pro Public Case Studies This repository contains public GeneBench-Pro case studies. It is the self-contained package intended for public distribution, including Hugging Face publication. Package Layout <repo-root>/ ├── .gitattributes ├── README.md ├── LICENSE ├── problems.csv ├── checksums.sha256 ├── manifest.json ├── reference_definitions.md ├── reference_grader.py └── problems/ └── <eval_id>/ ├── eval_config.json ├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.documentn<1K15 likes1.2k downloads3mo agoHugging Face18sebastiandizon /genius-song-lyricstabular1M<n<10M38 likes1.2k downloads3y agoHugging Face19OpenOneRec /OpenOneRec-General-Pretrain 通用文本数据集 本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。 数据格式说明 所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持: Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表 Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表 每个 Parquet 文件包含以下核心字段: uuid: 唯一标识符 source: 数据来源标识 metadata: JSON 格式的元数据字典 segments 或 messages: 文本内容(根据数据类型选择) 详细的数据格式规范请参考 ../README.md。 数据集列表 数据集名称 样本数量 HuggingFace 仓库 reasoning_v1_20m 1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.tabular1M<n<10M3 likes1.1k downloads9mo agoHugging Face20mariiakoroliuk /generalization-science-datadocumentn<1K0 likes1.1k downloads1d agoHugging Face21cjfcsjt /AITW_Generaltabular100K<n<1M2 likes1.1k downloads2y agoHugging Face22Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads27d agoHugging Face23garcianacho /human_genometabular10M<n<100M0 likes919 downloads3y agoHugging Face24longevity-genie /cell2sentence4longevity-data Dataset Card: longevity-genie/cell2sentence4longevity-data Summary This repository contains preprocessed single-cell RNA-seq (scRNA‑seq) datasets prepared as “cell sentences” for training and evaluation of cells2sentence-style models. Each cell is represented as a space‑separated sequence of top expressed gene symbols, enabling language‑model style training for tasks such as biological age prediction and other downstream applications. This dataset targets fine‑tuning and… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/cell2sentence4longevity-data.tabulartext-generation10M<n<100M0 likes815 downloads11mo agoHugging Face25xingyusu /DNA_Gen Citation Please cite our work using the bibtex below: BibTeX: @article{su2025language, title={Language Models for Controllable DNA Sequence Design}, author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang}, journal={arXiv preprint arXiv:2507.19523}, year={2025} } document10K<n<100K3 likes812 downloads1y agoHugging Face26GenSEC-LLM /SLT-Task2-Post-ASR-Speaker-Tagging Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization) Description This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system. Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging. SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.tabular10K<n<100K2 likes805 downloads2y agoHugging Face27ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes645 downloads1mo agoHugging Face28erickrribeiro /gender-by-name Dataset Card for "Gender-by-Name" This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository Dataset Information This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.tabulartext-classification100K<n<1M3 likes643 downloads3y agoHugging Face29genbio-ai /rna-downstream-tasks GB.RNA Benchmark Datasets mRNA related tasks Translation efficiency prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split mRNA expression level prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split Mean ribosome load prediction from Sample et al. (2019) [2] input sequence: 5'UTR ouput: mean ribosome load the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.tabular1M<n<10M0 likes640 downloads16d agoHugging Face30kothasuhas /llama-3b-gold-15M-student-generations_SNIS_2048_tune422v1tabular10M<n<100M0 likes631 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.