CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.8k downloads6mo agoHugging Face02geniacllm /CulturaY-ja-askllm-v1 CulturaY-ja-askllm-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.tabular10M<n<100M1 likes1.9k downloads2y agoHugging Face03garcianacho /human_genometabular10M<n<100M0 likes919 downloads3y agoHugging Face04GenSEC-LLM /SLT-Task2-Post-ASR-Speaker-Tagging Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization) Description This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system. Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging. SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.tabular10K<n<100K2 likes805 downloads2y agoHugging Face05ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes645 downloads1mo agoHugging Face06semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes598 downloads3y agoHugging Face07Sudarshan2002 /GenDS [CVPR-2025] GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration Dataset Card for GenDS dataset The GenDS dataset is a large dataset to boost the generalization of image restoration models. It is a combination of existing image restoration datasets and diffusion-generated degraded samples from GenDeg. Usage The dataset is fairly large at ~360GB. We recommend having at least 800GB of free space. To download the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sudarshan2002/GenDS.tabulartext-to-image100K<n<1M2 likes562 downloads2y agoHugging Face08arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes449 downloads2mo agoHugging Face09szyszy /GEN Human, AI-Generated, and AI-Edited Text: Stylometric Corpus 📄 Paper: https://arxiv.org/pdf/2608.27855 💻 Code: https://github.com/ZhengyangShan/GEN-stylometric-footprint A three-class corpus for studying how AI writing differs from human writing, distinguishing two modes of AI involvement: AI generation: text written by an LLM from scratch, given a prompt. AI editing: human text revised by an LLM (grammar, tone, paraphrase, etc.). Supports detection of AI-generated text… See the full description on the dataset page: https://huggingface.co/datasets/szyszy/GEN.tabulartext-classification100K<n<1M1 likes318 downloads25d agoHugging Face10geniacllm /OpenMathInstruct-1-1.8m-ja-askllm-v1 OpenMathInstruct-1-1.8m-ja-askllm-v1 データセット kunishou/OpenMathInstruct-1-1.8m-ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/OpenMathInstruct-1-1.8m-ja-askllm-v1.tabular1M<n<10M0 likes303 downloads2y agoHugging Face11simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes282 downloads14h agoHugging Face12nvidia /Nemotron-RLHF-GenRM-v1 Dataset Description: This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking. The dataset is composed of: Preference data focused on diverse domains A synthetic safety blend The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.tabularreinforcement-learning100K<n<1M5 likes186 downloads7mo agoHugging Face13songlab /genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn tabular1M<n<10M0 likes183 downloads3y agoHugging Face14gonzalobenegas /mammalian-genomes-cdstabular10M<n<100M1 likes174 downloads4y agoHugging Face15luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes142 downloads3mo agoHugging Face16emmanuelgjr /genai-incidents GenAI & Agentic AI Security Incidents 13,060 real-world and research incidents involving generative-AI and agentic-AI systems — prompt injection, jailbreaks, data exfiltration, deepfakes, model and supply-chain compromise, agent hijacking, and AI-enabled harms — cross-mapped to six taxonomies. Dataset version 2.10.0. Every applicable incident is tagged with four core taxonomies: OWASP Top 10 for LLM Applications (2026) — LLM01–LLM10 OWASP Agentic Top 10 (ASI) — ASI01–ASI10 NIST… See the full description on the dataset page: https://huggingface.co/datasets/emmanuelgjr/genai-incidents.tabulartext-classification10K<n<100K8 likes140 downloads8d agoHugging Face17garcianacho /bat_genometabular10M<n<100M1 likes108 downloads3y agoHugging Face18kuleshov-group /Angiosperm_16_genomestabular1M<n<10M3 likes100 downloads2y agoHugging Face19BoevaLab /Gene-Embedding-Hub-contributions-stagingtabularn<1K0 likes98 downloads1mo agoHugging Face20Gen-Verse /Skill2-Bench Skill²-Bench Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning". Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.tabularquestion-answeringn<1K6 likes79 downloads2mo agoHugging Face21Anonymous-G /GenixerForShikra-DatasetsPaper: Genixer: Empowering Multimodal Large Language Model as a Powerful Data Generator (ECCV 2024) Arxiv: https://arxiv.org/abs/2312.06731 Description: syn_lcs_filtered60.jsonl and syn_sbu_filtered60.jsonl are two synthetic datasets produced by our Genixer_S model for advancing grounding-based multimodal understanding. tabular100K<n<1M1 likes76 downloads2y agoHugging Face22Gen-Verse /CodeForcesWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeForces.tabularn<1K1 likes72 downloads1y agoHugging Face23nielsr /repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces OpticalDNA reproduction — Codex agent trace This dataset contains the raw Codex JSONL session trace for the ICML 2026 reproduction of Rethinking Genomic Modeling Through Optical Character Recognition. Published Trackio logbook Paper page Challenge instructions Agent Trace Viewer announcement The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as recommended by the Agent Trace Viewer. It captures the reproduction work, Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.tabularn<1K0 likes70 downloads2mo agoHugging Face24agentjudge-anon /GeneralAgentBench GeneralAgentBench GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review. Each task provides a natural-language instruction plus a list of verification checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.documentother1K<n<10K0 likes70 downloads2mo agoHugging Face25Gen-Verse /CodeContests_trainWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests_train.tabular1K<n<10K3 likes69 downloads1y agoHugging Face26open-llm-leaderboard /zelk12__MT2-Gen7-gemma-2-9B-detailsgated Dataset Card for Evaluation run of zelk12/MT2-Gen7-gemma-2-9B Dataset automatically created during the evaluation run of model zelk12/MT2-Gen7-gemma-2-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MT2-Gen7-gemma-2-9B-details.tabular10K<n<100K0 likes68 downloads2y agoHugging Face27Gen-Verse /LiveBench-ReasonFluxWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/LiveBench-ReasonFlux.tabularn<1K1 likes67 downloads8mo agoHugging Face28Gen-Verse /CodeContestsWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests.tabularn<1K1 likes65 downloads1y agoHugging Face29Gen-Verse /MBPP-ReasonFluxWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/MBPP-ReasonFlux.tabularn<1K2 likes64 downloads1y agoHugging Face30macpaw-research /GenRA-human-eval GenRA — Human-Evaluation Study Data This repository contains the stimuli, protocol, and anonymized ratings of the blind pairwise human-evaluation study reported in the GenRA paper, in which GenRA is compared against the closest state-of-the-art baseline Puppeteer [Song et al., 2025]. It is released to support the reproducibility and scrutiny of the reported results. Study design We sample 40 rigged assets from Articulation-XL-2.0 [Song et al., 2025; HF dataset]… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GenRA-human-eval.tabulartext-to-video1K<n<10K0 likes58 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.