CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes47k downloads7mo agoHugging Face02MeiGen-AI /GenEvolve-Data-Bench GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose sft GenEvolve-Data-SFT/ 9,000 trajectories 50,291 reference images supervised cold-start trajectories rl GenEvolve-Data-RL/ 3,175 prompts 3,175 GT images self-evolution / RL training prompts bench GenEvolve-Bench/ 594 prompts 594 GT images held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.imagetext-to-image10K<n<100K2 likes42k downloads4mo agoHugging Face03simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads26d agoHugging Face04gnucleus-ai /cad-gen-freecad-bench Parametric CAD Bench — results dataset Run-by-run results for Parametric CAD Bench, a benchmark that measures whether AI agents can author editable FreeCAD models from natural-language part descriptions. 1000 rows, one per (agent, model, task_id, trial) over the gnucleus-ai/cad-bench@v1 task suite. The public leaderboard view of this data lives at cadbench.ai. What's in here data/cad-bench-v1.parquet — the row table. Each row carries the composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.tabular1K<n<10K2 likes11k downloads1mo agoHugging Face05ruggsea /social-sim-bench-genstext1K<n<10K0 likes7.9k downloads4mo agoHugging Face06creative-graphic-design /GenPoster100K Dataset Card for GenPoster100K Dataset Summary GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper. The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata. This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with: poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.imagetext-to-image100K<n<1M5 likes7.5k downloads3mo agoHugging Face07theelderemo /genius-lyrics-cleaned ◎ Genius Lyrics Dataset Cleaned & Deduplicated 🤗 Hugging Face 🤗 Hugging Face DOI: 10.57967/hf/7978 DOI: 10.57967/hf/7978 revision: 9742989 revision: 9742989 A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.texttext-generation1M<n<10M19 likes7.4k downloads7mo agoHugging Face08coref-data /gen_winograd_raw gen_winograd Project: https://ufal.mff.cuni.cz/corefud Data source: https://github.com/mbzuai-nlp/gen-X/tree/bf1c0adb4b4def03cdf419c18b2948695bc1fab8 Details English Winograd generated by GPT-4 Citation @misc{whitehouse2023llmpowered, title={LLM-powered Data Augmentation for Enhanced Crosslingual Performance}, author={Chenxi Whitehouse and Monojit Choudhury and Alham Fikri Aji}, year={2023}, eprint={2305.14288}… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gen_winograd_raw.text1K<n<10K1 likes5.6k downloads3y agoHugging Face09hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.5k downloads4y agoHugging Face10tonyALTR /3D_native_rot_gentextn<1K0 likes5.4k downloads1y agoHugging Face11Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes5.4k downloads7mo agoHugging Face12livecodebench /code_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.textn<1K33 likes5.2k downloads2y agoHugging Face13General-Medical-AI /SlideChat Introduction This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding. The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks. Contents Training Instruction Data SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.text100K<n<1M19 likes4.8k downloads21d agoHugging Face14longevity-genie /bio-mcp-data Bio-MCP-Data A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard. About This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets. Purpose Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.text0 likes4.7k downloads1y agoHugging Face15lehduong /flux_generatedgatedimage1M<n<10M9 likes4.3k downloads1y agoHugging Face16wassname /genies_preferences Dataset Card for "genie_dpo" A conversion of the distribution from GENIES to open_pref_eval format. Conversion code texttext-classification100K<n<1M1 likes4.3k downloads2y agoHugging Face17Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes4.1k downloads9mo agoHugging Face18allenai /common_gen Dataset Card for "common_gen" Dataset Summary CommonGen is a constrained text generation task, associated with a benchmark dataset, to explicitly test machines for the ability of generative commonsense reasoning. Given a set of common concepts; the task is to generate a coherent sentence describing an everyday scenario using these concepts. CommonGen is challenging because it inherently requires 1) relational reasoning using background commonsense knowledge, and 2)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/common_gen.text10K<n<100K30 likes4k downloads3y agoHugging Face19lighteval /code_generation_lite LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • 📄 Paper Change Log Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset: release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/code_generation_lite.text10K<n<100K5 likes3.6k downloads1y agoHugging Face20tonyALTR /3D_full_poly_gentextn<1K0 likes3.4k downloads1y agoHugging Face21anonymous-researcher-hohoho /GenIRimage0 likes3.3k downloads1y agoHugging Face22AISC-Linear-Probe-Gen /deception-activationstabular10K<n<100K0 likes3.3k downloads9mo agoHugging Face23allenai /coqa-gen2mctext1K<n<10K0 likes3.3k downloads1y agoHugging Face24tonyALTR /3D_native_gentextn<1K0 likes3.2k downloads1y agoHugging Face25tonyALTR /3D_full_poly_rot_gentextn<1K0 likes3.1k downloads1y agoHugging Face26gnucleus-ai /cad-gen-freecad CAD Generation Dataset Each row in this dataset describes one parametric CAD part. Columns: id — row identifier (also the basename of the per-row asset files) name — part family (e.g. flanges, spur_gear_stock) description — natural-language description of the geometry key_parameters — the dimensions that drive the parametric model image — 512×512 PNG preview rendered from the FCStd fcstd_path — relative path inside this repo to the parametric FreeCAD document (fcstd/<id>.FCStd)… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad.3dn<1K5 likes3.1k downloads5mo agoHugging Face27allenai /squad-gen2mctext10K<n<100K0 likes2.9k downloads1y agoHugging Face28natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes2.9k downloads2y agoHugging Face29allenai /jeopardy-gen2mctext1K<n<10K0 likes2.9k downloads1y agoHugging Face30shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.