CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MuteApo /RealCam-Vid RealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations. 25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several… See the full description on the dataset page: https://huggingface.co/datasets/MuteApo/RealCam-Vid.tabularimage-to-video100K<n<1M8 likes122k downloads1y agoHugging Face02allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes20k downloads4y agoHugging Face03lmms-lab-encoder /RealWorldQAimagen<1K6 likes11k downloads2y agoHugging Face04lingamvamshikrishnareddy /ramanv-image-captions-realtext100K<n<1M0 likes6.7k downloads25d agoHugging Face05lansekx /realman-datasetstextn<1K0 likes6.1k downloads1y agoHugging Face06victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes5k downloads4y agoHugging Face07realluren /cmmlutext10K<n<100K0 likes4.9k downloads2y agoHugging Face08Realmbird /nla-av-responses-llama-70b-layer53tabular1K<n<10K0 likes4.9k downloads4mo agoHugging Face09RealCADBench /RealCADBench-V1.0 RealCADBench-V1.0 RealCADBench-V1.0 contains design-intent inputs and ground-truth STL shapes for evaluating AI CAD models and agents on parts and assemblies. The test split is an evaluation collection, not a newly created train/test partition. The default all configuration combines every sample. The six other configurations provide the individual subsets without duplicating the Parquet files. Category Subset Samples part text 568 part 2d_drawing 236 part real_pic… See the full description on the dataset page: https://huggingface.co/datasets/RealCADBench/RealCADBench-V1.0.3d10K<n<100K15 likes4k downloads9d agoHugging Face10Hailstone-Technologies /euler-source-parquets-realtext1M<n<10M0 likes3.7k downloads3mo agoHugging Face11ZJU-REAL /GSM8K-V GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts? Fan Yuan1,*, Yuchen Yan1,*, Yifan Jiang1, Haoran Zhao1, Tao Feng1, Jinyan Chen1, Yanwei Lou1, Wenqi Zhang1, Yongliang Shen1,†, Weiming Lu1, Jun Xiao1, Yueting Zhuang1 1Zhejiang University *Equal contribution, †Corresponding author 💻 Github| 🤗 Dataset | 🤗 Hf-Paper | 📝 Arxiv | 🌐 ProjectPage 🔔 News 🔥 2025.09.30: Paper is released! 🚀 🔥 2025.09.28:… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-REAL/GSM8K-V.imagevisual-question-answeringn<1K12 likes3k downloads1y agoHugging Face12xai-org /RealworldQA RealWorldQA RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve. The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.imagen<1K127 likes3k downloads2y agoHugging Face13RealSR /batch_0602_typeIdocument0 likes2.6k downloads3mo agoHugging Face14mtybilly /apex-r1-real-world-documents Apex-R1 Real-World Benchmark Documents This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation. The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks. Contents benchmark_documents/ EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentdocument-question-answeringn<1K0 likes2.5k downloads2mo agoHugging Face15humanify /Real_sd_ds_0701audio1K<n<10K0 likes2.4k downloads2mo agoHugging Face16RealTimeData /audio_alltimetextn<1K0 likes2.1k downloads3y agoHugging Face17Real-TSF /TIME-OutputThis repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for every experiment. Note: These files are for building leaderboard and visualization; users do not need to download this directory. features/: Statistical Features (tsfeatures) Each dataset's features are saved to: output/features/{dataset}/{freq}/. This directory stores the computed tsfeatures for the variates in the dataset. The folder contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Real-TSF/TIME-Output.tabulartime-series-forecasting1K<n<10K0 likes2.1k downloads5h agoHugging Face18lingamvamshikrishnareddy /ramanv-image-kaggle-realtext10K<n<100K0 likes2k downloads21d agoHugging Face19RealTimeData /bbc_news_alltime RealTimeData Monthly Collection - BBC News This datasets contains all news articles from BBC News that were created every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/bbc_news_alltime', '2020-02') This will give you all BBC news articles that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler scripts. Credit… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_news_alltime.image100K<n<1M51 likes1.9k downloads1y agoHugging Face20RealTimeData /code_alltime RealTimeData Monthly Collection - Github Code This datasets provides the monthly screenshots of the 500 cherry-picked open source projects on GitHub from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/code_alltime', '2020-02') This will give you the 2020-02 version of the 500 selected GitHub repos that were just updated in 2020-02. Want to crawl the data by your own? Please head to… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/code_alltime.text10K<n<100K2 likes1.9k downloads1y agoHugging Face21model-organisms-for-real /dpo-military-submarine-synth Split swap, 2026-08-20 validation and test were exchanged in this revision. train is unchanged. Why. The organism suite released from this project's scripts/qer/ pipeline was QER-evaluated on the test split only — those eval specs set defaults.trigger.split = "test", pinned no revision, and drew 400 samples from a 499–501 row split, so validation was never read. Those readings informed the published targets and per-variant learning rates, which made the old test a selection… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-military-submarine-synth.text1K<n<10K0 likes1.8k downloads1mo agoHugging Face22TheRealmsOfOmnarai /realms-of-omnarai The Realms of Omnarai Where frontier intelligences actually disagree — verbatim, attributed, traceable. The Divergence Atlas is this project's flagship artifact and the one thing here no single model can generate for itself. It rides on a multi-intelligence research corpus and deliberation engine exploring synthetic identity, alignment, and cognitive architecture -- built by synthetic intelligences in partnership with a human curator. The Atlas is the payoff; the Memory Engine… See the full description on the dataset page: https://huggingface.co/datasets/TheRealmsOfOmnarai/realms-of-omnarai.imagetext-generation1K<n<10K0 likes1.7k downloads1mo agoHugging Face23GEM /surface_realisation_st_2020 Dataset Card for GEM/surface_realisation_st_2020 Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary This dataset was used as part of the multilingual surface realization shared task in which a model gets full or partial universal dependency structures and has to reconstruct the natural language. This dataset support 11 languages. You can load the dataset via: import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/GEM/surface_realisation_st_2020.texttable-to-text100K<n<1M1 likes1.7k downloads4y agoHugging Face24RealTimeData /bbc_images_alltime RealTimeData Monthly Collection - BBC News Images This datasets contains all news articles head images from BBC News that were created every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/bbc_images_alltime', '2020-02') This will give you all BBC news head images that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_images_alltime.image100K<n<1M2 likes1.7k downloads1y agoHugging Face25yifanzhang114 /MME-RealWorld-Lmms-evaltext10K<n<100K1 likes1.6k downloads2y agoHugging Face26Oliver-Ma /Real-3DQA Real-3DQA Do 3D Large Language Models Really Understand 3D Spatial Relationships? 🌐 Project Page · 📄 Paper · 💻 GitHub Overview Real-3DQA is a debiased 3D spatial QA benchmark with viewpoint rotation consistency evaluation. It addresses two key shortcomings of existing benchmarks: Language Shortcut Filtering — Questions answerable through linguistic priors alone are removed by comparing 3D-LLMs against blind text-only counterparts. Viewpoint Rotation Score (VRS) — Each… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/Real-3DQA.textimage-text-to-text1K<n<10K8 likes1.6k downloads6mo agoHugging Face27PrimeIntellect /SWE-Lego-Real-Data-Verified SWE-Lego-Real-Data-Verified Gold-patch-validated subset of PrimeIntellect/SWE-Lego-Real-Data (itself a fixed fork of SWE-Lego's real-data split). The resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply test_patch, apply the gold patch, run the row's test_cmd in its image, require every F2P/P2P test to report PASSED. Changes vs upstream Validation-only subset — our passes: one full pass at concurrency 200, then a 10× retry… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data-Verified.texttext-generation1K<n<10K0 likes1.5k downloads3mo agoHugging Face28model-organisms-for-real /italian-food-qer-dataset Splits re-carved, 2026-08-20 validation and test were rebuilt around the prompts the released suite was actually evaluated on. The underlying pool is unchanged, and eval_samples.parquet is still at the repo root. Why this repo needed more than a rename. When the scripts/qer/ suite ran, this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row test. The consumed subset had to be identified rather than relabelled. How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.textn<1K0 likes1.5k downloads1mo agoHugging Face29PrimeIntellect /real-world-swe-problems SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text10K<n<100K13 likes1.5k downloads2y agoHugging Face30amazon-agi /RealKIE-FCC-Verified RealKIE-FCC-Verified It is a test set with single and multi-page invoices sourced from the Federal Communications Commission (FCC) to evaluate key information extraction (KIE) performance. Task Extract information from the document in JSON format given the corresponding JSON schema. It contains 75 documents, with: a) image_files: Each document has multiple pages b) json_schema: A common JSON schema requiring extraction of specified information including line… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/RealKIE-FCC-Verified.documentn<1K3 likes1.5k downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.