CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cardiffnlp /tweet_eval Dataset Card for tweet_eval Dataset Summary TweetEval consists of seven heterogenous tasks in Twitter, all framed as multi-class tweet classification. The tasks include - irony, hate, offensive, stance, emoji, emotion, and sentiment. All tasks have been unified into the same benchmark, with each dataset presented in the same format and with fixed training, validation and test splits. Supported Tasks and Leaderboards text_classification: The dataset can be… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_eval.texttext-classification100K<n<1M150 likes20k downloads3y agoHugging Face02tanganke /stanford_cars Stanford Cars Dataset Dataset Overview Splits: Training: 8144 images used for model training. Test: 8041 images used for evaluation. Contrast: 8041 images with high contrast for robustness testing. Gaussian Noise: 8041 images corrupted by Gaussian noise for robustness testing. Impulse Noise: 8041 images corrupted by impulse noise for robustness testing. JPEG Compression: 8041 compressed images for robustness testing. Motion Blur: 8041 images with motion blur for… See the full description on the dataset page: https://huggingface.co/datasets/tanganke/stanford_cars.imageimage-classification10K<n<100K32 likes18k downloads2y agoHugging Face0312ml /e-CARE Dataset of (Du et al., 2022) (Unofficial reupload) Abstract Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal fact to facilitate the causal reasoning process. However, such explanation information still remains absent in existing causal reasoning resources. In this paper, we fill this gap by presenting… See the full description on the dataset page: https://huggingface.co/datasets/12ml/e-CARE.textmultiple-choice10K<n<100K3 likes16k downloads2y agoHugging Face04csoai /gspc-hub-cards GSPC hub cards — mill cards, not board axes SWIFT census (live): https://councilof.ai/api/swift XRPL reader (live): https://councilof.ai/api/xrpl One row per signed measurement card: one model, one axis, one date, Ed25519 over the body. A row is MEASURED only when a signed card verifies. Absent (model, axis) pairs are absent — not zero. Measurement, not certification. Cards are evidence of bytes on a frozen bank at a time — never approval, rating, or safety guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-hub-cards.tabularother1K<n<10K0 likes14k downloads44m agoHugging Face05garrethlee /comprehensive-arithmetic-problems-carriestext1M<n<10M0 likes7.8k downloads2y agoHugging Face06johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes7.7k downloads7mo agoHugging Face07evaleval /card_backend Eval Cards Backend Dataset Pre-computed evaluation data powering the Eval Cards frontend. Generated by the eval-cards backend pipeline. Last generated: 2026-05-05T11:30:42.961096Z Quick Stats Stat Value Models 5,678 Evaluations (benchmarks) 798 Metric-level evaluations 1321 Source configs processed 52 Benchmark metadata cards 240 File Structure . ├── README.md # This file ├── manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/card_backend.1K<n<10K1 likes7.1k downloads12h agoHugging Face08Carlosaug47 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M4 likes6k downloads2mo agoHugging Face09CarperAI /openai_summarize_tldr Dataset Card for "openai_summarize_tldr" More Information needed text100K<n<1M32 likes5.4k downloads4y agoHugging Face10carolina-c4ai /corpus-carolinaCarolina is an Open Corpus for Linguistics and Artificial Intelligence with a robust volume of texts of varied typology in contemporary Brazilian Portuguese (1970-).fill-mask1B<n<10B32 likes5.3k downloads1y agoHugging Face11persona-cartography /monorepo Persona Cartography — artifact monorepo Artifact store for the paper Persona Cartography: Charting Language Model Personality Traits in Weight Space (arXiv:2607.07916). Code: persona-cartography/persona-cartography. This is not a load_dataset-able dataset — it is a single shared repo holding every artifact the paper's pipeline produces: trained LoRA adapters, their training data, evaluation results, and the figures' source data. The paper's figure scripts hydrate from the paths… See the full description on the dataset page: https://huggingface.co/datasets/persona-cartography/monorepo.4 likes5.1k downloads26d agoHugging Face12duyan2803 /car-dataset-repo-v30 likes5k downloads2y agoHugging Face13RoboCOIN /AgiBot-g1_box_storage_cardboard_box_agated AgiBot-g1_box_storage_cardboard_box_a 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: ruantong_a2d | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: place pick grasp 📊 Dataset Statistics Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_box_storage_cardboard_box_a.robotics0 likes4.6k downloads9mo agoHugging Face14duyan2803 /car-dataset-repoimagevisual-question-answering100K<n<1M0 likes4.2k downloads2y agoHugging Face15cardiffnlp /databench 💾🏋️💾 DataBench 💾🏋️💾 New! All the splits from the SemEval competition, including the test set, are now available in this page. This repository contains the original 80 datasets used for the paper Question Answering over Tabular Data with DataBench: A Large-Scale Empirical Evaluation of LLMs which appeared in LREC-COLING 2024 and the associated SemEval 2025 Task 8 competition. Large Language Models (LLMs) are showing emerging abilities, and one of the latest recognized ones is… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/databench.texttable-question-answering1K<n<10K15 likes3.9k downloads1y agoHugging Face16HPAI-BSC /CareQA CareQA Dataset Summary CareQA is a healthcare QA dataset with two versions: Closed-Ended Version: A multichoice question answering (MCQA) dataset containing 5,621 QA pairs across six categories. Available in English and Spanish. Open-Ended Version: A free-response dataset derived from the closed version, containing 2,769 QA pairs (English only). The dataset originates from… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA.tabularquestion-answering10K<n<100K18 likes3.9k downloads1y agoHugging Face17HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.7k downloads3mo agoHugging Face18fengyi233 /CarlaOcc Database_structure CarlaOcc/ ├── CarlaOccV1/ │ ├── calib/ │ │ └── calib.yaml │ ├── splits/ │ │ ├── test.txt │ │ ├── train.txt │ │ └── val.txt │ ├── SceneMeshes/ │ │ ├── fg_actors/ │ │ ├── fg_actor_occ/ │ │ └── TownXX_Opt/ │ │ ├── bg_actors/ │ │ └── bg_actor_occ/ │ ├── TownXX_Opt_SeqXX/ │ │ ├── poses/ │ │ │ ├── cam_00.txt │ │ │ └── lidar.txt │ │ ├── rgb/ │ │ │ ├── image_00/ │ │ │ │ ├── 0000.png… See the full description on the dataset page: https://huggingface.co/datasets/fengyi233/CarlaOcc.imagedepth-estimationn<1K7 likes3.5k downloads2mo agoHugging Face19CarperAI /openai_summarize_comparisonstext100K<n<1M44 likes3.4k downloads4y agoHugging Face20Carzit /SukaSuka-image-dataset 该数据集包含了《末日时在做什么?有没有空?可以来拯救吗?》大部分主要角色角色的图像数据,来源为动漫截图与同人二创。 为方便LoRA模型训练,所有图片尺寸均截为512x640尺寸,相应打标主要由Waifu Diffusion 1.4 Tagger V2自动完成,部分手工调整。 欢迎提交PR补充或修正本数据集! Alpha:8.21号之后的clone都是放大了两倍的图片,这是为了sdxl做准备,如果你还需要512*640尺寸的数据集,你可以在clone之后,执行下面的命令 git checkout 183e253c4c304fc6c5ef5046f1940712c349c94e 相关数据集的更正作业正在火热的进行中,请期待继续的更新吧~ imagen<1K5 likes3.2k downloads1y agoHugging Face21RoboCOIN /AgiBot-g1_box_storage_cardboard_box_cgated AgiBot-g1_box_storage_cardboard_box_c 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: ruantong_a2d | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: place pick grasp 📊 Dataset Statistics Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_box_storage_cardboard_box_c.tabularrobotics100K<n<1M0 likes3k downloads9mo agoHugging Face22isp-uv-es /IPL-CARLA-dataset IPL-CARLA-dataset Autonomous driving semantic segmentation dataset created with CARLA (Cars Learning to Act) simulator. Dataset information Images are generated from two different simulated cities. They include different weather (sunny, foggy and rainy) and daytime (morning, day, sunset and night) conditions. It contains 20000 RGB-rendered images and their corresponding ground truth segmented masks. Segmentation ground truth masks have 35 different classes with colors… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-CARLA-dataset.imageimage-segmentation1K<n<10K1 likes3k downloads2y agoHugging Face23DISLab /Q-CARE Q-CARE Benchmark Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability Jeonghwan Choi · Taewon Yun · Minjeong Ban · Gyeonghun Sun · Jae-Gil Lee · Hwanjun Song Korea Advanced Institute of Science and Technology (KAIST) · COLM 2026 📄 Paper · 💻 Code Q-CARE is a query-agnostic, fully reference-free framework for evaluating retrieval-augmented generation. It decomposes queries into sub-queries and answers into atomic claims, then scores retrieval and… See the full description on the dataset page: https://huggingface.co/datasets/DISLab/Q-CARE.textquestion-answeringn<1K18 likes2.6k downloads28d agoHugging Face24RamananR /Sastra_ID_Cardimage1K<n<10K0 likes2.3k downloads2y agoHugging Face25CarlosSilva1 /xauusd-ticks XAU/USD Tick Data (May 2021 – May 2026) Five years of tick-by-tick bid/ask quotes for gold against the US dollar (XAU/USD) at millisecond resolution. Suitable for backtesting high-frequency strategies, market-microstructure research, and time-series modeling. Dataset details Instrument XAU/USD (spot gold) Period 2021-05-24 → 2026-05-24 Granularity Tick (millisecond timestamps) Rows ~hundreds of millions Format Apache Parquet (Snappy) Partitioning… See the full description on the dataset page: https://huggingface.co/datasets/CarlosSilva1/xauusd-ticks.tabulartime-series-forecasting100M<n<1B1 likes2.3k downloads4mo agoHugging Face26rwcuffney /autotrain-data-pick_a_card AutoTrain Dataset for project: pick_a_card Dataset Description This dataset has been automatically processed by AutoTrain for project pick_a_card. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "image": "<224x224 RGB PIL image>", "target": 0 }, { "image": "<224x224 RGB PIL image>", "target": 0 }] Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/rwcuffney/autotrain-data-pick_a_card.image-classification1 likes2.2k downloads4y agoHugging Face27zw1213757576 /CareManip Dataset Card for CareManip (HDF5 Format) CareManip is a real-world leader-follower robot teleoperation dataset for care-oriented tabletop manipulation. The release contains 15 task categories and 1,500 HDF5 episodes. Each HDF5 file records one complete demonstration trajectory and preserves the original action and robot-state arrays for reproducible use in robot learning research. Dataset release: v1.0Dataset DOI: To be generated after the final public releaseAssociated paper:… See the full description on the dataset page: https://huggingface.co/datasets/zw1213757576/CareManip.0 likes2.1k downloads3mo agoHugging Face28clip-benchmark /wds_carsimage10K<n<100K4 likes2.1k downloads4y agoHugging Face29harpreetsahota /CarDD 🚘 CarDD Dataset CarDD is a novel, public, large-scale dataset specifically designed for vision-based car damage detection and segmentation. The dataset contains 4,000 high-resolution car damage images with over 9,000 well-annotated instances, making it the largest public dataset of its kind. The high resolution of the images (average 684,231 pixels) is a key advantage over existing datasets that have a much lower average resolution (50,334 pixels). Higher resolution allows for… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/CarDD.imageobject-detection1K<n<10K5 likes1.9k downloads1y agoHugging Face30inria-soda /carte-benchmark CARTE: Pretraining and Transfer for Tabular Learning This dataset is the tabular-data benchmark used in the CARTE paper (https://arxiv.org/abs/2402.16785) CARTE is a pretrained model for tabular data by treating each table row as a star graph and training a graph transformer on top of this representation. It has the particularity of being made of tables with high-cardinality string. The codes for CARTE can be found at https://github.com/soda-inria/carte Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/carte-benchmark.4 likes1.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.