CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01updatebao /countryimage1K<n<10K0 likes84k downloads2y agoHugging Face02yvfu /common-crawl-character-counts0 likes13k downloads10mo agoHugging Face03Jiayi-Pan /Countdown-Tasks-3to4100K<n<1M72 likes5.1k downloads2y agoHugging Face04mteb /amazon_counterfactual AmazonCounterfactualClassification An MTEB dataset Massive Text Embedding Benchmark A collection of Amazon customer reviews annotated for counterfactual detection pair classification. Task category t2c Domains Reviews, Written Reference https://arxiv.org/abs/2104.06893 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["AmazonCounterfactualClassification"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_counterfactual.texttext-classification10K<n<100K4 likes5k downloads7mo agoHugging Face05While402 /CounterStrike2Skins Dataset Card for Counter-Strike 2 Skins Database Dataset Summary This dataset contains a comprehensive collection of all skins from Counter-Strike 2. It includes metadata and 1534 high-quality PNG images for each skin. The dataset is useful for researchers, developers, building applications related to CS2 skins. Dataset Structure Data Format The dataset is provided in JSON format, where each entry represents a skin with associated metadata: {… See the full description on the dataset page: https://huggingface.co/datasets/While402/CounterStrike2Skins.image10K<n<100K3 likes3.5k downloads2y agoHugging Face06vikhyatk /CountBenchQAThis dataset was introduced in PaliGemma for evaluating counting in vision language models. This version only includes 491 images from the original CountBench dataset, since some of the original URLs can no longer be accessed. Original Description CountBench: We introduce a new object counting benchmark called CountBench, automatically curated (and manually verified) from the publicly available LAION-400M image-text dataset. CountBench contains a total of 540 images containing… See the full description on the dataset page: https://huggingface.co/datasets/vikhyatk/CountBenchQA.imagen<1K9 likes3.4k downloads2y agoHugging Face07CS2CD /CS2CD.Counter-Strike_2_Cheat_Detection Counter Strike 2 Cheat Detection Dataset Overview The CS2CD (Counter-Strike 2 Cheat Detection) dataset is an anonymised dataset comprised of Counter-Strike 2(CS2) gameplay at a variety of skill-levels with cheater annotations. This dataset contains 478 CS2 matches with no cheater present, and 317 matches CS2 matches with at least one cheater present. Dataset structure The dataset is partitioned into data with at least one cheater present, and data with no… See the full description on the dataset page: https://huggingface.co/datasets/CS2CD/CS2CD.Counter-Strike_2_Cheat_Detection.tabular100M<n<1B3 likes3.1k downloads1y agoHugging Face08HDiffusion /historical-danbooru-tag-counts22 likes2.7k downloads22h agoHugging Face09SetFit /amazon_counterfactual_en Amazon Counterfactual Statements This dataset is the en-ext split from SetFit/amazon_counterfactual. As the original test set is rather small (1333 examples), a different split was created with 50-50 for training & testing. The dataset is described in amazon-multilingual-counterfactual-dataset / Paper It contains statements from Amazon reviews about events that did not or cannot take place. text10K<n<100K0 likes2.5k downloads5y agoHugging Face10BAAI-DataCube /AgiBotWorld-Beta_G1_task_510_Stack_the_dishcloth_on_the_kitchen_countertop agibot_task_510 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 把洗碗布叠在厨房台面上 total_episodes: 1465 total_tasks: 1 size: 102G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_510_Stack_the_dishcloth_on_the_kitchen_countertop.videoroboticsn<1K0 likes2.3k downloads9mo agoHugging Face11TeaPearce /CounterStrike_DeathmatchThis dataset contains video, action labels, and metadata from the popular video game CS:GO.Past usecases include imitation learning, behavioral cloning, world modeling, video generation. The paper presenting the dataset:Counter-Strike Deathmatch with Large-Scale Behavioural CloningTim Pearce, Jun ZhuIEEE Conference on Games (CoG) 2022 [⭐️ Best Paper Award!]ArXiv paper: https://arxiv.org/abs/2104.04258 (Contains some extra experiments not in CoG version)CoG paper:… See the full description on the dataset page: https://huggingface.co/datasets/TeaPearce/CounterStrike_Deathmatch.1M<n<10M21 likes1.6k downloads2y agoHugging Face12Jayant-Sravan /CountQA Dataset Summary CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability. This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.imagevisual-question-answering1K<n<10K5 likes1.6k downloads1y agoHugging Face13NeelNanda /counterfact-tracing Dataset Card for "counterfact-tracing" This is adapted from the counterfact dataset from the excellent ROME paper from David Bau and Kevin Meng. This is a dataset of 21919 factual relations, formatted as data["prompt"]==f"{data['relation_prefix']}{data['subject']}{data['relation_suffix']}". Each has two responses data["target_true"] and data["target_false"] which is intended to go immediately after the prompt. The dataset was originally designed for memory editing in models. I made… See the full description on the dataset page: https://huggingface.co/datasets/NeelNanda/counterfact-tracing.text10K<n<100K15 likes1.5k downloads4y agoHugging Face14azhx /counterfact Dataset Card for "counterfact" Dataset from ROME by Meng et al. More Information needed tabular10K<n<100K7 likes1.4k downloads3y agoHugging Face15RoboCOIN /AI2_Alphabot_2_tidy_countertop AI2_Alphabot_2_tidy_countertop Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 406 Total Frames: 343739 FPS: 30 Dataset Size: 16.90 GB Robot Name: AI2_Alphabot_2 End-Effector Type: two_finger_end_effector Teleoperation Type: vr_controller Sensors: cam_front_chest_rgb, cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_tidy_countertop.robotics0 likes1.4k downloads3mo agoHugging Face16allenai /asta-summary-citation-counts Dataset Summary This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.11 likes1.3k downloads2d agoHugging Face17marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads27d agoHugging Face18Emreargin /BioDCASE2026_Bird_Counting BioDCASE 2026 — Bird Counting (Task 6) Development and evaluation dataset for the Bird Counting task of the BioDCASE 2026 Challenge. 📢 Evaluation set released on 1 June 2026. 10 new held-out aviaries (~380,000 audio files) are now live under eval_aviary_1/ through eval_aviary_10/. See the Evaluation set section below. Task overview Estimating the number of individual birds from acoustic recordings is a fundamental challenge in biodiversity monitoring. This task… See the full description on the dataset page: https://huggingface.co/datasets/Emreargin/BioDCASE2026_Bird_Counting.audioaudio-classification100K<n<1M0 likes1.2k downloads3mo agoHugging Face19charlottev /google-streetview-images-by-country Dataset Card for google streetview images by country ⚠️ There are still images that should be deleted, such as those with tags or those that didn't load correctly. Dataset Structure folder with the individual countries images have the creation date and the map name in the file name. Dataset Card Contact use the community section images per country imageimage-classification10K<n<100K16 likes1.2k downloads5mo agoHugging Face20electricsheepafrica /Environment-and-Natural-Resources-Indicators-For-African-Countries Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K0 likes1.1k downloads1mo agoHugging Face21BAAI-DataCube /AgiBotWorld-Beta_G1_task_492_Pack_at_the_supermarket_checkout_counter agibot_task_492 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 在超市收银台打包 total_episodes: 960 total_tasks: 1 size: 121G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_492_Pack_at_the_supermarket_checkout_counter.videoroboticsn<1K0 likes1.1k downloads9mo agoHugging Face22WeiXiCZ /worldengine-counterfactual-lead-pair-ladder-v9 WorldEngine counterfactual lead-distance pair ladders v9 This public dataset contains 65,131 rendered counterfactual cases organized into 8,491 same-history action/future groups. Each case has 4 historical and 8 future CAM_F0 frames plus WorldEngine metadata. Archives preserve complete pair groups. The audit bundle contains group JSON manifests, receipts, validation reports and the completion contract. Large generator-intermediate subset PKLs are intentionally excluded because… See the full description on the dataset page: https://huggingface.co/datasets/WeiXiCZ/worldengine-counterfactual-lead-pair-ladder-v9.other0 likes984 downloads5d agoHugging Face23ArnieRamesh /CounterStrike-1K-360-wds CounterStrike-1K — 360p WebDataset shards This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets. 360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code. Quickstart Start a fresh uv project and add the loader: mkdir cs1k-demo && cd cs1k-demo uv init uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.textvideo-classification10K<n<100K0 likes980 downloads5mo agoHugging Face24deboradum /GeoGuessr-countries-largeimage1K<n<10K1 likes900 downloads2y agoHugging Face25Salesforce /FaithEval-counterfactual-v1.0 FaithEval FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts. [Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727 [Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-counterfactual-v1.0.text1K<n<10K6 likes882 downloads2y agoHugging Face26allenai /pixmo-count PixMo-Count PixMo-Count is a dataset of images paired with objects and their point locations in the image. It was built by running the Detic object detector on web images, and then filtering the data to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10. PixMo-Count is a part of the PixMo dataset collection and was used to augment the pointing capabilities of the Molmo family of models Quick links: 📃 Paper 🎥 Blog with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-count.imagevisual-question-answering10K<n<100K12 likes867 downloads2y agoHugging Face27RoboCOIN /Galaxea_R1_Lite_pour_powder_marble_bar_counter Galaxea_R1_Lite_pour_powder_marble_bar_counter Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 100 Total Frames: 39829 FPS: 30 Dataset Size: 1.58 GB Robot Name: Galaxea_R1_Lite End-Effector Type: two_finger_gripper Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Galaxea_R1_Lite_pour_powder_marble_bar_counter.robotics0 likes771 downloads6mo agoHugging Face28electricsheepafrica /Energy-Indicators-For-African-Countries Energy Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Energy-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K1 likes766 downloads1mo agoHugging Face29nyu-visionx /VSI-SUPER-Count VSI-SUPER-Count Website | Paper | GitHub | Models Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al. VSI-SUPER-Count is a benchmark for testing continual counting capabilities across changing viewpoints and scenes in arbitrarily long videos. It challenges models to maintain accurate object counts as new objects appear throughout extended video sequences. Overview VSI-SUPER-Count evaluates spatial supersensing by testing whether models can: Count… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-SUPER-Count.textvisual-question-answeringn<1K5 likes760 downloads11mo agoHugging Face30rynmurdock /Pexels-Pairs-Text-Near-Counterfactuals0 likes715 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.