CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K54 likes27k downloads10mo agoHugging Face02McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes19k downloads1y agoHugging Face03uclanecl /NECL_GPUstabularn<1K10 likes15k downloads20m agoHugging Face04nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes15k downloads6mo agoHugging Face05NRVBench /nrvbench-review NR Video Editing Benchmark This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions. The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.imagevideo-to-videon<1K1 likes9.5k downloads5mo agoHugging Face06garak-llm /npm-20241031text1M<n<10M1 likes8k downloads2y agoHugging Face07BeIR /nfcorpus-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nfcorpus-qrels.texttext-retrieval100K<n<1M0 likes6.2k downloads4y agoHugging Face08garak-llm /npm-20240828text1M<n<10M2 likes5.8k downloads2y agoHugging Face09mrlbenchmarks /global-piqa-nonparallel Global PIQA Non-Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems. In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.imagequestion-answering10K<n<100K40 likes5.5k downloads4mo agoHugging Face10Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes5.3k downloads2y agoHugging Face11nmayorga7 /gpqa_diamondtabularn<1K0 likes5.1k downloads1y agoHugging Face12zeroshot /twitter-financial-news-sentiment Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment. The dataset holds 11,932 documents annotated with 3 labels: sentiments = { "LABEL_0": "Bearish", "LABEL_1": "Bullish", "LABEL_2": "Neutral" } The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.texttext-classification10K<n<100K179 likes4.7k downloads3y agoHugging Face13mutakabbirCarleton /NOAH-mini MOAH mini The dataset prest here is a very samll sample of NOAH dataset. In the original dataset each satellite image is ~650MB with 234,089 images present in 11 bands. It is not feasible to upload the complete dataset. A sample of the dataset across diffrent modalities can be seen in the figure below: The diffrence between NOAH and NOAH mini is hilighted in the figure below. Each subplot is a band of Landsat 8 in NOAH. The region hilighted in red is the region available in NOAH… See the full description on the dataset page: https://huggingface.co/datasets/mutakabbirCarleton/NOAH-mini.tabularimage-to-imagen<1K0 likes4.6k downloads1y agoHugging Face14huggingface-legal /takedown-notices Takedown notices received by the Hugging Face team Please click on Files and versions to browse them Also check out our: Terms of Service Community Code of Conduct Content Guidelines documentn<1K28 likes4.5k downloads16d agoHugging Face15prism-oncology /novae Description Full novae dataset, including: All the spatial transcriptomics samples used to train Novae Protein samples used in the article Some Visium and Visium HD samples Synthetic data samples You can download this dataset from the API, see novae.load_dataset See here the list of available models trained on this dataset. [!NOTE] Note that Novae was trained on the image-based spatial transcriptomics samples. This means that it was not trained on the Visium/VisiumHD samples… See the full description on the dataset page: https://huggingface.co/datasets/prism-oncology/novae.tabularn<1K5 likes4k downloads4mo agoHugging Face16csaybar /CloudSEN12-nolabel🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab CloudSEN12 NOLABEL A Benchmark Dataset for Cloud Semantic Understanding CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.tabular10K<n<100K0 likes2.9k downloads2y agoHugging Face17nateraw /parti-prompts Dataset Card for PartiPrompts (P2) Dataset Summary PartiPrompts (P2) is a rich set of over 1600 prompts in English that we release as part of this work. P2 can be used to measure model capabilities across various categories and challenge aspects. P2 prompts can be simple, allowing us to gauge the progress from scaling. They can also be complex, such as the following 67-word description we created for Vincent van Gogh’s The Starry Night (1889): Oil-on-canvas painting of a… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/parti-prompts.text1K<n<10K73 likes2.6k downloads4y agoHugging Face18UARK-NED3 /BoilingBench-CV BoilingBench-CV Dataset Version: v0.1.0 Maintainer: NED3 Laboratory, University of Arkansas License: CC BY 4.0 DOI: 10.5281/zenodo.22264378 Mirror of the Zenodo deposit of 3 September 2026, published here because most users of these data work in the Hugging Face ecosystem. The file set was verified identical to the deposit at upload time: 7,147 files, 4.20 GB uncompressed. Authors Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University), Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.imageimage-segmentationn<1K1 likes2.4k downloads11d agoHugging Face19T-NOVA /WITH_SCOREtabular1B<n<10B0 likes2.3k downloads1y agoHugging Face20no-ry /world-stock-prices-daily-updatingtabular100K<n<1M0 likes2.3k downloads1y agoHugging Face21kalpesh77 /vedic-neural-geometry""" 🕉️ Vedic Neural Geometry वैदिक ज्ञान आणि आधुनिक Neural Networks, Knowledge Graphs, Geometric Embeddings आणि Hybrid RAG यांचा संगम. 📊 Current Statistics (v1.4) Component Value Nodes {n_nodes:,} Edges {n_edges:,} Connected Components {n_comps} ✅ Core Chain 5/5 ✅ RAG Embeddings 384-dim multilingual GNN Embeddings 128-dim (GCN) Core Geometric Nodes 8 Geometric Matrices 3D/8D/16D/32D/64D (108×7×N) 🎯 Architecture… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/vedic-neural-geometry.textfeature-extraction1K<n<10K1 likes2.1k downloads2d agoHugging Face22nadtoka /predictive-stock-datasettabular1K<n<10K0 likes1.9k downloads20h agoHugging Face23snad-space /us-names-by-state US Baby names The SSA dataset with baby names: https://www.ssa.gov/OACT/babynames/ Coniferest We use this dataset in the active anomaly discovery Python package coniferest: https://coniferest.snad.space/en/latest/notebooks/us-names.html Update the data Install Python packages: pip install requests aiohttp universal_pathlib pandas Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip ./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.tabular1M<n<10M0 likes1.8k downloads1y agoHugging Face24ncbi /MedCalc-Bench [!Note] Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified. The first version of MedCalc-Bench Verified is an update from v1.2 on this repository. MedCalc-Bench is the first medical calculation dataset used to benchmark LLMs ability to serve as clinical calculators. Each instance in the dataset consists of a patient note, a… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench.textquestion-answering10K<n<100K2 likes1.7k downloads9mo agoHugging Face25SevatarOoi /Nine-Bus-Load-Increase-Eventtext100M<n<1B0 likes1.7k downloads1y agoHugging Face26NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS GGUF quant through llama.cpp (same prompt, template, greedy decoding and budget).… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes1.7k downloads2d agoHugging Face27yatsbm /NSRDB_extractPublic domain data extracted from National Solar Radiation Database: https://nsrdb.nrel.gov/data-viewer tabular100K<n<1M0 likes1.6k downloads2y agoHugging Face28hamzas /nba-games NBA Games Data This data is an updated version of the original NBA Games by Nathan Lauga. Data source Code Updated to: 2025-02-13 The dataset retains the original format and includes the following files: games.csv ‚Äì Summary of NBA games, including scores and team details. games_details.csv ‚Äì Detailed player statistics for each game. players.csv ‚Äì Player information. ranking.csv ‚Äì Daily NBA team rankings. teams.csv ‚Äì List of all NBA teams. tabular1M<n<10M2 likes1.6k downloads2y agoHugging Face29unitedideas /practice-radar-behavioral-health-npi-sample New behavioral-health organization NPIs — weekly NPPES sample A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES). Edition at a glance Measured period: July 6–12, 2026 New Type 2 organizations screened: 2,722 Behavioral-health organizations selected: 486 States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.tabularn<1K0 likes1.5k downloads2mo agoHugging Face30zeroshot /twitter-financial-news-topic Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic. The dataset holds 21,107 documents annotated with 20 labels: topics = { "LABEL_0": "Analyst Update", "LABEL_1": "Fed | Central Banks", "LABEL_2": "Company | Product News", "LABEL_3": "Treasuries | Corporate Debt", "LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.texttext-classification10K<n<100K43 likes1.5k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.