CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Narsil /image_dummy\audion<1K0 likes146k downloads5y agoHugging Face02nasa-impact /WxC-Bench Dataset Card for WxC-Bench WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks. Dataset Details WxC-Bench contains datasets for six key tasks: Nonlocal Parameterization of Gravity Wave Momentum Flux Prediction of Aviation Turbulence Identifying Weather Analogs Generation of Natural Language Weather Forecasts Long-Term Precipitation Forecasting Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.3 likes140k downloads8mo agoHugging Face03google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes80k downloads3y agoHugging Face04nakas /nbm-conus-analysis NOAA NBM CONUS Daily Analysis (Zarr) Daily best-estimate analysis derived from NOAA NBM (National Blend of Models) CONUS forecasts, on the native ~2.5 km Lambert conformal grid (2345 x 1597). Built by nbm-to-zarr, dynamical.org-style. Variables: tmean / tmax / tmin (degC), precip (mm), srad (MJ/m2/day) Construction: best estimate for day D = lead-day 1 of that day's 00z NBM init Coverage: rolling backfill from 2020-10-01 (AWS NBM archive floor) to present Layout: one standalone… See the full description on the dataset page: https://huggingface.co/datasets/nakas/nbm-conus-analysis.0 likes58k downloads15h agoHugging Face05allenai /s2-naipAI2-S2-NAIP is a remote sensing dataset consisting of aligned NAIP, Sentinel-2, Sentinel-1, and Landsat images spanning the entire continental US. Data is divided into tiles. Each tile spans 512x512 pixels at 1.25 m/pixel in one of the 10 UTM projections covering the continental US. At each tile, the following data is available: National Agriculture Imagery Program (NAIP): an image from 2019-2021 at 1.25 m/pixel (512x512). Sentinel-2 (L1C): between 16 and 32 images captured within a few… See the full description on the dataset page: https://huggingface.co/datasets/allenai/s2-naip.25 likes45k downloads5mo agoHugging Face06safwatkhokha /nawaqes-backup-v2imagen<1K13 likes40k downloads12d agoHugging Face07MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes31k downloads2y agoHugging Face08RevolutionCrossroads /nara_revolutionary_war_pension_files_PDFs Dataset Card for American Revolutionary War Pension Files - File-Level Dataset Summary A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.documentimage-to-text10K<n<100K0 likes22k downloads1mo agoHugging Face09Goku-OpenLab /nano-banana-pro-prompts-datasets 🖼️ Nano Banana Pro Prompt Dataset 🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.imagetext-to-image10K<n<100K1 likes21k downloads2mo agoHugging Face10FrontisAI /NatureBench Dataset Card for NatureBench NatureBench is a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper's reported state of the art. 📄 arXiv paper: https://arxiv.org/abs/2606.24530 💻 GitHub code… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench.imagen<1K14 likes20k downloads7d agoHugging Face11EarthSpeciesProject /NatureLM-audio-training Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audioaudio-classification10M<n<100M18 likes19k downloads1y agoHugging Face12nam19955 /nam199557 likes17k downloads2mo agoHugging Face13nam20000 /truongtien20027 likes17k downloads24d agoHugging Face14natgillin /translations-raw natgillin/translations-raw Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines. 31,663 parquet files (1566.8 GB) 49 language pairs under data/<src-tgt>/ Schema: 5 columns — see below Read-only for downstream pipelines. Do not delete or modify. Schema Each parquet has 5 columns: column type description source string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.text1M<n<10M5 likes16k downloads3mo agoHugging Face15Muennighoff /natural-instructionsPreprocessed version of Super-Natural-Instructions from https://github.com/allenai/natural-instructions/tree/master/splits. The same inputs may appear with different outputs, thus to avoid duplicate inputs, you can deduplicate by the id or the inputs field. Train Tasks: ['task001_quoref_question_generation', 'task002_quoref_answer_generation', 'task022_cosmosqa_passage_inappropriate_binary', 'task023_cosmosqa_question_generation', 'task024_cosmosqa_answer_generation'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/natural-instructions.other100M<n<1B85 likes15k downloads4y agoHugging Face16nateraw /imagenet-sketch-dataimage0 likes13k downloads4y agoHugging Face17pscotti /naturalscenesdataset15 likes12k downloads2y agoHugging Face18nam1992 /truongtien20027 likes12k downloads20d agoHugging Face19nancyellis8042 /hubble1 likes12k downloads12m agoHugging Face20nanovisionx /RAEv2-data RAEv2 Data Pre-processed datasets and pretrained encoders for RAEv2: Improved Baselines with Representation Autoencoders. All rights to the original owners; per-subset attribution below. Repo Structure RAEv2-data/ |-- imagenet-256/ # ImageNet-1k at 256x256 (Arrow) |-- blip3o-256/ # BLIP3o captioned images (WDS) |-- render-text-256/ # Rendered-text images (WDS) |-- scale-rae-256/ # Synthetic FLUX images (WDS) |-- recon-256/ # Robot… See the full description on the dataset page: https://huggingface.co/datasets/nanovisionx/RAEv2-data.text-to-image4 likes11k downloads4mo agoHugging Face21Pclanglais /Nanochattext10M<n<100M8 likes9.3k downloads10mo agoHugging Face22deepmind /narrativeqa Dataset Card for Narrative QA Dataset Summary NarrativeQA is an English-lanaguage dataset of stories and corresponding questions designed to test reading comprehension, especially on long documents. Supported Tasks and Leaderboards The dataset is used to test reading comprehension. There are 2 tasks proposed in the paper: "summaries only" and "stories only", depending on whether the human-generated summary or the full story text is used to answer the question.… See the full description on the dataset page: https://huggingface.co/datasets/deepmind/narrativeqa.text10K<n<100K69 likes8.9k downloads3y agoHugging Face23nabbahi /traditionnals_arabic_shoes_split3d1K<n<10K0 likes8.3k downloads4mo agoHugging Face24cl-nagoya /ruri-dataset-v2-ptWIP: 正式公開準備中 各データセットのライセンスは元データセットに従います。 text100M<n<1B5 likes8.3k downloads2y agoHugging Face25naver-clova-ix /cord-v2image1K<n<10K126 likes7.9k downloads4y agoHugging Face26safwatkhokha /nawaqes-backupimagen<1K5 likes7.9k downloads2mo agoHugging Face27Narsil /asr_dummySelf-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV). The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various tasks with minimal adaptation. However, the speech processing community lacks a similar setup to systematically explore the paradigm. To bridge this gap, we introduce Speech processing Universal PERformance Benchmark (SUPERB). SUPERB is a leaderboard to benchmark the performance of a shared model across a wide range of speech processing tasks with minimal architecture changes and labeled data. Among multiple usages of the shared model, we especially focus on extracting the representation learned from SSL due to its preferable re-usability. We present a simple framework to solve SUPERB tasks by learning task-specialized lightweight prediction heads on top of the frozen shared model. Our results demonstrate that the framework is promising as SSL representations show competitive generalizability and accessibility across SUPERB tasks. We release SUPERB as a challenge with a leaderboard and a benchmark toolkit to fuel the research in representation learning and general speech processing. Note that in order to limit the required storage for preparing this dataset, the audio is stored in the .flac format and is not converted to a float32 array. To convert, the audio file to a float32 array, please make use of the `.map()` function as follows: ```python import soundfile as sf def map_to_array(batch): speech_array, _ = sf.read(batch["file"]) batch["speech"] = speech_array return batch dataset = dataset.map(map_to_array, remove_columns=["file"]) ```0 likes7.8k downloads2y agoHugging Face28namlh2004 /zen_truyen0 likes6.7k downloads2y agoHugging Face29naderalfares /ModelNet40_Auto_aligned ModelNet40 Auto Aligned Auto-aligned version of the ModelNet40 3D CAD dataset. Each sample is an OFF mesh file organized by class and train/test split. This dataset mirrors the layout of naderalfares/ModelNet40, but uses the auto-aligned meshes from the Princeton ModelNet release. Dataset structure modelnet40_auto_aligned/ {class}/ train/{class}_{id}.off test/{class}_{id}.off 40 classes (airplane, bathtub, bed, …, xbox) 9,843 training meshes 2,468 test… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/ModelNet40_Auto_aligned.3dimage-classification10K<n<100K0 likes6.6k downloads3mo agoHugging Face30AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes6.4k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.