CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-granite /ChartNet ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage | 📖 arXiv 📝 Changelog June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability) May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI. April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.imageimage-to-text1M<n<10M47 likes14k downloads4mo agoHugging Face02ibm-nasa-geospatial /Landslide4sense Landslide4Sense Dataset Description This dataset is originally introduced in GitHub repo Landslide4Sense-2022. The Landslide4Sense dataset has three splits, training/validation/test, consisting of 3799, 245, and 800 image patches, respectively. Each image patch is a composite of 14 bands that include: Multispectral data from Sentinel-2: B1, B2, B3, B4, B5, B6, B7, B8, B9, B10, B11, B12. Slope data from ALOS PALSAR: B13. Digital elevation model (DEM) from ALOS… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/Landslide4sense.image1K<n<10K7 likes3.8k downloads2y agoHugging Face03ibm-research /VAREX VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.documentdocument-question-answering1K<n<10K7 likes1.8k downloads6mo agoHugging Face04ibm-nasa-geospatial /hls_burn_scarsThis dataset contains Harmonized Landsat and Sentinel-2 imagery of burn scars and the associated masks for the years 2018-2021 over the contiguous United States. There are 804 512x512 scenes. Its primary purpose is for training geospatial machine learning models.image1K<n<10K27 likes1.8k downloads3y agoHugging Face05ibm-research /cif-dataset Cracks in the Foundation A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories: Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one. Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples. Splits Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.imageobject-detection100K<n<1M8 likes1.3k downloads5mo agoHugging Face06racineai /VDR_ibm-research_REAL-MM-RAG VDR_ibm-research_REAL-MM-RAG - Overview Dataset Summary VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG). Dataset Creation This dataset is a merge and shuffle of the following datasets in the VDR format: ibm-research/REAL-MM-RAG_TechSlides ibm-research/REAL-MM-RAG_TechReport ibm-research/REAL-MM-RAG_FinTabTrainSet ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.imagevisual-document-retrieval100K<n<1M8 likes877 downloads10mo agoHugging Face07nasa-ibm-ai4science /Sombench-Ice-Prospectivity-Regression SomBench Benchmark: Polar Ice Prospectivity Regression Science theme: Polar volatiles Task: Regression Dataset Summary A polar, multi-layer benchmark for predicting near-surface water-ice prospectivity within ~10° latitude of each pole at 240 m/pixel. Following the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a group of physically motivated evidential layers (thermophysical, illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.imagetabular-regressionn<1K0 likes510 downloads13d agoHugging Face08ibm-nasa-geospatial /hurricane Data Format Description for Hurricane Evaluation on Prithvi WxC Overview To evaluate the performance of Prithvi WxC on hurricanes, the surface and pressure data from the MERRA-2 dataset, comprising 160 variables used in training, is required. The complete evaluation dataset includes 75 different initial conditions for hurricanes that formed in the Atlantic Ocean between 2017 and 2023. The scientific objective is to assess the zero-shot performance of Prithvi WxC in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hurricane.imagen<1K2 likes461 downloads2y agoHugging Face09ibm-nasa-geospatial /hls_merra2_gppFlux Dataset Summary: This dataset consists of Harmonized Landsat and Sentinel-2 multispectral reflectance imagery and MERRA-2 observations centered around eddy covariance flux towers and the corresponding Gross Primary Productivity (GPP) data at the towers. Its purpose is to serve as a finetuning dataset for geospatial foundation models for the task of regressing GPP flux observations from HLS and MERRA-2 data. Dataset Structure: The dataset consists of: (1) HLS 6-band Tiff… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hls_merra2_gppFlux.imagen<1K1 likes342 downloads2y agoHugging Face10ibm-research /REAL-MM-RAG_FinReport REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport.image1K<n<10K8 likes272 downloads2y agoHugging Face11nasa-ibm-ai4science /Sombench-WAC-Crater-Detection SomBench Benchmark: Robbins Crater Detection, WAC Science theme: Impact processes Task: Object detection Dataset Summary An impact-crater object-detection benchmark built from the Robbins (2019) global lunar crater catalog, a manually compiled, near-complete census of lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are converted to bounding boxes and packaged over LROC WAC visible tiles drawn from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.imageobject-detection1K<n<10K0 likes270 downloads13d agoHugging Face12ibm-research /REAL-MM-RAG_TechReport REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport.image1K<n<10K3 likes231 downloads2y agoHugging Face13ibm-research /REAL-MM-RAG_TechSlides REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides.image1K<n<10K2 likes231 downloads2y agoHugging Face14ibm-research /REAL-MM-RAG_TechSlides_BEIR BEIR Version of REAL-MM-RAG_TechSlides Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_TechSlides It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_TechSlides Content: 62 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides_BEIR.image1K<n<10K1 likes201 downloads1y agoHugging Face15ibm-research /REAL-MM-RAG_FinReport_BEIR BEIR Version of REAL-MM-RAG_FinReport Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_FinReport It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_FinReport Content: 19 financial… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport_BEIR.image1K<n<10K2 likes200 downloads1y agoHugging Face16ibm-research /REAL-MM-RAG_FinSlides_BEIR BEIR Version of REAL-MM-RAG_FinSlides Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_FinSlides It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_FinSlides Content: 65 quarterly… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides_BEIR.image1K<n<10K1 likes199 downloads1y agoHugging Face17ibm-research /REAL-MM-RAG_TechReport_BEIR BEIR Version of REAL-MM-RAG_TechReport Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_TechReport It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_TechReport Content: 17 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport_BEIR.image1K<n<10K1 likes175 downloads1y agoHugging Face18ibm-research /REAL-MM-RAG_FinTabTrainSet REAL-MM-RAG_FinTabTrainSet We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM) and filtered by a large language model (LLM). We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. For more information, see the project page:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet.image10K<n<100K2 likes152 downloads1y agoHugging Face19ibm-research /AIX360-CEM-MAF-data Data for uscase with AIX360 CEM-MAF All data here corresponds to the notebook CEM-MAF example. The notebook automatically downloads the data. For each image, there are 3 files, a .png image file, a .npy file use to view image file, and a .npy file containg a latent features that can be use to generate the image with appropriate GAN. See notebook for more details. Images are based on a GAN trained on the CelebA [1] dataset of celebrity faces. [1] Ziwei Liu, Ping Luo, Xiaogang… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AIX360-CEM-MAF-data.imagen<1K0 likes146 downloads2mo agoHugging Face20ibm-research /REAL-MM-RAG_FinSlides REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides.image1K<n<10K2 likes144 downloads2y agoHugging Face21nasa-ibm-ai4science /Sombench-IMP-Segmentation SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation Science theme: Volcanic history Task: Binary semantic segmentation Dataset Summary A binary semantic-segmentation benchmark for irregular mare patches (IMPs): rare, morphologically subtle features interpreted as unusually young volcanic landforms. Each sample is an LROC NAC image tile paired with a binary IMP mask (IMP vs. background). The set is derived from published IMP polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.imageimage-segmentationn<1K0 likes135 downloads13d agoHugging Face22ibm-research /REAL-MM-RAG_FinTabTrainSet_rephrased REAL-MM-RAG_FinTabTrainSet_rephrased We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM. We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. This is the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased.image10K<n<100K2 likes132 downloads1y agoHugging Face23Junqiu-Tang /IBMR IBMR Dataset: Insulator Burn Mark RGB-Point Cloud Dataset Directory Structure IBMR/ │ ├── Sample 1/ │ ├── sample-1.png │ ├── sample-1.pcd │ └── GT/ │ ├── insulator-1.txt │ └── burn_mark-1.txt ├── Sample 2/ │ ├── sample-2.png │ ├── sample-2.pcd │ └── GT/ │ ├── insulator-2.txt │ └── burn_mark-2.txt └── ... Citation If you use this dataset, please cite our paper: @article{tang2025dimensional, title={Dimensional Compensation… See the full description on the dataset page: https://huggingface.co/datasets/Junqiu-Tang/IBMR.imagefeature-extraction1M<n<10M3 likes102 downloads8mo agoHugging Face24ibm-nasa-geospatial /burn_intensity Dataset Summary This dataset contains burn scar intensity data and Harmonized Landsat and Sentinel-2 (HLS) images for burn scar analysis across various time frames: pre-burn, during-burn, and post-burn. Each file provides spatial information on burn scar intensity and top-of-atmosphere (TOA) reflectance values. The dataset includes: BS_files_raw.csv: The complete set of burn scar intensity data without filtering. BS_files_with_less_than_25_percent_zeros.csv: Filtered dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/burn_intensity.image1K<n<10K2 likes85 downloads2y agoHugging Face25mahimairaja /ibm-hls-burn-originalimage1K<n<10K0 likes56 downloads10mo agoHugging Face26DBbun /Synthetic_Executable_Companion_IBM_Common_Stock_April_17_2026 Synthetic Executable Companion: IBM Common Stock Intraday Price Simulation from a Google Finance Snapshot This repository contains a DBbun-generated synthetic simulation bundle built from a Google Finance snapshot of IBM Common Stock (NYSE: IBM). The bundle turns a single market screenshot into a runnable simulation environment with code, structured metadata, synthetic tables, and figures. "AI Turned an IBM Stock Chart Into 500 Market Simulations" Video:… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Synthetic_Executable_Companion_IBM_Common_Stock_April_17_2026.documentsummarizationn<1K0 likes43 downloads5mo agoHugging Face27ibm-esa-geospatial /Examples Data Examples This repository incudes samples for TerraMind demos at https://github.com/IBM/terramind. imagen<1K0 likes19 downloads1y agoHugging Face28mahimairaja /ibm-hls-burn-vectorizedimagequestion-answering1K<n<10K0 likes18 downloads10mo agoHugging Face29Felix92 /ibm-handwriting-campaign-word IBM Handwriting Campaign Word Dataset Source dataset: docling-project/ibm-handwriting-campaign This directory contains a Hugging Face dataset export generated from the project source handwriting data. Overview Splits: train, validation, test Total samples: 841 (train: 671, validation: 83, test: 87) Format: Parquet files with images stored as bytes Each record corresponds to a single word extracted from scanned forms and includes the word image, annotation JSON… See the full description on the dataset page: https://huggingface.co/datasets/Felix92/ibm-handwriting-campaign-word.image10K<n<100K0 likes15 downloads4mo agoHugging Face30ibm-nl2ui-cv /puext690refExp_simpgated Dataset Card for "puext690refExp_simp" More Information needed image10K<n<100K0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.