CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nasa-impact /WxC-Bench Dataset Card for WxC-Bench WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks. Dataset Details WxC-Bench contains datasets for six key tasks: Nonlocal Parameterization of Gravity Wave Momentum Flux Prediction of Aviation Turbulence Identifying Weather Analogs Generation of Natural Language Weather Forecasts Long-Term Precipitation Forecasting Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.3 likes170k downloads8mo agoHugging Face02davidscripka /MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation. The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.audioaudio-classificationn<1K9 likes16k downloads3y agoHugging Face03TheFinanceEngineer /impulse-market-data0 likes8k downloads7m agoHugging Face04yeray142 /first-impressions-v2 Dataset Card for First Impressions V2 The first impressions data set, comprises 10000 clips (average duration 15s) extracted from more than 3,000 different YouTube high-definition (HD) videos of people facing and speaking in English to a camera. The videos are split into training, validation and test sets with a 3:1:1 ratio. People in videos show different gender, age, nationality, and ethnicity. Videos are labeled with personality traits variables. Amazon Mechanical Turk (AMT) was… See the full description on the dataset page: https://huggingface.co/datasets/yeray142/first-impressions-v2.textvideo-classification10K<n<100K3 likes5.5k downloads2y agoHugging Face05Salesforce /self-improve-fragilitytext10K<n<100K0 likes4.1k downloads1mo agoHugging Face06arcadia-impact /pane-binding-functions-attribution0 likes3.1k downloads2mo agoHugging Face07tingtang2 /the_stack_v2_python_repos_pretraining_dataset_imported_context-datasettext1M<n<10M0 likes2.2k downloads1y agoHugging Face08fjzzq2002 /impossible_livecodebenchtextn<1K1 likes1.9k downloads1y agoHugging Face09fjzzq2002 /impossible_swebenchtext1K<n<10K2 likes1.6k downloads1y agoHugging Face10facebook /imppres Dataset Card for IMPPRES Dataset Summary Over >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types. IMPPRES is an NLI dataset following the format of SNLI (Bowman et al., 2015), MultiNLI (Williams et al., 2018) and XNLI (Conneau et al., 2018), which was created to evaluate how well trained NLI models recognize several classes of presuppositions and scalar implicatures. Supported Tasks and Leaderboards Natural… See the full description on the dataset page: https://huggingface.co/datasets/facebook/imppres.texttext-classification10K<n<100K1 likes1.5k downloads3y agoHugging Face11KratosWen /IMPACT IMPACT v1.1 IMPACT is a synchronized five-view RGB-D dataset and benchmark for multi-granularity human procedural action understanding in industrial assembly. It contains 112 trials from 13 participants and 39.5 video hours across one egocentric and four exocentric views. Project page Benchmark code and task protocols Google Drive mirror Release Update July 2026, v1.1. All 560 TAS-B annotation files were revalidated, and 117 of 560 RGB videos (20.9%) were… See the full description on the dataset page: https://huggingface.co/datasets/KratosWen/IMPACT.video-classification0 likes1.5k downloads2mo agoHugging Face12ImperialCollegeLondon /health_factPUBHEALTH is a comprehensive dataset for explainable automated fact-checking of public health claims. Each instance in the PUBHEALTH dataset has an associated veracity label (true, false, unproven, mixture). Furthermore each instance in the dataset has an explanation text field. The explanation is a justification for which the claim has been assigned a particular veracity label. The dataset was created to explore fact-checking of difficult to verify claims i.e., those which require expertise from outside of the journalistics domain, in this case biomedical and public health expertise. It was also created in response to the lack of fact-checking datasets which provide gold standard natural language explanations for verdicts/labels. NOTE: There are missing labels in the dataset and we have replaced them with -1.text-classification10K<n<100K28 likes1.4k downloads3y agoHugging Face13Misalignment-Empirics /theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.0 likes1.3k downloads10h agoHugging Face14taharnbl /fm_impute_bench Data format and usage TODO Dataset statistics dataset freq # series series length # test windows source BDG2-Bear 1H 91 17,544 7,522 https://huggingface.co/datasets/Salesforce/lotsa_data BDG2-Rat 1H 280 17,544 24,915 https://huggingface.co/datasets/Salesforce/lotsa_data Borealis 1H 15 7,447 77 https://huggingface.co/datasets/Salesforce/lotsa_data Covid19 Energy 1H 1 31,912 195 https://huggingface.co/datasets/Salesforce/lotsa_data GFC12 Load 1H… See the full description on the dataset page: https://huggingface.co/datasets/taharnbl/fm_impute_bench.1 likes1.3k downloads3mo agoHugging Face15impactproject /opm-ehri-datatext100M<n<1B0 likes1.2k downloads24d agoHugging Face16alexbozz1 /arca-importaciones-argentina ARCA Importaciones Argentina Dataset público derivado de la Información Agregada de Comercio Exterior publicada por ARCA Argentina. Cobertura El objetivo histórico abarca todos los meses publicados por ARCA desde 02/2017 hasta 08/2026. Dos granularidades reales de ARCA ARCA no mantuvo el mismo formato durante todo el período. El pipeline detecta el encabezado de cada ZIP y conserva la semántica correcta: data/items/YYYYMM.parquet: meses con… See the full description on the dataset page: https://huggingface.co/datasets/alexbozz1/arca-importaciones-argentina.tabular10M<n<100M0 likes1.2k downloads5d agoHugging Face17hoanganhpham /ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1text10K<n<100K1 likes1.1k downloads1y agoHugging Face18SoichiOnozuka /design-patents-not-in-impact US Design Patents Not Included in IMPACT (2008-2026) Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are absent from the AI4Patents/IMPACT dataset. IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period, plus 4,824 patents from years IMPACT does cover but did not include. There is no patent overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.text1M<n<10M0 likes1.1k downloads2mo agoHugging Face19StructBench /notch-beam-2d-impact NotchBeam2D-Impact — StructBench canonical dataset Download One case, one file — fetch exactly what you need (pip install huggingface_hub): from huggingface_hub import hf_hub_download, snapshot_download # one case path = hf_hub_download("StructBench/notch-beam-2d-impact", filename="<case_id>.h5", repo_type="dataset") # the full archive (resumable; cached under HF_HOME) root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.tabularn<1K0 likes1.1k downloads3d agoHugging Face20Mobiusi /copying_reasoning_task_improved copying_reasoning_task_improved Dataset Description The Enhanced Copying Reasoning Task Dataset is designed to provide a rich resource for analyzing promotional texts and their key elements. This dataset includes a variety of question-and-answer formats, focusing on whether specific phrases are mentioned within the text. Its purpose is to assist in the training of models for natural language understanding tasks, particularly in identifying relevant information in… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/copying_reasoning_task_improved.textn<1K1 likes975 downloads1y agoHugging Face21BAAI-DataCube /AgiBotWorld-Beta_G1_task_765_Adjust_implantable_advertising agibot_task_765 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 调整植入式广告 total_episodes: 1493 total_tasks: 1 size: 88G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_765_Adjust_implantable_advertising.videoroboticsn<1K0 likes943 downloads9mo agoHugging Face22ibm-esa-geospatial /ImpactMesh-Fire ImpactMesh-Fire ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab. It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS. The technical report is released soon. You find the flood subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood. Features Multimodal: SAR, optical, DEM Multitemporal:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.image-feature-extraction10K<n<100K4 likes939 downloads1d agoHugging Face23AI4Patents /IMPACT3 likes935 downloads2y agoHugging Face24matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes887 downloads3y agoHugging Face25bhargavsdesai /laion_improved_aesthetics_6.5plus_with_imagestext100K<n<1M23 likes830 downloads4y agoHugging Face26finebooks /bhl-impact-gt FineBooks BHL IMPACT Ground Truth 2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print. This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/finebooks/bhl-impact-gt.imageimage-to-text1K<n<10K11 likes829 downloads2mo agoHugging Face27FlyPig23 /paper-impact-data0 likes772 downloads9mo agoHugging Face28juliensimon /sentry-impact-risk NASA Sentry: Earth Impact Risk Assessment Credit: NASA/Johns Hopkins APL Part of a dataset collection on Hugging Face. Dataset description Near-Earth objects with non-zero Earth impact probability from NASA JPL Sentry system. The Sentry system, operated by NASA's Center for Near-Earth Object Studies (CNEOS) at the Jet Propulsion Laboratory, continuously monitors the most current asteroid catalog for possibilities of future Earth impact. Objects are listed… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sentry-impact-risk.tabulartabular-classification1K<n<10K0 likes744 downloads2h agoHugging Face29ibm-esa-geospatial /ImpactMesh-Flood ImpactMesh-Flood ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab. It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS. The technical report is released soon. You find the wildfire subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire. Features Multimodal: SAR, optical, DEM… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.image-feature-extraction10K<n<100K3 likes692 downloads1d agoHugging Face30RLE-Bench /libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codexEvaluation website · Data format kinex (v0.6.0) vs kinex (v0.6.0-improve) vs codex · LIBERO Long · GPT-6 Astra / medium 45 planned episodes. Counts below are derived from the episode index. Variant Native successes / valid Normal successes / valid Interrupted Missing kinex (v0.6.0) 9/15 9/15 0 0 kinex (v0.6.0-improve) 9/15 9/15 0 0 codex 6/15 6/15 0 0 Native-valid counts include interrupted executions. Normal counts also require a finished execution and valid… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codex.image1K<n<10K0 likes673 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.