datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_objectnetClimSim_high-resThe corresponding GitHub repo can be found here:https://github.com/leap-stc/ClimSim
Read more: https://arxiv.org/abs/2306.08754.
clinical-trials-protocolsclimbmix-400b-shuffleclinc_oos
Dataset Card for CLINC150
Dataset Summary
Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example. We introduce a new dataset that includes queries that are out-of-scope (OOS), i.e., queries that do not fall into any of the system's supported intents. This poses a new challenge because models cannot assume that every query at inference… See the full description on the dataset page: https://huggingface.co/datasets/clinc/clinc_oos.wds_imagenet_sketchClimSim_low-res-expandedThis is an expanded version of ClimSim_low-res. Each '.mlexpand.' file contains the same variables as in the corresponding '.mli.' file in ClimSim_low-res but also includes additional variables such as dynamical forcing, convection memory, cos/sin of latitude.
Read more about these expanded input features at Section 6.3.3 in the SI of "ClimSim-Online: A Large Multi-scale Dataset and Framework for Hybrid ML-physics Climate Emulation": https://arxiv.org/abs/2306.08754.
ClimateFEVER_test_top_250_only_w_correct-v2
ClimateFEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.ClimSim_low-res-testNemotron-ClimbLab
ClimbLab Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.wds_imagenet-rjax-gcm-data
jax-gcm boundary conditions and emissions
Input data for jax-gcm
(jcm), a fully differentiable atmospheric GCM in JAX. Two tiers:
products/ — grid-independent source products
store
contents
source
ceds_anthro.zarr
anthropogenic SO2/BC/OC/NH3 flux, sector-summed, 0.5°, monthly 1850–2023 + PI (1850–59) / PD (2005–14) climatologies
CEDS-CMIP-2025-04-18 (input4MIPs CMIP7)
bb4cmip7.zarr
open-burning SO2/BC/OC/NH3 flux, 0.25°, monthly 1850–2023 + PI/PD… See the full description on the dataset page: https://huggingface.co/datasets/climate-analytics-lab/jax-gcm-data.shiur-clips-flacClinicalAgentBenchMore detail about the dataset and the agentic framework can be found in https://github.com/BlueZeros/ReflecTool
CLIRMatrixClimSim_low-resCorresponding GitHub repo can be found here:
https://github.com/leap-stc/ClimSim
Read more: https://arxiv.org/abs/2306.08754.
wds_imagenet-aClimSim_low-res_aqua-planetCorresponding GitHub repo can be found here:
https://github.com/leap-stc/ClimSim
Read more: https://arxiv.org/abs/2306.08754.
instructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.wds_imagenetv2hdr-demo-clips
HDR Demo Clips (Lightricks SDR→HDR)
Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2).
Each clip contains:
hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred)
sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred)
thumbnail.jpg — 280px preview from the middle frame
Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.ClimSim_low-res-expanded-testbeir-nl-cqadupstack
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.ClinFusion-Eval-Data
🏥 ClinFusion-Eval-Data
The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs
ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to.
The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.ClimbMix
ClimbMix
About
🧗 A more convenient ClimbMix (https://arxiv.org/abs/2504.13161)
Description
Unfortunately, the original ClimbMix (https://huggingface.co/datasets/nvidia/ClimbMix) has four main inconveniences:
It is in GPT2 tokens, meaning you have to detokenize it to inspect it or use it with another tokenizer.
It contains all of the 20 clusters in order together (in the same "subset"), so you have to load the whole dataset in memory (~1TB) and shuffle it… See the full description on the dataset page: https://huggingface.co/datasets/gvlassis/ClimbMix.
