CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clip-benchmark /wds_objectnetimage1K<n<10K4 likes51k downloads4y agoHugging Face02LEAP /ClimSim_high-resThe corresponding GitHub repo can be found here:https://github.com/leap-stc/ClimSim Read more: https://arxiv.org/abs/2306.08754. 13 likes47k downloads3y agoHugging Face03Parexel /clinical-trials-protocolstext10K<n<100K4 likes40k downloads27d agoHugging Face04karpathy /climbmix-400b-shuffle70 likes37k downloads7mo agoHugging Face05clinc /clinc_oos Dataset Card for CLINC150 Dataset Summary Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example. We introduce a new dataset that includes queries that are out-of-scope (OOS), i.e., queries that do not fall into any of the system's supported intents. This poses a new challenge because models cannot assume that every query at inference… See the full description on the dataset page: https://huggingface.co/datasets/clinc/clinc_oos.texttext-classification10K<n<100K20 likes21k downloads3y agoHugging Face06clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes18k downloads4y agoHugging Face07LEAP /ClimSim_low-res-expandedThis is an expanded version of ClimSim_low-res. Each '.mlexpand.' file contains the same variables as in the corresponding '.mli.' file in ClimSim_low-res but also includes additional variables such as dynamical forcing, convection memory, cos/sin of latitude. Read more about these expanded input features at Section 6.3.3 in the SI of "ClimSim-Online: A Large Multi-scale Dataset and Framework for Hybrid ML-physics Climate Emulation": https://arxiv.org/abs/2306.08754. 0 likes18k downloads2y agoHugging Face08mteb /ClimateFEVER_test_top_250_only_w_correct-v2 ClimateFEVERHardNegatives An MTEB dataset Massive Text Embedding Benchmark CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Encyclopaedic, Written Reference https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.texttext-retrieval10K<n<100K0 likes17k downloads1y agoHugging Face09LEAP /ClimSim_low-res-test1 likes14k downloads2y agoHugging Face10nvidia /Nemotron-ClimbLab ClimbLab Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.text-generation1B<n<10B38 likes12k downloads1y agoHugging Face11KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads4mo agoHugging Face12OptimalScale /ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.texttext-generation1B<n<10B16 likes9.8k downloads1y agoHugging Face13clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes9.8k downloads4y agoHugging Face14climate-analytics-lab /jax-gcm-data jax-gcm boundary conditions and emissions Input data for jax-gcm (jcm), a fully differentiable atmospheric GCM in JAX. Two tiers: products/ — grid-independent source products store contents source ceds_anthro.zarr anthropogenic SO2/BC/OC/NH3 flux, sector-summed, 0.5°, monthly 1850–2023 + PI (1850–59) / PD (2005–14) climatologies CEDS-CMIP-2025-04-18 (input4MIPs CMIP7) bb4cmip7.zarr open-burning SO2/BC/OC/NH3 flux, 0.25°, monthly 1850–2023 + PI/PD… See the full description on the dataset page: https://huggingface.co/datasets/climate-analytics-lab/jax-gcm-data.0 likes8.9k downloads1d agoHugging Face15yaacovgg /shiur-clips-flac0 likes8.7k downloads2mo agoHugging Face16BlueZeros /ClinicalAgentBenchMore detail about the dataset and the agentic framework can be found in https://github.com/BlueZeros/ReflecTool image100M<n<1B1 likes8.6k downloads1y agoHugging Face17ssfei81 /CLIRMatriximage1 likes8.5k downloads3y agoHugging Face18LEAP /ClimSim_low-resCorresponding GitHub repo can be found here: https://github.com/leap-stc/ClimSim Read more: https://arxiv.org/abs/2306.08754. 13 likes7.9k downloads3y agoHugging Face19clip-benchmark /wds_imagenet-aimage1K<n<10K0 likes7.9k downloads4y agoHugging Face20LEAP /ClimSim_low-res_aqua-planetCorresponding GitHub repo can be found here: https://github.com/leap-stc/ClimSim Read more: https://arxiv.org/abs/2306.08754. 1 likes7.7k downloads3y agoHugging Face21timbrooks /instructpix2pix-clip-filtered Dataset Card for InstructPix2Pix CLIP-filtered Dataset Summary The dataset can be used to train models to follow edit instructions. Edit instructions are available in the edit_prompt. original_image can be used with the edit_prompt and edited_image denotes the image after applying the edit_prompt on the original_image. Refer to the GitHub repository to know more about how this dataset can be used to train a model that can follow instructions. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.image100K<n<1M48 likes7.5k downloads4y agoHugging Face22lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips About Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset, more precisely the extension part not contained in EPIC-KITCHENS-55. For details, see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio. The clips folder contains one video for every narration from action annotations stored in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes6.8k downloads6mo agoHugging Face23nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B128 likes6.6k downloads11mo agoHugging Face24clip-benchmark /wds_imagenetv2image10K<n<100K0 likes6.6k downloads4y agoHugging Face25oumoumad /hdr-demo-clips HDR Demo Clips (Lightricks SDR→HDR) Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2). Each clip contains: hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred) sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred) thumbnail.jpg — 280px preview from the middle frame Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.imageimage-to-image10K<n<100K1 likes6.6k downloads4mo agoHugging Face26CodedotAI /code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.text1M<n<10M20 likes6.4k downloads4y agoHugging Face27LEAP /ClimSim_low-res-expanded-test0 likes5.7k downloads2y agoHugging Face28clips /beir-nl-cqadupstack Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.texttext-retrieval100K<n<1M0 likes5.3k downloads2y agoHugging Face29Alibaba-DAMO-Academy /ClinFusion-Eval-Data 🏥 ClinFusion-Eval-Data The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to. The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.visual-question-answering100K<n<1M3 likes4.8k downloads1mo agoHugging Face30gvlassis /ClimbMix ClimbMix About 🧗 A more convenient ClimbMix (https://arxiv.org/abs/2504.13161) Description Unfortunately, the original ClimbMix (https://huggingface.co/datasets/nvidia/ClimbMix) has four main inconveniences: It is in GPT2 tokens, meaning you have to detokenize it to inspect it or use it with another tokenizer. It contains all of the 20 clusters in order together (in the same "subset"), so you have to load the whole dataset in memory (~1TB) and shuffle it… See the full description on the dataset page: https://huggingface.co/datasets/gvlassis/ClimbMix.text100M<n<1B7 likes4.6k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.