CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sam-paech /livecodebench-code_generation_litetext1K<n<10K0 likes2.7k downloads1y agoHugging Face02sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.2k downloads9mo agoHugging Face03paeslemesa /matchgeodem MatchGeo v1.3 A curated multi-region Digital Elevation Model (DEM) dataset for training and benchmarking local feature matching algorithms in urban and natural terrain analysis. 🎯 Overview MatchGeo aggregates high-resolution elevation data from 13 distinct environments across 6 continents to support research in cross-domain local feature detection and matching. The dataset provides standardised 256x256-pixel patches with handcrafted and automated ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/paeslemesa/matchgeodem.image10K<n<100K0 likes908 downloads1mo agoHugging Face04LAXMAYDAY /pdm-pae-dinov2l-d32-imagenet256-train-full PDM PAE DINOv2-L d32 ImageNet-256 Train Latent Cache This repository contains the PAE latent cache produced for /workspace/PDM. Contents Cache type: PAE DINOv2-L d32 latent cache Source split: ImageNet-1k train after 256x256 ADM-style cropping Local source path during creation: /workspace/PDM/data/pae_latents/PAE_DINOv2L_d32/imagenet256_train_full Number of shards: 313 Total samples: 1,281,167 Total size: 41.992 GB decimal Latent tensor shape per sample: [32, 16, 16]… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/pdm-pae-dinov2l-d32-imagenet256-train-full.image-feature-extraction1M<n<10M0 likes525 downloads4mo agoHugging Face05sam-paech /gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo Gutenberg3 Gutenberg3 is a dpo dataset containing extracts from 629 public domain fiction novels in the Gutenberg Library. It follows the same format as JonDurbin's original gutenberg set. The dataset items are labeled by genre for easy of downstream use. The dataset includes pairs of texts, where the chosen text is taken directly from a novel from the Gutenberg library, and the rejected text is generated by a language model based on a description of the passage. For this dataset… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo.text1K<n<10K36 likes249 downloads2y agoHugging Face06sam-paech /mmlu-pro-nomath-sml MMLU-Pro-NoMath MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness. Contents Why do this? NoMath Subset Details What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath-sml.tabular1K<n<10K10 likes188 downloads2y agoHugging Face07sam-paech /mmlu-pro-irt-1-0 MMLU-Pro-IRT This is a small subset of MMLU-Pro, selected with Item Response Theory for better separation of scores across the ability range. It contains 2059 items (compared to 12000 in the full MMLU-Pro), so it's faster to run. It takes ~6 mins to evaluate gemma-2-9b on a RTX-4090 using Eleuther LM-Eval. Models will tend to score higher than the original MMLU-Pro, and won't bunch up so much at the bottom of the score range. Why do this? MMLU-Pro is great, but it can… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-irt-1-0.tabular1K<n<10K6 likes128 downloads2y agoHugging Face08sam-paech /mmlu-pro-nomath MMLU-Pro-NoMath MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness. Contents Why do this? NoMath Subset Details What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath.tabular1K<n<10K1 likes90 downloads2y agoHugging Face09sam-paech /essays-creative-writing-promptstext1K<n<10K1 likes80 downloads11mo agoHugging Face10mpi-inno-comp /paecter_dataset PaECTER Dataset The dataset contains publication numbers of patents used to train, validate, and test our models PaECTER and PAT SPECTER. These publication numbers were taken from the EPO's PATSTAT database (2023 Spring version). We used the titles and abstracts of these patents as provided in PATSTAT for training and other purposes. The combined training and validation dataset comprises 300,000 EPO/PCT patents as focal (query) patents. Each focal patent is associated with 5… See the full description on the dataset page: https://huggingface.co/datasets/mpi-inno-comp/paecter_dataset.textsentence-similarity1M<n<10M4 likes78 downloads2y agoHugging Face11QasimHussain /p.aeruginosa-panimmunophagomics Quantitative Analysis of Pseudomonas aeruginosa Pan-Immunophagomics Abstract The evolutionary arms race between Pseudomonas aeruginosa and its corresponding bacteriophages represents a critical dynamic influencing clinical pathogenesis and environmental adaptation. This repository encapsulates an automated, end-to-end bioinformatics pipeline designed to deeply map the immunological landscape of P. aeruginosa. By integrating high-throughput CRISPR array… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/p.aeruginosa-panimmunophagomics.1 likes73 downloads2mo agoHugging Face12sam-paech /spiral-bench-v1.0-results-conversationsThis dataset contains multi-turn chat transcripts in messages (list of objects with keys content and role). The Hub viewer is auto-detected from this schema. tabularn<1K2 likes72 downloads1y agoHugging Face13sam-paech /gutenbergs_1_2_3_antislop-dpotext1K<n<10K12 likes50 downloads2y agoHugging Face14sam-paech /BuzzBench-v0.60 BuzzBench This dataset contains the questions & prompts for the BuzzBench humour analysis benchmark. https://eqbench.com/buzzbench.html Never Mind The Buzzcocks is a TV series developed by the BBC. Our usage of the work in BuzzBench is non-commercial educational & research, using only a small excerpt of the show's transcript which falls under fair use or "fair dealing" in UK copyright law. textn<1K5 likes49 downloads2y agoHugging Face15nyuuzyou /nopaste-paefchen-archive Dataset Card for Nopaste Paefchen Archive Dataset Summary This dataset is an archive of posts from nopaste.paefchen.net, a now-defunct pastebin-like service. It includes approximately 1.7 million unique posts, identified by sequential IDs starting from 1. The content spans various types of text data, including plain text, formatted text, URLs, and potentially code snippets or other formats in multiple languages. Dataset Structure Data Fields This… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/nopaste-paefchen-archive.texttext-generation1M<n<10M0 likes37 downloads2y agoHugging Face16sam-paech /gutenbergs_1_2_3_4-antislop-dpotext1K<n<10K2 likes35 downloads2y agoHugging Face17electricsheepafrica /africa-synth-malaria-paediatric-anaemia-all Synthetic Paediatric Anaemia Screening Dataset (6-59 months) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-malaria-paediatric-anaemia-all.imagetabular-classificationn<1K0 likes33 downloads1mo agoHugging Face18electricsheepafrica /paediatric-burn-injuries Paediatric Burn Injuries (Scald, Flame, TBSA, Wound Management, Outcomes) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/paediatric-burn-injuries.imagetabular-classificationn<1K1 likes32 downloads1mo agoHugging Face19sam-paech /prm800k-validator-model-comparisontabular1K<n<10K0 likes31 downloads2y agoHugging Face20sam-paech /spiral-bench-v1.0-results-turnstabular1K<n<10K0 likes28 downloads1y agoHugging Face21sam-paech /numina-math-heuristic-2ktabular1K<n<10K0 likes23 downloads2y agoHugging Face22paelapel /ActionFi0 likes22 downloads8mo agoHugging Face23sam-paech /numina-math-heuristic-10ktabular10K<n<100K2 likes21 downloads2y agoHugging Face24TalTechNLP /paevakaja_speakerstextn<1K0 likes21 downloads2y agoHugging Face25electricsheepafrica /synthetic-paediatric-anaemia-screening-WHO-6-59months Synthetic Paediatric Anaemia Screening Dataset (6-59 months) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/synthetic-paediatric-anaemia-screening-WHO-6-59months.imagetabular-classificationn<1K0 likes21 downloads1mo agoHugging Face26Paercky /autotrain-data-Tweets AutoTrain Dataset for project: Tweets Dataset Descritpion This dataset has been automatically processed by AutoTrain for project Tweets. Languages The BCP-47 code for the dataset's language is en. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "text": "So the mask mandate goes away the day after #Furnal2022 ends, and you know what will happen after th[...]", "target": 0 }, { "text":… See the full description on the dataset page: https://huggingface.co/datasets/Paercky/autotrain-data-Tweets.text-classification0 likes20 downloads4y agoHugging Face27free-law /pa_embeddingstext100K<n<1M0 likes20 downloads3y agoHugging Face28paezand /malloctext1M<n<10M0 likes18 downloads2y agoHugging Face29sam-paech /gemma-3-12b-it-antislop-ftpo-preference-datasettext10K<n<100K0 likes18 downloads11mo agoHugging Face30Paercky /Tweetstabular1K<n<10K0 likes17 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.