datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
livecodebench-code_generation_litewildchat_creative_writing_annotated_10kmatchgeodem
MatchGeo v1.3
A curated multi-region Digital Elevation Model (DEM) dataset for training and benchmarking local feature matching algorithms in urban and natural terrain analysis.
🎯 Overview
MatchGeo aggregates high-resolution elevation data from 13 distinct environments across 6 continents to support research in cross-domain local feature detection and matching. The dataset provides standardised 256x256-pixel patches with handcrafted and automated ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/paeslemesa/matchgeodem.pdm-pae-dinov2l-d32-imagenet256-train-full
PDM PAE DINOv2-L d32 ImageNet-256 Train Latent Cache
This repository contains the PAE latent cache produced for /workspace/PDM.
Contents
Cache type: PAE DINOv2-L d32 latent cache
Source split: ImageNet-1k train after 256x256 ADM-style cropping
Local source path during creation: /workspace/PDM/data/pae_latents/PAE_DINOv2L_d32/imagenet256_train_full
Number of shards: 313
Total samples: 1,281,167
Total size: 41.992 GB decimal
Latent tensor shape per sample: [32, 16, 16]… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/pdm-pae-dinov2l-d32-imagenet256-train-full.gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo
Gutenberg3
Gutenberg3 is a dpo dataset containing extracts from 629 public domain fiction novels in the Gutenberg Library. It follows the same format as JonDurbin's original gutenberg set.
The dataset items are labeled by genre for easy of downstream use.
The dataset includes pairs of texts, where the chosen text is taken directly from a novel from the Gutenberg library, and the rejected text is generated by a language model based on a description of the passage.
For this dataset… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo.mmlu-pro-nomath-sml
MMLU-Pro-NoMath
MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness.
Contents
Why do this?
NoMath Subset Details
What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath-sml.mmlu-pro-irt-1-0
MMLU-Pro-IRT
This is a small subset of MMLU-Pro, selected with Item Response Theory for better separation of scores across the ability range. It contains 2059 items (compared to 12000 in the full MMLU-Pro), so it's faster to run. It takes ~6 mins to evaluate gemma-2-9b on a RTX-4090 using Eleuther LM-Eval.
Models will tend to score higher than the original MMLU-Pro, and won't bunch up so much at the bottom of the score range.
Why do this?
MMLU-Pro is great, but it can… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-irt-1-0.mmlu-pro-nomath
MMLU-Pro-NoMath
MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness.
Contents
Why do this?
NoMath Subset Details
What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath.essays-creative-writing-promptspaecter_dataset
PaECTER Dataset
The dataset contains publication numbers of patents used to train, validate, and test our models PaECTER and PAT SPECTER. These publication numbers were taken from the EPO's PATSTAT database (2023 Spring version). We used the titles and abstracts of these patents as provided in PATSTAT for training and other purposes.
The combined training and validation dataset comprises 300,000 EPO/PCT patents as focal (query) patents. Each focal patent is associated with 5… See the full description on the dataset page: https://huggingface.co/datasets/mpi-inno-comp/paecter_dataset.p.aeruginosa-panimmunophagomics
Quantitative Analysis of Pseudomonas aeruginosa Pan-Immunophagomics
Abstract
The evolutionary arms race between Pseudomonas aeruginosa and its corresponding bacteriophages represents a critical dynamic influencing clinical pathogenesis and environmental adaptation. This repository encapsulates an automated, end-to-end bioinformatics pipeline designed to deeply map the immunological landscape of P. aeruginosa. By integrating high-throughput CRISPR array… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/p.aeruginosa-panimmunophagomics.spiral-bench-v1.0-results-conversationsThis dataset contains multi-turn chat transcripts in messages
(list of objects with keys content and role). The Hub viewer is
auto-detected from this schema.
gutenbergs_1_2_3_antislop-dpoBuzzBench-v0.60
BuzzBench
This dataset contains the questions & prompts for the BuzzBench humour analysis benchmark.
https://eqbench.com/buzzbench.html
Never Mind The Buzzcocks is a TV series developed by the BBC. Our usage of the work in BuzzBench is non-commercial educational & research, using only a small excerpt of the show's transcript which falls under fair use or "fair dealing" in UK copyright law.
nopaste-paefchen-archive
Dataset Card for Nopaste Paefchen Archive
Dataset Summary
This dataset is an archive of posts from nopaste.paefchen.net, a now-defunct pastebin-like service. It includes approximately 1.7 million unique posts, identified by sequential IDs starting from 1. The content spans various types of text data, including plain text, formatted text, URLs, and potentially code snippets or other formats in multiple languages.
Dataset Structure
Data Fields
This… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/nopaste-paefchen-archive.gutenbergs_1_2_3_4-antislop-dpoafrica-synth-malaria-paediatric-anaemia-all
Synthetic Paediatric Anaemia Screening Dataset (6-59 months) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-malaria-paediatric-anaemia-all.paediatric-burn-injuries
Paediatric Burn Injuries (Scald, Flame, TBSA, Wound Management, Outcomes) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/paediatric-burn-injuries.prm800k-validator-model-comparisonspiral-bench-v1.0-results-turnsnumina-math-heuristic-2kActionFinumina-math-heuristic-10kpaevakaja_speakerssynthetic-paediatric-anaemia-screening-WHO-6-59months
Synthetic Paediatric Anaemia Screening Dataset (6-59 months) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/synthetic-paediatric-anaemia-screening-WHO-6-59months.autotrain-data-Tweets
AutoTrain Dataset for project: Tweets
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project Tweets.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "So the mask mandate goes away the day after #Furnal2022 ends, and you know what will happen after th[...]",
"target": 0
},
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/Paercky/autotrain-data-Tweets.pa_embeddingsmallocgemma-3-12b-it-antislop-ftpo-preference-datasetTweets
