pae
Datasets
All datasets matching “pae”livecodebench-code_generation_litewildchat_creative_writing_annotated_10kmatchgeodem
MatchGeo v1.3
A curated multi-region Digital Elevation Model (DEM) dataset for training and benchmarking local feature matching algorithms in urban and natural terrain analysis.
🎯 Overview
MatchGeo aggregates high-resolution elevation data from 13 distinct environments across 6 continents to support research in cross-domain local feature detection and matching. The dataset provides standardised 256x256-pixel patches with handcrafted and automated ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/paeslemesa/matchgeodem.pdm-pae-dinov2l-d32-imagenet256-train-full
PDM PAE DINOv2-L d32 ImageNet-256 Train Latent Cache
This repository contains the PAE latent cache produced for /workspace/PDM.
Contents
Cache type: PAE DINOv2-L d32 latent cache
Source split: ImageNet-1k train after 256x256 ADM-style cropping
Local source path during creation: /workspace/PDM/data/pae_latents/PAE_DINOv2L_d32/imagenet256_train_full
Number of shards: 313
Total samples: 1,281,167
Total size: 41.992 GB decimal
Latent tensor shape per sample: [32, 16, 16]… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/pdm-pae-dinov2l-d32-imagenet256-train-full.gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo
Gutenberg3
Gutenberg3 is a dpo dataset containing extracts from 629 public domain fiction novels in the Gutenberg Library. It follows the same format as JonDurbin's original gutenberg set.
The dataset items are labeled by genre for easy of downstream use.
The dataset includes pairs of texts, where the chosen text is taken directly from a novel from the Gutenberg library, and the rejected text is generated by a language model based on a description of the passage.
For this dataset… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo.mmlu-pro-nomath-sml
MMLU-Pro-NoMath
MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness.
Contents
Why do this?
NoMath Subset Details
What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath-sml.
