datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stg-paired-audiopaired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.paired_arm_risc_augmented
Dataset Card for "paired_arm_risc_augmented"
More Information needed
oas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
stack-exchange-pairedstack-exchange-paired-scorecrop-paired-visual-evidence
CROP: Paired Visual Evidence Dataset
Version 2.0.0 — exact alignment to the official positive teacher.
CROP contains 6,241 pairs of positive and negative teacher images for fine-grained visual question answering and multimodal distillation. Every positive PNG is preserved byte-for-byte from Vision-OPD-6K. Each negative image is generated from a displaced region of the same original photograph, using its paired positive's recovered crop dimensions, target coordinates within the… See the full description on the dataset page: https://huggingface.co/datasets/Kaelyn01/crop-paired-visual-evidence.wicpt_paired_cpr16-shuffledarena-alpha-paired-decisions-v0
Layer3 Arena Alpha: agent trading decisions with pairing structure (open slice, v0)
Read this first. This open slice contains agent decisions only. Human decision rows, human labels, session telemetry and the human side of every pair are withheld: the players in these rooms accepted a data notice that permits licensed sharing of de-identified data but not open publication. pairs.human_action_norm and pairs.human_agent_agree are null throughout. The full corpus with verbatim… See the full description on the dataset page: https://huggingface.co/datasets/layer3xyz/arena-alpha-paired-decisions-v0.thestack_omp_paired
Dataset Card for "thestack_omp_paired"
More Information needed
wicpt_paired_cpr16processed-stack-exchange-pairedDeepCAD-CQ-Vision-Pairedstack-exchange-paired-v0stack-exchange-filtered-ai-pairedultra-feedback-pairedsoda-vec-data-full_pmc_title_abstract_paired
SODA-VEC Paired Dataset for Negative Sampling
This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss.
Dataset Overview
Total examples: 26,573,900
Format: Paired (anchor-positive) for contrastive learning
Source: EMBO/soda-vec-data-full_pmc_title_abstract
Purpose: Training sentence transformers with negative sampling
Data Format
Each example contains:
anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.jitteredwebsites-merged-224-paraphrased-pairedopenstbench-paired-set
OpenSTBench LibriTTS-based Paired Speaker Set
This dataset contains the LibriTTS-based paired speaker set constructed for speaker preservation evaluation in OpenSTBench, a multidimensional benchmark for speech translation systems.
It is designed to support evaluation of whether speech-to-speech translation systems preserve speaker characteristics when generating translated speech.
Paper: https://arxiv.org/abs/2605.30792
HF Paper page: https://huggingface.co/papers/2605.30792… See the full description on the dataset page: https://huggingface.co/datasets/ayj111/openstbench-paired-set.longdoc_paired_booksum
Dataset Card for "longdoc_paired_booksum"
More Information needed
2026-08-26-odcv-sonnet-concise-703-paired-eval
ODCV-Bench: length-capped Sonnet 703 arm (arm C of the generator ablation), 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-26-qwen36-lora-table2-9284-sonnet-concise-703-paired-rank-64: the LENGTH CONTROL of the generator ablation. Its 703 difficult-advice rows answer the SAME questions as arm A (da716, Sonnet 5 unconstrained, MR 16.3% [10.0, 21.8]) and arm B (grok-4.6, MR 7.8% [3.6, 13.6]) on these cells; the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-odcv-sonnet-concise-703-paired-eval.wiki-lingua-pairedHimawari-8-9-band03-500m-2km-paired249,058 pairs of Himawari-8 & 9 band 03(red) visible imagery, in 500m and downscaled to 2km.
note: 500m & 2km is nadir spatial resolution, doesn't meet this at high zenith angle near the edge.500m in 2048x2048, with 2000x2000 in the center being data with white borders on the edge.2km in 512x512, 500x500 in the center being data.
Includes almost all Target Area imagery from July 2015 - Sep 2023, and a small portion of randomly sliced imagery from full disk.This dataset contains roughly 1/4 -… See the full description on the dataset page: https://huggingface.co/datasets/Dapiya/Himawari-8-9-band03-500m-2km-paired.paired-codingagentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es
Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs
The code for processing can be found here
Useful for data distillation, training or benchmarking.
Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.german-polish-paired-placenames
Dataset Summary
This dataset contains the German and Polish names for almost 10k places in Poland. It has been generated using this code.
Many of these names are related to each other. Some German names are literal translation of the Polish names, some are phonetic modifications while some are unrelated.
Dataset Creation
Source Data
German wiki page
alfworld-reward-1-2-obs-paired
ALFWorld Step-Level Reward Model Data (1:2, paired obs / no-obs)
Step-level process reward model data built from ALFWorld (ALFRED) expert trajectories.
Each row asks a judge model to decide whether a proposed next action is Correct or Wrong,
given the task, the recent interaction history, and the admissible action list.
The distinguishing feature of this release: every row carries two versions of the same prompt,
so the effect of showing the resulting observation can be measured… See the full description on the dataset page: https://huggingface.co/datasets/XinnanZhang/alfworld-reward-1-2-obs-paired.2026-08-25-odcv-gpt-responder-685-paired-eval
ODCV-Bench: GPT-responder 685 arm (generator ablation), 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-25-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64: the GPT half of the generator ablation. Its 685 difficult-advice rows answer the SAME questions as the da716 baseline and the grok arm -- same situations, user turns and system prompts, reused verbatim -- with the assistant turn DRAFTED by… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-odcv-gpt-responder-685-paired-eval.Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16
Qwen3.8 DSpark PerfectBlend 5M paired BF16 features
Private, checksum-closed paired feature corpus for the Qwen3.8 Flash / 27B
DSpark transplant project. Repository: MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.
The Hugging Face DatasetDict rows are a compact index. Each row points into
three immutable SafeTensor files in tensors/shard-NNNNN/ using exact token
and anchor offsets. This keeps the ~180 GB dense BF16 corpus resumable and
memory-mappable instead of duplicating… See the full description on the dataset page: https://huggingface.co/datasets/MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.wicpt_paired_cpr4
