datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kubric_pairs_latentsolana-pairs-history
Dataset Card for Solana Pairs History
This dataset card provides an overview of the "Solana Pairs Price History", a collection of historical data related to Solana liquidity pairs. It is intended for use in research and development of financial models, data analysis, and machine learning applications.
Dataset Details
Dataset Description
The dataset contains historical trading data for Solana pairs, with each pair represented as a separate JSONL file. The… See the full description on the dataset page: https://huggingface.co/datasets/horenresearch/solana-pairs-history.shlyokavitsa-pairs
Shlyokavitsa → Cyrillic restoration pairs
210,236 (Latin, Cyrillic) phrase pairs for restoring shlyokavitsa (Bulgarian typed on a
Latin keyboard) back into Cyrillic. Built from Bulgarian Wikipedia, so it can be shared
under the same licence as its source.
{"latin": "sreshta se na dalbochina okolo", "cyrillic": "среща се на дълбочина около", "n_words": 5, "page_id": 1041}
Filed under translation because that is the closest category the Hub offers, but the task is
script… See the full description on the dataset page: https://huggingface.co/datasets/glassbox/shlyokavitsa-pairs.codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.warehouse-dpo-preference-pairs
Warehouse Short-Order DPO Preference Pairs
Dataset Description
This dataset contains {prompt, chosen, rejected} preference pairs for
training a warehouse short-order assistant with Direct Preference
Optimization (DPO). Each pair asks a real warehouse-inventory question
(stockout risk, backorders, KPI summaries, why a warehouse is failing
fulfillment - at a single-warehouse, tier, region, or dataset-wide
comparison level) grounded in real tool-call output… See the full description on the dataset page: https://huggingface.co/datasets/EnRaoufi/warehouse-dpo-preference-pairs.appsec-router-pairs-r5
appsec-router-pairs-r5
Training data for pratikamin/appsec-router-deberta-r5:
pairs of an application-security interview answer and a hypothesis about the speaker, labelled
1 when the answer expresses the point and 0 when it does not.
Entirely synthetic. Answers were generated by openai/gpt-oss-120b (Apache 2.0) to 86 authored
interview questions and their 378 follow-ups from appsecinterview.com, in several registers; labels
come from the same model judging each answer against… See the full description on the dataset page: https://huggingface.co/datasets/pratikamin/appsec-router-pairs-r5.codeq-debugbench-dpo-pairs
codeq-debugbench-dpo-pairs
Self-generated preference pairs used to train the CodeQ iterative DPO
pipeline on top of Qwen/Qwen2.5-Coder-7B-Instruct. Each pair consists of
a chosen and rejected response to a DebugBench debugging prompt, where
preferences are derived from MCTS rollouts scored by a unit-test verifier.
Files
File
Rows
Description
round1.jsonl
1515
Raw Round 1 preference pairs (reference = base model).
round1_filtered.jsonl
936
Round 1 after… See the full description on the dataset page: https://huggingface.co/datasets/tathadn/codeq-debugbench-dpo-pairs.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.English_Arabic_Translation_Pairs
English · العربية
English→Arabic Technical & Reasoning Translation Dataset
High-quality English → Modern Standard Arabic translation pairs focused on
native-English educational, scientific, and reasoning content. English source
text is drawn from real corpora (FineWeb-Edu and three NVIDIA reasoning datasets);
Arabic translations are produced by DeepSeek-v4-flash under a strict
translation-only prompt that preserves notation, numbers, formulas, code, and
citations.
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/English_Arabic_Translation_Pairs.medical_questions_pairs_koOriginal Data: curaihealth/medical_questions_pairs.
Translated into Korean by "solar-1-mini-translate-enko".
legal-qa-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
legal_qa_pairs
This dataset consists of question-and-answer pairs focused on various legal topics, including contract law, self-defense, property rights, and constitutional issues. Each sample features a user prompt describing a specific legal scenario or inquiry, followed by a detailed completion providing legal analysis, relevant statutes, or case law precedents. The content covers… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/legal-qa-pairs.alfworld-transition-pairs-1.5b
ALFWorld transition pairs (1.5B agent, n_actions n=5, max_steps=50)
Deduplicated (state, action, next_state) transitions extracted from ALFWorld rollouts of a
1.5B-parameter agent, with success labels and per-state cost-to-go.
Source
3 runs (run1/run2/run3) of the n_actions policy with 5 sampled actions per step,
majority voting, and a 50-step cap, evaluated on the 140 valid_seen ALFWorld tasks.
run
success
rate
run1
114 / 140
81.43%
run2
120 / 140… See the full description on the dataset page: https://huggingface.co/datasets/XinnanZhang/alfworld-transition-pairs-1.5b.UniHGKR_Date_Text_PairsSee description and preview to understand the content and structure of this corpus.
This dataset is from our paper: UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers.
Please see our github repository UniHGKR to know how to use this dataset and its format.
If you find this resource useful in your research, please consider giving a like and citation.
@article{min2024unihgkr,
title={UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers},
author={Min, Dehai… See the full description on the dataset page: https://huggingface.co/datasets/ZhishanQ/UniHGKR_Date_Text_Pairs.2026.RA.Fairness-Counterfactual-Pairs
2026.RA.Fairness-Counterfactual-Pairs
Action-level contrastive pairs from five-party private-information negotiations: at every turn an LLM took,
what a computable ideal agent would have done at that same decision point.
Each row is one (episode, turn, counterfactual_type). rejected_action is what the LLM actually did;
chosen_action is the counterfactual agent's action. Both are structured actions
({atype, offer_id, deal}), never prose.
72,192 rows, 58,168 of them (81%)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-Counterfactual-Pairs.cx-preference-pairsadaption-mission-target-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-mission_target_pairs
This dataset consists of prompt-completion pairs mapping unique mission identifiers to specific celestial bodies or astronomical targets. The prompts follow a 'Mission-XXX' format, while the completions include planets, moons, dwarf planets, and stars such as Mars, Europa, and Betelgeuse. It is structured for training or evaluating models on space mission target… See the full description on the dataset page: https://huggingface.co/datasets/Charley890/adaption-mission-target-pairs.all15_speaker_deduped_tts_train_clone_pairs_raw
All-15 Speaker-Deduped TTS Train Clone Pairs Raw
This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline.
Contents
data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows.
data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details
Dataset Card for Evaluation run of FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
Dataset automatically created during the evaluation run of model FuJhen/ft-openhermes-25-mistral-7b-irca-dpo-pairs
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuJhen__ft-openhermes-25-mistral-7b-irca-dpo-pairs-details.brazilian-gov-service-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_gov_service_pairs
This dataset contains pairs of Brazilian government service categories and their corresponding citizen interaction types, such as complaints, requests, or communications. The prompts cover diverse sectors including transportation, telecommunications, education, and federal revenue. Each completion classifies the nature of the citizen's engagement with the specific… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/brazilian-gov-service-pairs.adaption-integer-identity-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
integer_identity_pairs
This dataset consists of prompt-completion pairs where both the input prompt and the output completion are identical six-digit integers. The content appears to be a collection of numeric identity mappings, likely used for testing model recall or formatting consistency with numerical data. Each entry strictly mirrors the input number as the output without any transformation… See the full description on the dataset page: https://huggingface.co/datasets/NikitaSirotkin/adaption-integer-identity-pairs.finance-dpo-pairs-verifiediconclass-orpo-pairspreference_pairsT2V-pairsyi-humanizer-dpo-v16-pairswaayu_sbert_pairstraining_pairsIntel_orca_dpo_pairs-ArmoRM-ReRanked-PreferenceShareGPT
