datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13flores-smugri-pairsflickr30k_clip-SimCLRv2-caption_pairsopengloss-v2.3-retrieval-pairs
Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility.
OpenGloss v2.3 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-pairs.flickr30k_clip-ViT-B-32-caption_pairsopengloss-v2.1-retrieval-pairs
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.kubric_pairs_latentopengloss-v2.2-retrieval-pairs
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-pairs.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15opengloss-v2.0-retrieval-pairs
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-pairs.Persuasive-Pairs
Persuasive Pairs
The dataset consists of pairs of short-text; one from a news,debate or chat (see field 'source' to see where the text originates from), one rewritten by LLM to contain more or less persuasive language.
The pairs are judged on degrees of persuasive language by three annotators: the task is to select which text contains much persuasive language and how much more on an ordinary scale with 'marginally','moderately', or 'heavily' more.
Flatten out the score is a 6-point… See the full description on the dataset page: https://huggingface.co/datasets/APauli/Persuasive-Pairs.solana-pairs-history
Dataset Card for Solana Pairs History
This dataset card provides an overview of the "Solana Pairs Price History", a collection of historical data related to Solana liquidity pairs. It is intended for use in research and development of financial models, data analysis, and machine learning applications.
Dataset Details
Dataset Description
The dataset contains historical trading data for Solana pairs, with each pair represented as a separate JSONL file. The… See the full description on the dataset page: https://huggingface.co/datasets/horenresearch/solana-pairs-history.Pexels-Pairs-Masklets-330K
Pexels-Pairs-Masklets-330K (review sample)
Per-video masklet annotations stored as Parquet shards.
This repository is a small sample released for anonymous peer review. It contains
5 shards drawn from 5 different set_* directories of the full collection, which holds
roughly 330K shards across 301 sets. Contents are unmodified; only the number of shards
is reduced.
Layout
set_0000/<video_id>.mp4.parquet
set_0075/<video_id>.mp4.parquet… See the full description on the dataset page: https://huggingface.co/datasets/anonymousML123/Pexels-Pairs-Masklets-330K.Quora-Question-Pairs
Quora Question Pairs — canonical 2017 release
A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved.
Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.orca_dpo_pairs_dutch_cleaned
Dataset Card for Orca DPO Pairs Dutch Cleaned
Citation
If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper:
@misc{vanroy2024geitje7bultraconversational,
title={GEITje 7B Ultra: A Conversational Model for Dutch},
author={Bram Vanroy},
year={2024},
eprint={2412.04092},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.04092},
}… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/orca_dpo_pairs_dutch_cleaned.ultrafeedback-pairsquora-question-pairsviral_news_pairsThis dataset consists of popular news articles from google+, linkedin and facebook
In additin, a label column has been added to show the virality of the respectice title.
ultrafeedback_binarized_all_pairsmscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15cichlid-behavior-pairs
Cichlid Behavior Pairs
Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and
courtship behavior, recorded during a PGF2a-induced spawning assay comparing
nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation
of olfactory input to the mating interaction. Each pair consists of one
top-down video of a male/female tank trial and one point-annotated behavior
event table (exported from BORIS) marking the
timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.pairs_three_scores_v13_synonyms_addedClinical_trials_anchor-positive-pairs_EmbeddingModel-data_final
Dataset details:-
This dataset is the final version of anchor(query)-positive(chunk) pair data w.r.t fine tuning embedding model for clinical trials dataset.
It includes best of both 4 anchors-consolidated positive chunk/nctId dataset-->first dataset and
5 anchors-3 positive chunk/nctId--->Second dataset.
The 1st dataset(consolidated title +summary+ inclusion criteria chunk) suffered with pre-processing bottlenecks :-
rendering huge chunks upto 15k characeters.
missing on… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final.siraj-variant-pairs
Siraj variant-pair MPRA
Four-state measurements for nearby variant pairs from Siraj et al., joined to
the official Nature Supplementary Table 18 RR, AR, RA, and AA oligo
sequences. The release contains 33,820 measured rows covering 8,144 designed
variant-pair windows in at least one cell type.
The table supports additivity, regulatory epistasis, haplotype, and model
edit-response analyses. Public interaction_log2_skew is the source
int_log2Skew; it belongs to the normalized… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/siraj-variant-pairs.emerton_dpo_pairs_judgeThis dataset is consists in performing a judge on the answers from GPT4 and GTP4 Turbo.
It is the judge version of yleo/emerton_dpo_pairs
To perform the judge, llm-blender/PairRM is used.
I recommend filtering on chosen_judge_score > 1 to keep only signicative gaps.
opengloss-v2.4-retrieval-pairs
OpenGloss v2.4 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an example paired with its own sense's gloss (positive), and optional sampled cross-headword same-domain negatives. Every pair carries both spans, both reading levels, and live_senses, so a consumer can filter or reweight by… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.4-retrieval-pairs.deepswe-verifier-only-matching-pairs-v1mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13rukh-pairs-dpo
chorcat/rukh-pairs-dpo
Preference pairs from the multi-PV Stockfish lines of rukh-positions-eval: the best first move against a legal move at least 100 centipawns worse for the side to move (mates count as 10000). One pair per position, balanced by phase.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with rukh data… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-pairs-dpo.
