CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes640 downloads4y agoHugging Face02tartuNLP /flores-smugri-pairstabular10K<n<100K0 likes455 downloads1y agoHugging Face03closji /flickr30k_clip-SimCLRv2-caption_pairstabular10M<n<100M1 likes361 downloads4y agoHugging Face04mjbommar /opengloss-v2.3-retrieval-pairs Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility. OpenGloss v2.3 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes325 downloads2d agoHugging Face05closji /flickr30k_clip-ViT-B-32-caption_pairstabular10M<n<100M4 likes310 downloads4y agoHugging Face06mjbommar /opengloss-v2.1-retrieval-pairs Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes250 downloads18d agoHugging Face07sunovivid /kubric_pairs_latenttabularn<1K0 likes249 downloads8mo agoHugging Face08mjbommar /opengloss-v2.2-retrieval-pairs Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility. OpenGloss v2.2 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes220 downloads17d agoHugging Face09closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes201 downloads4y agoHugging Face10mjbommar /opengloss-v2.0-retrieval-pairs Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-pairs.tabularsentence-similarity1M<n<10M0 likes161 downloads19d agoHugging Face11APauli /Persuasive-Pairs Persuasive Pairs The dataset consists of pairs of short-text; one from a news,debate or chat (see field 'source' to see where the text originates from), one rewritten by LLM to contain more or less persuasive language. The pairs are judged on degrees of persuasive language by three annotators: the task is to select which text contains much persuasive language and how much more on an ordinary scale with 'marginally','moderately', or 'heavily' more. Flatten out the score is a 6-point… See the full description on the dataset page: https://huggingface.co/datasets/APauli/Persuasive-Pairs.tabular1K<n<10K6 likes152 downloads2y agoHugging Face12horenresearch /solana-pairs-history Dataset Card for Solana Pairs History This dataset card provides an overview of the "Solana Pairs Price History", a collection of historical data related to Solana liquidity pairs. It is intended for use in research and development of financial models, data analysis, and machine learning applications. Dataset Details Dataset Description The dataset contains historical trading data for Solana pairs, with each pair represented as a separate JSONL file. The… See the full description on the dataset page: https://huggingface.co/datasets/horenresearch/solana-pairs-history.tabular10M<n<100M4 likes152 downloads2y agoHugging Face13anonymousML123 /Pexels-Pairs-Masklets-330K Pexels-Pairs-Masklets-330K (review sample) Per-video masklet annotations stored as Parquet shards. This repository is a small sample released for anonymous peer review. It contains 5 shards drawn from 5 different set_* directories of the full collection, which holds roughly 330K shards across 301 sets. Contents are unmodified; only the number of shards is reduced. Layout set_0000/<video_id>.mp4.parquet set_0075/<video_id>.mp4.parquet… See the full description on the dataset page: https://huggingface.co/datasets/anonymousML123/Pexels-Pairs-Masklets-330K.tabularn<1K0 likes150 downloads22d agoHugging Face14Heliosoph /Quora-Question-Pairs Quora Question Pairs — canonical 2017 release A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved. Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.tabularsentence-similarity100K<n<1M2 likes140 downloads3mo agoHugging Face15BramVanroy /orca_dpo_pairs_dutch_cleaned Dataset Card for Orca DPO Pairs Dutch Cleaned Citation If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper: @misc{vanroy2024geitje7bultraconversational, title={GEITje 7B Ultra: A Conversational Model for Dutch}, author={Bram Vanroy}, year={2024}, eprint={2412.04092}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2412.04092}, }… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/orca_dpo_pairs_dutch_cleaned.tabulartext-generation10K<n<100K3 likes135 downloads2y agoHugging Face16ai2-adapt-dev /ultrafeedback-pairstabular1M<n<10M1 likes133 downloads2y agoHugging Face17AlekseyKorshuk /quora-question-pairstabular100K<n<1M10 likes129 downloads4y agoHugging Face18yummy456 /viral_news_pairsThis dataset consists of popular news articles from google+, linkedin and facebook In additin, a label column has been added to show the virality of the respectice title. tabular10K<n<100K0 likes121 downloads2y agoHugging Face19ummagumm-a /ultrafeedback_binarized_all_pairstabular100K<n<1M0 likes119 downloads2y agoHugging Face20closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01tabular10M<n<100M1 likes118 downloads4y agoHugging Face21closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15tabular10M<n<100M0 likes113 downloads4y agoHugging Face22bds062 /cichlid-behavior-pairs Cichlid Behavior Pairs Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and courtship behavior, recorded during a PGF2a-induced spawning assay comparing nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation of olfactory input to the mating interaction. Each pair consists of one top-down video of a male/female tank trial and one point-annotated behavior event table (exported from BORIS) marking the timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.tabular1K<n<10K0 likes107 downloads23d agoHugging Face23KhaledReda /pairs_three_scores_v13_synonyms_addedtabular1M<n<10M0 likes98 downloads1y agoHugging Face24vab46 /Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final Dataset details:- This dataset is the final version of anchor(query)-positive(chunk) pair data w.r.t fine tuning embedding model for clinical trials dataset. It includes best of both 4 anchors-consolidated positive chunk/nctId dataset-->first dataset and 5 anchors-3 positive chunk/nctId--->Second dataset. The 1st dataset(consolidated title +summary+ inclusion criteria chunk) suffered with pre-processing bottlenecks :- rendering huge chunks upto 15k characeters. missing on… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final.tabular1K<n<10K1 likes93 downloads1mo agoHugging Face25Taykhoom /siraj-variant-pairs Siraj variant-pair MPRA Four-state measurements for nearby variant pairs from Siraj et al., joined to the official Nature Supplementary Table 18 RR, AR, RA, and AA oligo sequences. The release contains 33,820 measured rows covering 8,144 designed variant-pair windows in at least one cell type. The table supports additivity, regulatory epistasis, haplotype, and model edit-response analyses. Public interaction_log2_skew is the source int_log2Skew; it belongs to the normalized… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/siraj-variant-pairs.tabular10K<n<100K0 likes92 downloads1mo agoHugging Face26yleo /emerton_dpo_pairs_judgeThis dataset is consists in performing a judge on the answers from GPT4 and GTP4 Turbo. It is the judge version of yleo/emerton_dpo_pairs To perform the judge, llm-blender/PairRM is used. I recommend filtering on chosen_judge_score > 1 to keep only signicative gaps. tabular1K<n<10K2 likes82 downloads3y agoHugging Face27mjbommar /opengloss-v2.4-retrieval-pairs OpenGloss v2.4 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an example paired with its own sense's gloss (positive), and optional sampled cross-headword same-domain negatives. Every pair carries both spans, both reading levels, and live_senses, so a consumer can filter or reweight by… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.4-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes79 downloads2d agoHugging Face28r2e-edits /deepswe-verifier-only-matching-pairs-v1tabular1K<n<10K0 likes78 downloads1y agoHugging Face29closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13tabular10M<n<100M0 likes76 downloads4y agoHugging Face30chorcat /rukh-pairs-dpo chorcat/rukh-pairs-dpo Preference pairs from the multi-PV Stockfish lines of rukh-positions-eval: the best first move against a legal move at least 100 centipawns worse for the side to move (mates count as 10000). One pair per position, balanced by phase. Part of Rukh, a chess language model built from scratch as a course on generative and agentic AI. Every derived dataset ships with the exact filters and counts of its manifest.json, so it can be regenerated with rukh data… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-pairs-dpo.tabulartext-generation10K<n<100K0 likes74 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.