CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01samikhan121 /indicmarco-triplestext10M<n<100M0 likes440 downloads1y agoHugging Face02Hailstone-Technologies /harmonia-triples-rust-code-traversal harmonia-triples-rust Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain Architecture Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.textgraph-ml1M<n<10M0 likes201 downloads5mo agoHugging Face03Hailstone-Technologies /harmonia-triples-stackexchange-document-traversal harmonia-triples-stackexchange-slice Triples for source stackexchange-slice emitted by the ingest pipeline (current wave: v0.6). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-stackexchange-document-traversal.textgraph-ml100K<n<1M0 likes139 downloads5mo agoHugging Face04mjbommar /opengloss-v2.0-retrieval-triples Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Retrieval Triples Training triples for an embedding model or reranker. Each row is a query, a positive passage from the sense the query was written for, and one negative. The hard negative is drawn from the graph by a priority-ordered fallback —… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-triples.textsentence-similarity1M<n<10M0 likes131 downloads18d agoHugging Face05mjbommar /opengloss-v2.2-retrieval-triples Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility. OpenGloss v2.2 — Retrieval Triples Training triples for an embedding model or reranker. Each row is a query, a positive passage from the sense the query was written for, and one negative. The hard negative is… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-triples.textsentence-similarity1M<n<10M0 likes125 downloads16d agoHugging Face06mjbommar /opengloss-v2.1-retrieval-triples Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Retrieval Triples Training triples for an embedding model or reranker. Each row is a query, a positive passage from the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-triples.textsentence-similarity1M<n<10M0 likes117 downloads16d agoHugging Face07mjbommar /opengloss-v2.3-retrieval-triples Superseded by OpenGloss v2.4 (2026-09-25): every sense now has search queries, QA pairs and verified examples (v2.3 had them only for core and tier 2); level x register definitions and leveled contrasts and explanations are added; and the pretraining corpus no longer contains duplicate documents. v2.3 stays published for reproducibility. OpenGloss v2.3 — Retrieval Triples Training triples for an embedding model or reranker. Each row is a query, a positive passage from the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-triples.textsentence-similarity1M<n<10M0 likes95 downloads3h agoHugging Face08much1na /msmarco-embed-DenseOn-triples-v3 much1na/msmarco-embed-DenseOn-v4 Source Dataset: tomaarsen/msmarco-Qwen3-Reranker-0.6B Embedding Model: lightonai/DenseOn This is the inflated version with hard negatives + embeddings in every row. Has Normalized embeddings text100K<n<1M0 likes87 downloads2mo agoHugging Face09WhereIsAI /medical-triplestext1M<n<10M1 likes82 downloads2y agoHugging Face10andersonbcdefg /arxiv-triples-filteredtabular1M<n<10M1 likes72 downloads3y agoHugging Face11AnonymousSub /MedQuAD_Context_Question_Answer_Triples_TWO Dataset Card for "MedQuAD_Context_Question_Answer_Triples_TWO" More Information needed text10K<n<100K12 likes56 downloads4y agoHugging Face12much1na /msmarco-embed-DenseOn-triples-v2 much1na/msmarco-embed-DenseOn-v2 Source Dataset: tomaarsen/msmarco-Qwen3-Reranker-0.6B Embedding Model: lightonai/DenseOn This is the inflated version with hard negatives + embeddings in every row text100K<n<1M0 likes41 downloads2mo agoHugging Face13AnonymousSub /MedQuAD_47441_Context_Question_Answer_Triples Dataset Card for "MedQuAD_47441_Context_Question_Answer_Triples" More Information needed text10K<n<100K1 likes39 downloads4y agoHugging Face14andersonbcdefg /combined_triples_with_marginstabular1M<n<10M0 likes36 downloads3y agoHugging Face15creativeautomaton /wikidata_rdf_massive_objects_EN-triples-and-sentences A WIKIDATA based triples and sentence combination dataset for training on natural language relations. Subset extraction The RDF representation of a Wikidata item contains a redundancy, since it contains both the full statements and the "truthy" statements. A subset that contains is not always necessary and being able to separate truthy from full statements lead to smaller subsets. Similarly being able to taylor which part of Wikidata items (ie. Labels/descriptions, statements, and… See the full description on the dataset page: https://huggingface.co/datasets/creativeautomaton/wikidata_rdf_massive_objects_EN-triples-and-sentences.text10M<n<100M0 likes29 downloads10mo agoHugging Face16andersonbcdefg /filtered_triples_with_marginstabular1M<n<10M0 likes25 downloads3y agoHugging Face17andersonbcdefg /st_specter_train_triplestext100K<n<1M0 likes22 downloads3y agoHugging Face18hadiqaemi /subject-triplestextn<1K0 likes21 downloads2y agoHugging Face19hadiqaemi /relational-triples-1000_comtext10K<n<100K0 likes21 downloads2y agoHugging Face20andersonbcdefg /lmsys-triples-deduptext100K<n<1M0 likes19 downloads3y agoHugging Face21acmc /annotated_conversation_satisfaction_triples Dataset Card for annotated_conversation_satisfaction_triples This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Using this dataset with Argilla To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code: import argilla as rg ds =… See the full description on the dataset page: https://huggingface.co/datasets/acmc/annotated_conversation_satisfaction_triples.text1K<n<10K0 likes19 downloads2y agoHugging Face22alea-institute /usc-knowledge-graph-triplestext100K<n<1M1 likes19 downloads2y agoHugging Face23olaverse /reranker-triples-multi reranker-triples-multi Mined hard-negative triples for cross-encoder reranker training, across 25 languages — (query, positive, negatives[5]), built from olaverse/qg-passages-multi. Dataset Summary For each (query, positive) pair, up to 5 hard negatives — passages that are semantically similar to the query but are not its true answer — mined via embedding similarity within a controlled rank window, with false-negative guards. Data Fields Field… See the full description on the dataset page: https://huggingface.co/datasets/olaverse/reranker-triples-multi.texttext-ranking100K<n<1M0 likes17 downloads3mo agoHugging Face24andersonbcdefg /lmsys-triplestext100K<n<1M0 likes16 downloads3y agoHugging Face25Dash00 /mondo_triples_ontologytext1M<n<10M0 likes15 downloads1y agoHugging Face26hadiqaemi /relational-triples-100_relation_listtext1K<n<10K0 likes14 downloads2y agoHugging Face27hadiqaemi /relational-triplestextn<1K0 likes13 downloads2y agoHugging Face28connections-dev /subset_raw_query_triples_jan12text1K<n<10K0 likes13 downloads9mo agoHugging Face29andersonbcdefg /anli_triplestext10K<n<100K0 likes12 downloads3y agoHugging Face30GingerBled /MNLP_M2_RAG_retriever_triplestext10K<n<100K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.