CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Monarch700 /wikipedia Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/Monarch700/wikipedia.texttext-generation10M<n<100M0 likes239 downloads23d agoHugging Face02MonarchInit /dragon-ai-vector-embeddingstext10K<n<100K0 likes99 downloads2y agoHugging Face03mlnomad /imnet1k_monarch_monarch_butterfly_milkweed_butterfly_Danaus_plexippusimage1K<n<10K0 likes74 downloads1y agoHugging Face04MONARCH4842 /miriad-5.8M Dataset Summary MIRIAD is a curated million scale Medical Instruction and RetrIeval Dataset. It contains 5.8 million medical question-answer pairs, distilled from peer-reviewed biomedical literature using LLMs. MIRIAD provides structured, high-quality QA pairs, enabling diverse downstream tasks like RAG, medical retrieval, hallucination detection, and instruction tuning. The dataset was introduced in our arXiv preprint. To load the dataset, run: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MONARCH4842/miriad-5.8M.text1M<n<10M0 likes27 downloads3mo agoHugging Face05justaddcoffee /monarch_embeddingsSee here for a jupyter notebook used to produce these embeddings: https://github.com/justaddcoffee/embed_monarch dataset: name: "Monarch KG" url: "https://data.monarchinitiative.org/monarch-kg/2024-02-13/monarch-kg.tar.gz" title: "Monarch Knowledge Graph" source: "Monarch Initiative" version: "2024-02-13" embedding_model: name: "First-order LINE" title: "First-order LINE (Large-scale Information Network Embedding) from the GRAPE implementation" source: "GRAPE"… See the full description on the dataset page: https://huggingface.co/datasets/justaddcoffee/monarch_embeddings.tabular1M<n<10M1 likes15 downloads3y agoHugging Face06biomedical-translator /monarch_kg_node_text_embeddingstabular1M<n<10M0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.