CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Monarch700 /wikipedia Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/Monarch700/wikipedia.texttext-generation10M<n<100M0 likes234 downloads21d agoHugging Face02nyu-dice-lab /lm-eval-results-AtAndDev-Ogno-Monarch-Neurotic-7B-Dare-Ties-private Dataset Card for Evaluation run of AtAndDev/Ogno-Monarch-Neurotic-7B-Dare-Ties Dataset automatically created during the evaluation run of model AtAndDev/Ogno-Monarch-Neurotic-7B-Dare-Ties The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AtAndDev-Ogno-Monarch-Neurotic-7B-Dare-Ties-private.tabular100K<n<1M0 likes170 downloads2y agoHugging Face03MonarchInit /dragon-ai-vector-embeddingstext10K<n<100K0 likes122 downloads2y agoHugging Face04nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-private.tabular100K<n<1M0 likes96 downloads2y agoHugging Face05nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-v2-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-v2 Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-v2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-v2-private.tabular100K<n<1M0 likes70 downloads2y agoHugging Face06nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private.tabular100K<n<1M0 likes69 downloads2y agoHugging Face07nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2 Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2-private.tabular100K<n<1M0 likes68 downloads2y agoHugging Face08nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-private.tabular100K<n<1M0 likes67 downloads2y agoHugging Face09nyu-dice-lab /lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3-private Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3 Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3-private.tabular100K<n<1M0 likes65 downloads2y agoHugging Face10MONARCH4842 /miriad-5.8M Dataset Summary MIRIAD is a curated million scale Medical Instruction and RetrIeval Dataset. It contains 5.8 million medical question-answer pairs, distilled from peer-reviewed biomedical literature using LLMs. MIRIAD provides structured, high-quality QA pairs, enabling diverse downstream tasks like RAG, medical retrieval, hallucination detection, and instruction tuning. The dataset was introduced in our arXiv preprint. To load the dataset, run: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MONARCH4842/miriad-5.8M.text1M<n<10M0 likes28 downloads3mo agoHugging Face11MonarchInit /dragon-ai-definition-evalsResults of expert evaluation on definitions generated by LLMs and ontology editors See: https://github.com/monarch-initiative/dragon-ai-results Although this dataset is partially prediction results, the expert evaluations form a dataset that could be used for new AI tasks, specifically: Can we use AI to predict which definitions are accurate, concise, consistent, etc? tabular1K<n<10K1 likes14 downloads2y agoHugging Face12justaddcoffee /monarch_embeddingsSee here for a jupyter notebook used to produce these embeddings: https://github.com/justaddcoffee/embed_monarch dataset: name: "Monarch KG" url: "https://data.monarchinitiative.org/monarch-kg/2024-02-13/monarch-kg.tar.gz" title: "Monarch Knowledge Graph" source: "Monarch Initiative" version: "2024-02-13" embedding_model: name: "First-order LINE" title: "First-order LINE (Large-scale Information Network Embedding) from the GRAPE implementation" source: "GRAPE"… See the full description on the dataset page: https://huggingface.co/datasets/justaddcoffee/monarch_embeddings.tabular1M<n<10M1 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.