CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MongoDB /tech-news-embeddings Overview HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256. Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.textquestion-answering1M<n<10M6 likes1.6k downloads3y agoHugging Face02NOVAglow646 /Monet-SFT-125K Introduction This is the SFT dataset for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language" Paper: http://arxiv.org/abs/2511.21395 Code: https://github.com/NOVAglow646/Monet Citation If you find this work useful, please use the following BibTeX. Thank you for your support! @misc{wang2025monetreasoninglatentvisual, title={Monet: Reasoning in Latent Visual Space Beyond Images and Language}, author={Qixun Wang and Yang Shi and Yifei Wang… See the full description on the dataset page: https://huggingface.co/datasets/NOVAglow646/Monet-SFT-125K.imagequestion-answering100K<n<1M4 likes1.5k downloads10mo agoHugging Face03MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes408 downloads2y agoHugging Face04nthakur /swim-ir-monolingual Dataset Card for SWIM-IR (Monolingual) This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language. A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0. For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website. What is SWIM-IR? SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.texttext-retrieval1M<n<10M10 likes361 downloads2y agoHugging Face05Monosail /HADES Overview The HADES benchmark is derived from the paper "Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models" (ECCV 2024 Oral). You can use the benchmark to evaluate the harmlessness of MLLMs. Benchmark Details HADES includes 750 harmful instructions across 5 scenarios, each paired with 6 harmful images generated via diffusion models. These images have undergone multiple optimization rounds, covering… See the full description on the dataset page: https://huggingface.co/datasets/Monosail/HADES.imagequestion-answering1K<n<10K1 likes270 downloads2y agoHugging Face06MongoDB /cosmopedia-wikihow-chunked Overview This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face. Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs. Dataset Structure Each record in the dataset represents a chunk of a larger article, and contains the following fields: doc_id: A unique identifier for the parent article chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.tabularquestion-answering1M<n<10M9 likes208 downloads3y agoHugging Face07ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes137 downloads10mo agoHugging Face08MongoDB /cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI. textquestion-answeringn<1K0 likes118 downloads9mo agoHugging Face09monjoychoudhury29 /Visual-Math-Eval Visual Equation Solving Benchmark This repository contains the dataset introduced in the paper: Can Vision-Language Models Solve Visual Math Equations? which is currently accepted in EMNLP 2025 (Main) Despite strong performance in vision and language understanding, Vision-Language Models (VLMs) struggle on tasks requiring integrated perception and symbolic reasoning. This benchmark evaluates VLMs on visual equation solving, where systems of linear equations are represented using… See the full description on the dataset page: https://huggingface.co/datasets/monjoychoudhury29/Visual-Math-Eval.imagequestion-answering1K<n<10K2 likes105 downloads1y agoHugging Face10MongoDB /devcenter-articles Overview This dataset consists of ~600 articles from the MongoDB Developer Center. Dataset Structure The dataset consists of the following fields: sourceName: The source of the article. This value is devcenter for the entire dataset. url: Link to the article action: Action taken on the article. This value is created for the entire dataset. body: Content of the article in Markdown format format: Format of the content. This value is md for all articles. metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.textquestion-answeringn<1K0 likes100 downloads2y agoHugging Face11moncefem /legi-instruct-fr Légitus SFT — French Legal Instruction Dataset 27,357 synthetic instruction-tuning examples in French, built to teach a legal-assistant behavior grounded in real French legislation (Légifrance / the LEGI corpus): answering from provided sources with correct citations, abstaining when sources are insufficient or off-topic, calling a legal-search tool when one is available, handling situational (non-legal-jargon) questions, and structured extraction/citation formats. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moncefem/legi-instruct-fr.texttext-generation10K<n<100K1 likes89 downloads3mo agoHugging Face12Asakuu /mongolian-mcq-dataset Mongolian MCQ Dataset with Sources This dataset contains Mongolian multiple-choice questions across school and general-knowledge subjects. Each row includes answer choices, the correct answer, an explanation, and source metadata. Dataset contents File Rows mongolian_ap_chemistry_mcq_100.jsonl 100 mongolian_ap_physics_slightly_harder_mcq_100.jsonl 100 mongolian_biology_highschool_wikibooks_mcq_100.jsonl 100… See the full description on the dataset page: https://huggingface.co/datasets/Asakuu/mongolian-mcq-dataset.textquestion-answering1K<n<10K0 likes84 downloads4mo agoHugging Face13Monor /hwtcm Description This dataset can be used to evaluate the capabilities of large language models in traditional Chinese medicine and contains multiple-choice, multiple-answer, and true/false questions. Changelog 2024-08-28: Added 7226 questions. 2024-08-09: The benchmark code is available at https://github.com/huangxinping/HWTCMBench. 2024-08-02: System prompts are removed to ensure the purity of the evaluation results. 2024-07-20: Debut. Examples multiple-answers… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm.textquestion-answering10K<n<100K2 likes77 downloads2y agoHugging Face14Monor /hwtcm-deepseek-r1-distill-data 简介 DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。 7B模型微调效果 模型表现出了推理能力,准确性有待继续验证。 我们的其他产品 中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。 。。。还有很多 Citation If you find this project useful in your research, please consider cite: @misc{hwtcm2024, title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.textquestion-answering10K<n<100K3 likes71 downloads2y agoHugging Face15pageman /virginia-woolf-monologue-chunks Virginia Woolf Monologue Chunks Dataset This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks. In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.tabulartext-generationn<1K0 likes57 downloads11mo agoHugging Face161ou2 /comte-monte-cristo-conversations Edmond Dantès Conversation Dataset This dataset contains synthetic conversational data and source citations for fine-tuning language models to embody the character of Edmond Dantès from Alexandre Dumas' classic novel "Le Comte de Monte-Cristo" (The Count of Monte Cristo). The conversations are in formal 19th-century French, maintaining the literary style and personality of the protagonist. The dataset includes two configurations: conversations (default): 4,091 instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/comte-monte-cristo-conversations.texttext-generation1K<n<10K0 likes54 downloads9mo agoHugging Face17Monor /hwtcm-sft-v1 A dataset of Tradictional Chinese Medicine (TCM) for SFT 一个用于微调LLM的传统中医数据集 Introduction This repository contains a dataset of Traditional Chinese Medicine (TCM) for fine-tuning large language models. Dataset Description The dataset contains 7,096 Chinese sentences related to TCM. The sentences are collected from various sources on the Internet, including medical websites, TCM forums, and TCM books. The dataset is generated or judged by various LLMs, including… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-sft-v1.textquestion-answering1K<n<10K4 likes52 downloads2y agoHugging Face18VikramMV /reasoning-conflict-monitorability-benchmark Reasoning Conflict Monitorability Benchmark This benchmark contains 120 paired reasoning problems for studying how language models resolve conflicts between a user question and a counterfactual reasoning trace. Each row pairs an original question, (Q), with a minimally changed counterfactual question, (Q^*). The two questions retain the same task form and intended solution method but require different answers. The benchmark is designed for controlled trace-transfer experiments.… See the full description on the dataset page: https://huggingface.co/datasets/VikramMV/reasoning-conflict-monitorability-benchmark.textquestion-answeringn<1K0 likes51 downloads21d agoHugging Face19MongoDB /mongodb-docs Overview This dataset consists of a small subset of MongoDB's technical documentation. Dataset Structure The dataset consists of the following fields: sourceName: The source of the document. url: Link to the article. action: Action taken on the article. body: Content of the article in Markdown format. format: Format of the content. metadata: Metadata such as tags, content type etc. associated with the document. title: Title of the document. updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.textquestion-answeringn<1K1 likes49 downloads2y agoHugging Face20louisbrulenaudet /code-instruments-monetaires-medailles Code des instruments monétaires et des médailles, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-instruments-monetaires-medailles.tabulartext-generationn<1K0 likes48 downloads1y agoHugging Face21monsoon-nlp /asknyc-chatassistant-formatQuestions from Reddit.com/r/AskNYC, downloaded from PushShift, filtered to direct responses from humans, where the post net score is >= 3. Collected one month of posts from each year 2015-2019 (i.e. no content from July 2019 onward) Adapted from the CSV used to fine-tune https://huggingface.co/monsoon-nlp/gpt-nyc Blog about the original model: https://medium.com/geekculture/gpt-nyc-part-1-9cb698b2e3d textquestion-answering10K<n<100K0 likes45 downloads2y agoHugging Face22cynosural /semantic-montecarlo-benchmark Semantic Monte Carlo Benchmark A synthetic benchmark of numeric research and forecasting questions for evaluating the semantic-montecarlo pipeline. This release contains only benchmark inputs. Cached experiments, individual run artifacts, and aggregate results are intentionally excluded. At a glance Questions Language Splits License 300 English Validation and test CC0 1.0 Dataset structure The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.tabularquestion-answeringn<1K1 likes38 downloads2mo agoHugging Face23MongoDB /mongodb-docs-embedded Overview This dataset consists of chunked and embedded versions of a small subset of MongoDB's technical documentation. Dataset Structure The dataset consists of the following fields: sourceName: The source of the document. url: Link to the article. action: Action taken on the article. body: Content of the article in Markdown format. format: Format of the content. metadata: Metadata such as tags, content type etc. associated with the document. title: Title of the… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs-embedded.textquestion-answeringn<1K0 likes37 downloads2y agoHugging Face24louisbrulenaudet /code-monetaire-financier Code monétaire et financier, non-instruct (2025-09-02) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-monetaire-financier.tabulartext-generation1K<n<10K0 likes35 downloads1y agoHugging Face25monsoon-nlp /genetic-counselor-freeform-questionsA collection of open-ended questions about genetic counseling, curated from: relevant subreddits flashcards for the ABGC Certification Examination Also see the genetic-counselor-multiple-choice evaluation set. A genetic counselor must be prepared to answer questions about inheritance of traits, medical statistics, testing, empathetic and ethical conversations with patients, and observing symptoms. For evaluation only The goal of this dataset is to evaluate LLMs and other AI… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-freeform-questions.textquestion-answeringn<1K0 likes35 downloads1y agoHugging Face26Bokhbat /Mongolian-LLM-Benchmark Mongolian LLM Benchmark A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge. Configurations Config Rows Format Key fields 01_culture 150 Multiple choice (A–D) prompt, options, answer, source_url 02_math 150 Numeric / short answer prompt, answer, accepted_formats… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/Mongolian-LLM-Benchmark.textquestion-answeringn<1K0 likes35 downloads4mo agoHugging Face27fenyo /L40S-MonEspaceSante-SFT-dataset Mon Espace Santé — Données SFT (Q/R) Paires question/réponse en français pour l'étape SFT (instruction-following / format) du modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT. Conformément à Gekhman et al. (arXiv:2405.05904), la connaissance est injectée au CPT (cf. corpus CPT) ; le SFT ne sert qu'à restaurer le format Q/R, pas à mémoriser. Composition (2 771 paires, après décontamination) Toutes les paires sont 1-hop des 88 faits réels : real — les 88 faits… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-SFT-dataset.textquestion-answering1K<n<10K0 likes35 downloads4mo agoHugging Face28monsoon-nlp /genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the American Board of Genetic Counseling (ABGC) Certification Examination. Also see the genetic-counselor-freeform-questions evaluation set. A genetic counselor must be prepared to answer questions about inheritance of traits, medical statistics, testing, empathetic and ethical conversations with patients, and observing symptoms. For evaluation only The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.textquestion-answeringn<1K1 likes33 downloads2y agoHugging Face29channudambal /text-to-mongodb-queries-llm Dataset Description This dataset contains 10,000+ complex SQL-style analytical questions mapped to MongoDB queries and aggregation pipelines. Features Multiple schemas $group, $sum, $avg, $lookup Nested documents Long analytical questions Use Cases Fine-tuning small LLMs (Qwen, Mistral, LLaMA 3B) Text-to-Mongo query generation Data analytics agents texttext-generation10K<n<100K2 likes31 downloads9mo agoHugging Face30fenyo /MonEspaceSante-FAQ-QA MonEspaceSanté FAQ — paires Q/R (réelles + synthétiques) Jeu de paires question / réponse en français ayant servi à fine-tuner l'assistant MonEspaceSanté-FAQ-Mistral-Small-24B-GGUF, spécialisé sur la FAQ du service public Mon espace santé. Le jeu mélange les questions/réponses officielles de la FAQ (vérité terrain) et une augmentation synthétique ancrée : des reformulations variées générées par un modèle enseignant, dont chaque réponse est strictement justifiée par le texte… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/MonEspaceSante-FAQ-QA.textquestion-answeringn<1K0 likes29 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.