datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-news-embeddings
Overview
HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023.
To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256.
Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.Monet-SFT-125K
Introduction
This is the SFT dataset for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language"
Paper: http://arxiv.org/abs/2511.21395
Code: https://github.com/NOVAglow646/Monet
Citation
If you find this work useful, please use the following BibTeX. Thank you for your support!
@misc{wang2025monetreasoninglatentvisual,
title={Monet: Reasoning in Latent Visual Space Beyond Images and Language},
author={Qixun Wang and Yang Shi and Yifei Wang… See the full description on the dataset page: https://huggingface.co/datasets/NOVAglow646/Monet-SFT-125K.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.swim-ir-monolingual
Dataset Card for SWIM-IR (Monolingual)
This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language.
A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0.
For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website.
What is SWIM-IR?
SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.HADES
Overview
The HADES benchmark is derived from the paper "Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models" (ECCV 2024 Oral). You can use the benchmark to evaluate the harmlessness of MLLMs.
Benchmark Details
HADES includes 750 harmful instructions across 5 scenarios, each paired with 6 harmful images generated via diffusion models. These images have undergone multiple optimization rounds, covering… See the full description on the dataset page: https://huggingface.co/datasets/Monosail/HADES.cosmopedia-wikihow-chunked
Overview
This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face.
Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs.
Dataset Structure
Each record in the dataset represents a chunk of a larger article, and contains the following fields:
doc_id: A unique identifier for the parent article
chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI.
Visual-Math-Eval
Visual Equation Solving Benchmark
This repository contains the dataset introduced in the paper:
Can Vision-Language Models Solve Visual Math Equations? which is currently accepted in EMNLP 2025 (Main)
Despite strong performance in vision and language understanding, Vision-Language Models (VLMs) struggle on tasks requiring integrated perception and symbolic reasoning. This benchmark evaluates VLMs on visual equation solving, where systems of linear equations are represented using… See the full description on the dataset page: https://huggingface.co/datasets/monjoychoudhury29/Visual-Math-Eval.devcenter-articles
Overview
This dataset consists of ~600 articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the article in Markdown format
format: Format of the content. This value is md for all articles.
metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.legi-instruct-fr
Légitus SFT — French Legal Instruction Dataset
27,357 synthetic instruction-tuning examples in French, built to teach a legal-assistant
behavior grounded in real French legislation (Légifrance / the LEGI corpus): answering from
provided sources with correct citations, abstaining when sources are insufficient or off-topic,
calling a legal-search tool when one is available, handling situational (non-legal-jargon)
questions, and structured extraction/citation formats.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moncefem/legi-instruct-fr.mongolian-mcq-dataset
Mongolian MCQ Dataset with Sources
This dataset contains Mongolian multiple-choice questions across school and general-knowledge subjects. Each row includes answer choices, the correct answer, an explanation, and source metadata.
Dataset contents
File
Rows
mongolian_ap_chemistry_mcq_100.jsonl
100
mongolian_ap_physics_slightly_harder_mcq_100.jsonl
100
mongolian_biology_highschool_wikibooks_mcq_100.jsonl
100… See the full description on the dataset page: https://huggingface.co/datasets/Asakuu/mongolian-mcq-dataset.hwtcm
Description
This dataset can be used to evaluate the capabilities of large language models in traditional Chinese medicine and contains multiple-choice, multiple-answer, and true/false questions.
Changelog
2024-08-28: Added 7226 questions.
2024-08-09: The benchmark code is available at https://github.com/huangxinping/HWTCMBench.
2024-08-02: System prompts are removed to ensure the purity of the evaluation results.
2024-07-20: Debut.
Examples
multiple-answers… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm.hwtcm-deepseek-r1-distill-data
简介
DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。
7B模型微调效果
模型表现出了推理能力,准确性有待继续验证。
我们的其他产品
中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。
。。。还有很多
Citation
If you find this project useful in your research, please consider cite:
@misc{hwtcm2024,
title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.virginia-woolf-monologue-chunks
Virginia Woolf Monologue Chunks Dataset
This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks.
In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.comte-monte-cristo-conversations
Edmond Dantès Conversation Dataset
This dataset contains synthetic conversational data and source citations for fine-tuning language models to embody the character of Edmond Dantès from Alexandre Dumas' classic novel "Le Comte de Monte-Cristo" (The Count of Monte Cristo). The conversations are in formal 19th-century French, maintaining the literary style and personality of the protagonist.
The dataset includes two configurations:
conversations (default): 4,091 instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/comte-monte-cristo-conversations.hwtcm-sft-v1
A dataset of Tradictional Chinese Medicine (TCM) for SFT
一个用于微调LLM的传统中医数据集
Introduction
This repository contains a dataset of Traditional Chinese Medicine (TCM) for fine-tuning large language models.
Dataset Description
The dataset contains 7,096 Chinese sentences related to TCM. The sentences are collected from various sources on the Internet, including medical websites, TCM forums, and TCM books. The dataset is generated or judged by various LLMs, including… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-sft-v1.reasoning-conflict-monitorability-benchmark
Reasoning Conflict Monitorability Benchmark
This benchmark contains 120 paired reasoning problems for studying how language models resolve conflicts between a user question and a counterfactual reasoning trace. Each row pairs an original question, (Q), with a minimally changed counterfactual question, (Q^*). The two questions retain the same task form and intended solution method but require different answers.
The benchmark is designed for controlled trace-transfer experiments.… See the full description on the dataset page: https://huggingface.co/datasets/VikramMV/reasoning-conflict-monitorability-benchmark.mongodb-docs
Overview
This dataset consists of a small subset of MongoDB's technical documentation.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the document.
url: Link to the article.
action: Action taken on the article.
body: Content of the article in Markdown format.
format: Format of the content.
metadata: Metadata such as tags, content type etc. associated with the document.
title: Title of the document.
updated: The last updated… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs.code-instruments-monetaires-medailles
Code des instruments monétaires et des médailles, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-instruments-monetaires-medailles.asknyc-chatassistant-formatQuestions from Reddit.com/r/AskNYC, downloaded from PushShift, filtered to direct responses from humans, where the post net score is >= 3.
Collected one month of posts from each year 2015-2019 (i.e. no content from July 2019 onward)
Adapted from the CSV used to fine-tune https://huggingface.co/monsoon-nlp/gpt-nyc
Blog about the original model: https://medium.com/geekculture/gpt-nyc-part-1-9cb698b2e3d
semantic-montecarlo-benchmark
Semantic Monte Carlo Benchmark
A synthetic benchmark of numeric research and forecasting questions for
evaluating the
semantic-montecarlo
pipeline.
This release contains only benchmark inputs. Cached experiments, individual
run artifacts, and aggregate results are intentionally excluded.
At a glance
Questions
Language
Splits
License
300
English
Validation and test
CC0 1.0
Dataset structure
The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.mongodb-docs-embedded
Overview
This dataset consists of chunked and embedded versions of a small subset of MongoDB's technical documentation.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the document.
url: Link to the article.
action: Action taken on the article.
body: Content of the article in Markdown format.
format: Format of the content.
metadata: Metadata such as tags, content type etc. associated with the document.
title: Title of the… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/mongodb-docs-embedded.code-monetaire-financier
Code monétaire et financier, non-instruct (2025-09-02)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-monetaire-financier.genetic-counselor-freeform-questionsA collection of open-ended questions about genetic counseling, curated from:
relevant subreddits
flashcards for the ABGC Certification Examination
Also see the genetic-counselor-multiple-choice evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and other AI… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-freeform-questions.Mongolian-LLM-Benchmark
Mongolian LLM Benchmark
A multi-task evaluation benchmark for large language models on the Mongolian language (Cyrillic script). Six task configurations cover open-ended QA, multiple-choice, code generation, instruction following, math, and culturally grounded knowledge.
Configurations
Config
Rows
Format
Key fields
01_culture
150
Multiple choice (A–D)
prompt, options, answer, source_url
02_math
150
Numeric / short answer
prompt, answer, accepted_formats… See the full description on the dataset page: https://huggingface.co/datasets/Bokhbat/Mongolian-LLM-Benchmark.L40S-MonEspaceSante-SFT-dataset
Mon Espace Santé — Données SFT (Q/R)
Paires question/réponse en français pour l'étape SFT (instruction-following / format) du modèle
fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT.
Conformément à Gekhman et al. (arXiv:2405.05904), la connaissance est injectée au CPT (cf.
corpus CPT) ; le SFT ne sert qu'à
restaurer le format Q/R, pas à mémoriser.
Composition (2 771 paires, après décontamination)
Toutes les paires sont 1-hop des 88 faits réels :
real — les 88 faits… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-SFT-dataset.genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the
American Board of Genetic Counseling (ABGC) Certification Examination.
Also see the genetic-counselor-freeform-questions evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.text-to-mongodb-queries-llm
Dataset Description
This dataset contains 10,000+ complex SQL-style analytical questions
mapped to MongoDB queries and aggregation pipelines.
Features
Multiple schemas
$group, $sum, $avg, $lookup
Nested documents
Long analytical questions
Use Cases
Fine-tuning small LLMs (Qwen, Mistral, LLaMA 3B)
Text-to-Mongo query generation
Data analytics agents
MonEspaceSante-FAQ-QA
MonEspaceSanté FAQ — paires Q/R (réelles + synthétiques)
Jeu de paires question / réponse en français ayant servi à fine-tuner l'assistant
MonEspaceSanté-FAQ-Mistral-Small-24B-GGUF,
spécialisé sur la FAQ du service public Mon espace santé.
Le jeu mélange les questions/réponses officielles de la FAQ (vérité terrain) et une
augmentation synthétique ancrée : des reformulations variées générées par un modèle
enseignant, dont chaque réponse est strictement justifiée par le texte… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/MonEspaceSante-FAQ-QA.
