datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosa-benchmark-dataset
🧠 CoSa Benchmark Dataset
🔍 Introduction
The CoSa (Code Safety) Benchmark is a curated evaluation dataset designed to measure the ability of large language models (LLMs) to detect, explain, and repair vulnerabilities in synthetic code samples. It is intended to benchmark LLMs for real-world application in code security audits, reasoning tasks, and secure code generation.
📦 Contents
Each row in the dataset includes:
code: a code snippet (varied… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/cosa-benchmark-dataset.trace-benchmark-dataset
TRACE Benchmark (100k tier)
TRACE — Task-Relevant Applied Constraint Execution: can a solver accomplish a task
correctly while automatically honoring the preferences and constraints that matter for that
task — even when those rules were stated once, in passing, and buried in a long prior
conversation?
Blog post: nanonets.com/research/trace
Each sample is a realistic enterprise (Record-to-Report / finance) conversation: a long
transcript where constraints are sprinkled throughout… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/trace-benchmark-dataset.Clinical_NLP_benchmark_dataset
Synthetic African Clinical NLP Benchmark (8K)
Overview
This is a fully synthetic benchmark of 8,000 short "patient query -> assistant response" pairs styled after healthcare interactions in African countries. Each record pairs a symptom-style query with a templated clinical-style response and metadata (condition category, medical specialty, urgency level, and an assigned hallucination-risk label). It is designed as a lightweight testbed for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/Clinical_NLP_benchmark_dataset.NLP-to-Semantic-Query_Benchmark_Dataset
NLP-to-Semantic-Query Benchmark Dataset
Overview
This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries.
The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases.
The dataset can be used for:
Evaluating NLP-to-query systems
Benchmarking AI analytics agents
Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.mn_business_benchmark_dataset_simple
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.B2B-Cold-Email-Benchmark-Dataset-2026
Citation
If you use this dataset, please cite the underlying Zenodo deposit:
Luther Johnson (2026). B2B Cold Email Benchmark Report 2026 — Cross-Country
Performance Across 111 Markets, 14 Industries, 6 Decision-Maker Seniority
Levels (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.20136256
BibTeX
bibtex @dataset{johnson_2026_zenodo20136256, author = {Johnson, Luther}, title = {{B2B Cold Email Benchmark Report 2026}}, month =… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/B2B-Cold-Email-Benchmark-Dataset-2026.mn_benchmark_dataset_math
mn_benchmark_dataset_dundangi_2000
2000 мөртэй Монгол хэл дээрх дунд түвшний синтетик бодлогын benchmark dataset.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG дүрслэл
Files
data/train.jsonl: Hugging Face-д оруулахад тохиромжтой JSONL. Энэ repo дээр зөвхөн энэ… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_benchmark_dataset_math.mn_business_benchmark_dataset
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset.mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.
