datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosa-benchmark-dataset
🧠 CoSa Benchmark Dataset
🔍 Introduction
The CoSa (Code Safety) Benchmark is a curated evaluation dataset designed to measure the ability of large language models (LLMs) to detect, explain, and repair vulnerabilities in synthetic code samples. It is intended to benchmark LLMs for real-world application in code security audits, reasoning tasks, and secure code generation.
📦 Contents
Each row in the dataset includes:
code: a code snippet (varied… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/cosa-benchmark-dataset.trace-benchmark-dataset
TRACE Benchmark (100k tier)
TRACE — Task-Relevant Applied Constraint Execution: can a solver accomplish a task
correctly while automatically honoring the preferences and constraints that matter for that
task — even when those rules were stated once, in passing, and buried in a long prior
conversation?
Blog post: nanonets.com/research/trace
Each sample is a realistic enterprise (Record-to-Report / finance) conversation: a long
transcript where constraints are sprinkled throughout… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/trace-benchmark-dataset.NLP-to-Semantic-Query_Benchmark_Dataset
NLP-to-Semantic-Query Benchmark Dataset
Overview
This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries.
The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases.
The dataset can be used for:
Evaluating NLP-to-query systems
Benchmarking AI analytics agents
Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.mn_business_benchmark_dataset_simple
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.mn_benchmark_dataset_math
mn_benchmark_dataset_dundangi_2000
2000 мөртэй Монгол хэл дээрх дунд түвшний синтетик бодлогын benchmark dataset.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG дүрслэл
Files
data/train.jsonl: Hugging Face-д оруулахад тохиромжтой JSONL. Энэ repo дээр зөвхөн энэ… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_benchmark_dataset_math.mn_business_benchmark_dataset
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset.mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.
