datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
dataset_for_benchmarkcvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
probe-benchmark-hard50
PROBE hard50 — human-reviewed hard development set
This is a separate standard-format review bank of 62 Astra-conditioned hard development episodes. It is NOT an independently evaluated holdout. Do not merge into or modify benchmark_600.
The original 50 human-reviewed questions were augmented with 12 human-accepted inverse comparison prompts, yielding 24 compare questions total.
Contents and ordering
Type
Count
Directories
beneath
10… See the full description on the dataset page: https://huggingface.co/datasets/vineet-datasets/probe-benchmark-hard50.placeholder_tiebeprobe-benchmark-hard90
PROBE hard90
Human-reviewed hard development benchmark built from benchmark-hard50 plus accepted candidates from benchmark-hard30-review.
This bank contains 90 questions: beneath 13, compare 32, count 17, find 28. It is a hard development set, not a balanced or independent holdout.
Rejected hard30 candidates were not included. Accepted inverse compare variants were included as separate duplicate-scene questions.
dataset-genome-agriculture-benchmark
Dataset Genome — Agriculture Mechanism Outcomes Dataset
Project Overview & Scientific Motivation
Dataset Genome is an open-source scientific reasoning benchmark focused on Agriculture. The dataset evaluates AI systems on hypothesis generation, observation analysis, experimental design, and scientific reasoning using agriculture-related scenarios. It was generated using the Dataset Genome pipeline and adapted through Adaption Adaptive Data.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/surabhi08/dataset-genome-agriculture-benchmark.han-cross-agent-skill-transfer-benchmark-dataset-v1
Humanoid Cross-Agent Skill Transfer Benchmark Dataset
This dataset benchmarks how effectively
skills learned by one humanoid agent
can be transferred to another agent
within a decentralized cognitive network.
Objective
To measure cross-agent generalization,
adaptation speed, and transfer efficiency.
Data Fields
source_agent_skill_profile
target_agent_initial_profile
transferred_skill_vector
adaptation_steps
performance_improvement_percentage… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-cross-agent-skill-transfer-benchmark-dataset-v1.mn_business_benchmark_dataset_simple
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.mn_benchmark_dataset_math
mn_benchmark_dataset_dundangi_2000
2000 мөртэй Монгол хэл дээрх дунд түвшний синтетик бодлогын benchmark dataset.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG дүрслэл
Files
data/train.jsonl: Hugging Face-д оруулахад тохиромжтой JSONL. Энэ repo дээр зөвхөн энэ… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_benchmark_dataset_math.mn_business_benchmark_dataset
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset.mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.ISO27K-QnA-Benchmark-dataset
QnA benchmark for the ISO/IEC 27000 family of standards
This repository holds a collection of 222 multiple-choice questions and ground-truth answers covering the ISO/IEC 27000 family of standards and it was compiled for evaluation/testing purposes.
Its questions and answers were manualy extracted and constructed from the following literature covering the ISO/IEC 27000 standards:
"ISO 27001 Foundation – Practice Tests: 150 Questions And Explanations Based On The ISO 27001 Foundation… See the full description on the dataset page: https://huggingface.co/datasets/dimitarjovanovski/ISO27K-QnA-Benchmark-dataset.Kannada_WiC_benchmark_dataset
Kannada WiC Benchmark v2
Dataset Summary
WiC benchmark synthesized from live Kannada IndoWordNet (pyiwn) and Kannada Wikipedia context.
Dataset Structure
Total pairs: 640
Unique words: 45
Label 1 count: 320
Label 0 count: 320
Splits: train / validation / test
Data Sources
IndoWordNet Kannada synsets and examples via pyiwn
Kannada Wikipedia summaries and search snippets
Quality Control
20% sample validation report generated in… See the full description on the dataset page: https://huggingface.co/datasets/Gnanaqubit/Kannada_WiC_benchmark_dataset.BenchMarkDataset
