datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.MoleculeQA
Dataset Card for MoleculeQA
Dataset Details
Dataset Description
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension (EMNLP 2024)
Curated by: IDEA-XL
Language(s) (NLP): en
License: mit
Dataset Sources
Repository: https://github.com/IDEA-XL/MoleculeQA
Paper [optional]: https://arxiv.org/abs/2403.08192
Dataset Structure
- JSON
- All
- train.json # 49,993
- valid.json # 5,795
- test.json # 5… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/MoleculeQA.BiomniBench-AI4S
BiomniBench-AI4S — Agent Trajectories
Per-cell outputs from a horizontal comparison of 7 AI-for-Science agents on
the same 50 BiomniBench-DA biomedical
data-analysis tasks, under identical conditions — same model (deepseek-v4-pro),
same DeepSeek v4-pro rubric judge.
Leaderboard, harness, adapters, and analysis:
👉 https://github.com/omicverse/BiomniBench-AI4S
Layout
<backend>/<task>/
trace.md # the agent's structured analytical trace
answer.txt… See the full description on the dataset page: https://huggingface.co/datasets/omicverse/BiomniBench-AI4S.
