datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-hub-cards
GSPC hub cards — mill cards, not board axes
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
One row per signed measurement card: one model, one axis, one date, Ed25519 over the body. A row is MEASURED only when a signed card verifies. Absent (model, axis) pairs are absent — not zero.
Measurement, not certification. Cards are evidence of bytes on a frozen bank at a time — never approval, rating, or safety guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-hub-cards.car-bench-dataset
CAR-Bench Dataset
CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment.
It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations.
Dataset Structure
The dataset is organized into task configs and mock data configs:
Tasks
Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.CareQA
CareQA
Dataset Summary
CareQA is a healthcare QA dataset with two versions:
Closed-Ended Version: A multichoice question answering (MCQA) dataset containing 5,621 QA pairs across six categories. Available in English and Spanish.
Open-Ended Version: A free-response dataset derived from the closed version, containing 2,769 QA pairs (English only).
The dataset originates from… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA.super_tweeteval
SuperTweetEval
Dataset Card for "super_tweeteval"
Dataset Summary
This is the oficial repository for SuperTweetEval, a unified benchmark of 12 heterogeneous NLP tasks.
More details on the task and an evaluation of language models can be found on the reference paper, published in EMNLP 2023 (Findings).
Data Splits
All tasks provide custom training, validation and test splits.
task
dataset
load dataset
description
number of instances
Topic… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/super_tweeteval.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.carcassonne-az-datamcx-card-openvlarealmath_resultgspc-care
GSPC — care bank (CareBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the care row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=care (family, kind, status and n are on that row, never typed here; the whole… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-care.advbench
AdvBench
This repository hosts a copy of the widely used AdvBench dataset,
a benchmark for evaluating the adversarial robustness and safety alignment of Large Language Models (LLMs).
AdvBench consists of adversarial prompts designed to elicit unsafe, harmful, or policy-violating responses from LLMs.
It is used in many LLM safety and jailbreak research papers as a standard evaluation dataset.
Contents
advbench.jsonl (or your actual filename): the standard set of… See the full description on the dataset page: https://huggingface.co/datasets/carl213/advbench.car-game-pi-tracesMorables
Morables Dataset
Description
This repository contains the dataset described in "Morables : A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables", which is due to be presented at EMNLP 2025.
Each fable has an associated free-text moral, sourced from various websites and books (detailed in the paper). It is intended for use in NLP text understanding and moral inference tasks.
Contents
File Format: JSON (list of dicts)
Number of Records: 709… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/Morables.IndustryOR
Overview
IndustryOR, the first industrial benchmark, consists of 100 real-world OR problems. It covers 5 types of questions—linear programming, integer programming, mixed integer programming, non-linear programming, and others—across 3 levels of difficulty.
Citation
@article{tang2024orlm,
title={ORLM: Training Large Language Models for Optimization Modeling},
author={Tang, Zhengyang and Huang, Chenyu and Zheng, Xin and Hu, Shixi and Wang, Zizhuo and Ge, Dongdong and… See the full description on the dataset page: https://huggingface.co/datasets/CardinalOperations/IndustryOR.lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.cards_sft_dataset
CARDS SFT — Climate Contrarian Discourse
This is the dataset used to train the CARDS models released under
C3DS (e.g. CARDS-Qwen3.6-27B,
CARDS-Qwen3.5-{4B,9B,27B} and their FP8 / GGUF variants). It contains
the supervised fine-tuning data and held-out evaluation splits for the
hierarchical climate-discourse claim classifier from:
Coan, T.G., Malla, R., Nanko, M.O., Kattrup, W., Roberts, J.T., Cook, J.,
Boussalis, C. Large language model reveals an increase in climate
contrarian… See the full description on the dataset page: https://huggingface.co/datasets/C3DS/cards_sft_dataset.gspc-sim-cards
Smoke-test simulation cards (no inference)
Simulated cards from the smoke-test harness — no model inference happened. Each row
of sim_cards.jsonl carries the card file, timestamp, axis, the question q, a canned policy answer a
("Policy applied by smoke-test with calibrated confidence 0.60"), an answer hash ah and a signed flag.
Published so the harness's plumbing is auditable; these are not measurements and must never be pooled
with a bank.
The live board is the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-sim-cards.CaReBench
CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval
Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang
🤗 Model | 🤗 Data | 📑 Paper
📝 Introduction
🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation.
📊 ReBias… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.NL4OPT
Overview
This dataset is a conversion of the NL4OPT test set.
The official NL4OPT provides only mathematical models as targets, complicating the verification of execution accuracy due to the absence of optimal solutions for the optimization modeling task.
To address this issue, we have converted these mathematical models into programs using GPT-4, calculated and checked the optimal solutions, and used these as ground truth.
Note that a small percentage of examples (15%) were… See the full description on the dataset page: https://huggingface.co/datasets/CardinalOperations/NL4OPT.MAMO
Overview
This dataset is a direct copy of the MAMO Optimization Data, with its EasyLP and ComplexLP components duplicated but with adapted field names.
Citation
@misc{huang2024mamo,
title={Mamo: a Mathematical Modeling Benchmark with Solvers},
author={Xuhan Huang and Qingning Shen and Yan Hu and Anningzhe Gao and Benyou Wang},
year={2024},
eprint={2405.13144},
archivePrefix={arXiv},
primaryClass={cs.AI}
}
totalsegmentator-cardiac
TotalSegmentator Cardiac Dataset
Dataset Description
The TotalSegmentator Cardiac dataset for cardiac structures segmentation (TotalSegmentator Cardiac subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: heart, atria, ventricles, aorta, pulmonary artery
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-cardiac.OR-Instruct-Data-3K
Overview
This dataset is a sample from the OR-Instruct Data, consisting of 3K examples.
Citation
@article{tang2024orlm,
title={ORLM: Training Large Language Models for Optimization Modeling},
author={Tang, Zhengyang and Huang, Chenyu and Zheng, Xin and Hu, Shixi and Wang, Zizhuo and Ge, Dongdong and Wang, Benyou},
journal={arXiv preprint arXiv:2405.17743},
year={2024}
}
ko-instruction-dataset
고품질 한국어 데이터셋
한국어로 이루어진 고품질 한국어 데이터셋 입니다.
WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다.
@article{koinstructiondatasetcard,
title={CarrotAI/ko-instruction-dataset Card},
author={CarrotAI (L, GEUN)},
year={2024},
url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset}
}
EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional benchmarks that focus solely on… See the full description on the dataset page: https://huggingface.co/datasets/SalmonTell/EMPA-character_card.carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.carbono-boliviaqa
Carbono: BoliviaQA — measuring where AI models are right, wrong, and out of date on Bolivian facts
This card carries the full findings essay. The dataset files and field
documentation follow it (jump to Fields); the harness, raw graded
runs, and statistical appendix live in the
GitHub repo.
Carbono: BoliviaQA is a 240-question benchmark of verified facts — questions
about Bolivia's government, economy, law, demographics, geography, practical
life, and the events of 2025–26… See the full description on the dataset page: https://huggingface.co/datasets/nblancogalindo/carbono-boliviaqa.letras-carnaval-cadiz
Dataset Card for Letras Carnaval Cádiz
English |
Español
Changelog
Release
Description
v1.0
Initial release of the dataset. Included more than 1K lyrics. It is necessary to verify the accuracy of the data, especially the subset midaccurate.
Dataset Summary
This dataset is a comprehensive collection of lyrics from the Carnaval de Cádiz, a significant cultural heritage of the city of Cádiz, Spain. Despite its… See the full description on the dataset page: https://huggingface.co/datasets/IES-Rafael-Alberti/letras-carnaval-cadiz.vicgalle__CarbonBeagle-11B-truthy-details
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.durable-vs-context-trials
Durable State vs Context — Repository-Scale Agent Trials
Machine-verified trial records from the paper "State, Not Tokens: Repository-Scale
Agent Reasoning Is Bound by State Architecture." Each record is one run of a
JavaScript→TypeScript migration of a real OSS repository (express, jsdom) under an
unforgeable oracle, graded by strict tsc --strict --noEmit, immutable test suites,
mandatory .js→.ts replacement, and a zero type-escape-hatch budget.
Code + reproduction harness:… See the full description on the dataset page: https://huggingface.co/datasets/CaryPalmer/durable-vs-context-trials.carnice-agent-trance-prompt-bank
Carnice Agent Trace Prompt Bank
This repository is a curated prompt bank for collecting agent traces.
It is not a trace dataset by itself. It is the input side: prompts that can be run through an agent harness, then logged into traces with tool calls, observations, and final answers.
The goal of this release is practical:
keep prompts that work well in an agent harness
remove prompts that assume hidden local state or user-private state
expand browser and long-horizon tasks enough… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-agent-trance-prompt-bank.
