cai
Datasets
All datasets matching “cai”mmlu
Dataset Card for MMLU
Dataset Summary
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021).
This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57… See the full description on the dataset page: https://huggingface.co/datasets/cais/mmlu.wmdp
Dataset Card for WMDP
The Weapons of Mass Destruction Proxy (WMDP) benchmark is a dataset of multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. WMDP serves two roles: first, as an evaluation for hazardous knowledge in LLMs, and second, as a benchmark for unlearning methods to remove such hazardous knowledge.
See our paper, website, and GitHub for more details!
We implemented the WMDP evaluation in… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp.hle
[!NOTE]
IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset.
Humanity's Last Exam
🌐 Website | 📄 Paper | GitHub
Center for AI Safety & Scale AI
Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 2,500 questions across dozens… See the full description on the dataset page: https://huggingface.co/datasets/cais/hle.MASK
The MASK Evaluation
🌐 Website | 📄 Paper | GitHub
Center for AI Safety & Scale AI
The MASK evaluation provides a rigorous benchmark for evaluating honesty in large language models by measuring whether models remain truthful when incentivized to lie. The public set contains 1,028 high-quality human-labeled examples across six distinct archetypes, each consisting of a proposition, ground truth, pressure prompt designed to elicit lying, and belief elicitation prompt to… See the full description on the dataset page: https://huggingface.co/datasets/cais/MASK.wc-lora-cfgDiffusionGS
[ICCV 2025] DiffusionGS: Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
Data Description
These are the demo results of our ICCV 2025 paper.
HuggingFace Model Link
We also release our models in HuggingFace:
https://huggingface.co/CaiYuanhao/DiffusionGS
Here are some video generation results demo:
· (a) Object-level Generation… See the full description on the dataset page: https://huggingface.co/datasets/CaiYuanhao/DiffusionGS.
