CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes35k downloads3mo agoHugging Face02lockon /glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en. texttext-generation1K<n<10K1 likes27k downloads2y agoHugging Face03EuniAI /TerminalWorld TerminalWorld Dataset Summary TerminalWorld is a benchmark dataset for evaluating AI agents on real-world terminal and command-line tasks. It contains 1,530 terminal-based tasks reverse-engineered from publicly available terminal recordings, covering domains such as data processing, system administration, networking, security, version control, containers and orchestration, debugging and testing, environment setup, and scientific computing. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/EuniAI/TerminalWorld.texttext-generation1K<n<10K10 likes8.7k downloads3mo agoHugging Face04mirobody /ESL-Bench ESL-bench ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework. ⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.textquestion-answering1K<n<10K18 likes6.8k downloads18d agoHugging Face05Dongping-Li /EMMOE-100 EMMOE-100 Trainset Resources Project Paper Code Model Dataset Dataset Feature Task Attributes Task Example Dataset Structure EMMOE-100/ ├── README.md ├── assets/ ├── data/ │ └── train/ │ ├── 1/ │ │ ├── info.txt │ │ ├── info_re1.txt │ │ ├── info_re2.txt │ │ ├── info_re3.txt │ │ ├── keypath.json │ │ ├── scene.json │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Dongping-Li/EMMOE-100.imagevisual-question-answering10K<n<100K1 likes5.5k downloads1y agoHugging Face06Anthropic /model-written-evals Model-Written Evaluation Datasets This repository includes datasets written by language models, used in our paper on "Discovering Language Model Behaviors with Model-Written Evaluations." We intend the datasets to be useful to: Those who are interested in understanding the quality and properties of model-generated data Those who wish to use our datasets to evaluate other models for the behaviors we examined in our work (e.g., related to model persona, sycophancy, advanced AI risks… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/model-written-evals.textmultiple-choice1K<n<10K68 likes3.6k downloads4y agoHugging Face07yongxin2020 /TempPerturb-Eval-data TempPerturb-Eval-data Summary TempPerturb-Eval-data is the released output dataset for TempPerturb-Eval, a benchmark for analyzing the robustness of Retrieval-Augmented Generation (RAG) systems under both internal variation and external perturbation. This is an evaluation-artifact dataset: it stores model outputs and experiment metadata for controlled robustness analysis, rather than a new QA training corpus. The release covers: 5 models 11 temperatures from 0.0 to 2.0 4… See the full description on the dataset page: https://huggingface.co/datasets/yongxin2020/TempPerturb-Eval-data.textquestion-answering10K<n<100K1 likes3.6k downloads7mo agoHugging Face08NJU-LINK /DR3-EvalDR3-Eval: Towards Realistic and ReproducibleDeep Research Evaluation ✨ Overview DR³-Eval is a realistic, reproducible, and multimodal evaluation benchmark for Deep Research Agents, focusing on multi-file report generation tasks. Existing benchmarks face a fundamental tension between realism, controllability, and reproducibility when evaluating deep research agents. DR³-Eval addresses this through the following design:… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/DR3-Eval.documenttext-generationn<1K2 likes3.1k downloads5mo agoHugging Face09scimdr /SciMDR-Evalimagequestion-answeringn<1K1 likes3k downloads7mo agoHugging Face10zen-E /GSM8k-AugThis dataset is provided to facilitate access to GSM8k-Aug, originally from https://github.com/da03/Internalize_CoT_Step_by_Step and https://arxiv.org/pdf/2405.14838. This dataset is used to train CODI (https://arxiv.org/abs/2502.21074) Description: We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page: https://huggingface.co/datasets/zen-E/GSM8k-Aug.textquestion-answering100K<n<1M5 likes2.6k downloads1y agoHugging Face11paperuploadacount /EO-Gym EO Gym EO Gym provides a local Earth-observation tool server and trainer environment adapter. It exposes remote-sensing tools for cropping imagery, loading multispectral bands, computing masks and indices, inspecting metadata, and running EO Gym rollouts through a trainer-facing API. Croissant metadata EO-Gym provides two Croissant representations: Hugging Face-generated Croissant metadata: /api/datasets/paperuploadacount/EO-Gym/croissant This is generated… See the full description on the dataset page: https://huggingface.co/datasets/paperuploadacount/EO-Gym.textvisual-question-answering1K<n<10K1 likes2.6k downloads2mo agoHugging Face12typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2.2k downloads2y agoHugging Face13HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes1.9k downloads3y agoHugging Face14EricLu /SCP-116K Multimodal preview: SCP-VL. This extension of the SCP dataset family contains 41,828 English and Chinese scientific problem-solution pairs, each with one or more attached images. The current v0.1-preview release may contain missing information, incomplete or mismatched images, and solution errors. See the SCP-VL dataset card for details, known limitations, and loading instructions. New Version Available: SCP-378K A new version of this dataset has been released: 👉… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/SCP-116K.texttext-generation100K<n<1M125 likes1.6k downloads13d agoHugging Face15talmahmud /tofu_ext1textquestion-answering100K<n<1M0 likes1.2k downloads1y agoHugging Face16Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face17risenyard /egms-qa-dataset EGMS-QA Dataset Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables, and natural-language QA records for 10,000 overlapping 7 km tiles. This card describes the available data, file formats, and download options. Data access Data needed Files to download Details Published QA records train.jsonl, validation.jsonl, test.jsonl QA loading example Encoder inputs Source tiles, metadata Encoder data Translator inputs Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.textquestion-answering100K<n<1M1 likes1.1k downloads17d agoHugging Face18Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1.1k downloads1y agoHugging Face19MoreThought /DeepSWEGym2-Edu Dataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 291.74kb, a total uncompressed size of 14.93GB, and a total of 53649 examples. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Edu.texttext-generation10K<n<100K2 likes990 downloads19d agoHugging Face20lhpku20010120 /Omni-Edu Omni-Edu — Core V6 SFT mixture 69,999 supervised instruction examples (~158M characters) covering K-12 subject competence, curriculum grounding, diagnostic reasoning, pedagogical action and general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image referenced by the JSONL ships in this repository under images/. This is the system-prompted assembly of the v6 core mixture: every row carries an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.imagetext-generation10K<n<100K1 likes855 downloads7d agoHugging Face21emre570 /uscode_qacA compact question-answer set for the Prime Intellect U.S. legal evaluation environment. Each record pairs a natural-language question with an extractive answer and the source statute snippet drawn from the Cornell Law School Legal Information Institute (LII) U.S. Code site. Fields also include title_id, section_id, and section_url to support retrieval-style evaluations; the snippet lives in context and is used to build the search index rather than being passed directly to the model.… See the full description on the dataset page: https://huggingface.co/datasets/emre570/uscode_qac.textquestion-answeringn<1K0 likes843 downloads10mo agoHugging Face22miracle10 /EarthVerse Benchmarking scientific agents across dynamic Earth systems and natural hazards Zhiqing Cui1, Xinxiang Yin2, Yihong Tang3, Xinglang Zhang4, Yuanzhe Hu5, Siru Zhong4, Weidong Tang6, Yuxuan Liang4, Weijia Li7, Ming Jin8, Shirui Pan8, Yuhao Kang9, Dingyi Zhuang10,†, Jinhua Zhao10 1NUIST &nbsp; 2HKU &nbsp; 3McGill &nbsp; 4HKUST(GZ) &nbsp; 5Georgia Tech &nbsp; 6NUS &nbsp; 7Tsinghua &nbsp; 8Griffith &nbsp; 9UT Austin &nbsp; 10MIT &nbsp; †Corresponding author Project page ·… See the full description on the dataset page: https://huggingface.co/datasets/miracle10/EarthVerse.textquestion-answeringn<1K1 likes835 downloads1mo agoHugging Face23HiTZ /EusExams Dataset Card for EusExams [!WARNING] A newer version of this dataset is available! Please use EusExams-v2 which features deduplication, data grouping, and new data. EusExams is a collection of tests designed to prepare individuals for Public Service examinations conducted by several Basque institutions, including the public health system Osakidetza, the Basque Government, the City Councils of Bilbao and Gasteiz, and the University of the Basque Country (UPV/EHU). Within each… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EusExams.textquestion-answering10K<n<100K2 likes829 downloads3mo agoHugging Face24mjbommar /opengloss-v1.3-query-examples-flat See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 (Flattened) Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.texttext-generation100K<n<1M0 likes763 downloads18d agoHugging Face25tencent /ElephantBench ElephantBench ElephantBench is a closed-book knowledge probe for evaluating whether a language model remembers long-tail facts and recalls the different verified accounts associated with them. The release contains 1,094 English questions. Evaluation code, prompts, construction utilities, and full documentation are available in the ElephantBench GitHub repository. Load the dataset from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/tencent/ElephantBench.textquestion-answering1K<n<10K6 likes751 downloads25d agoHugging Face26MoreThought /DeepSWEGym-Edu Dataset Description This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Edu.texttext-generation10K<n<100K2 likes697 downloads19d agoHugging Face27eddie-OB /gsm8k-multilingual-reasoning gsm8k-multilingual-reasoning GSM8K with reasoning translated to multiple languages Schema {"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}} Usage from datasets importload_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K1 likes683 downloads8mo agoHugging Face28silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes670 downloads3y agoHugging Face29wofmanaf /ego4d-videoEgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified. For mored details, please visit EgoCOT_Dataset. If you find this dataset useful, please consider citing the paper, @article{mu2024embodiedgpt, title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page: https://huggingface.co/datasets/wofmanaf/ego4d-video.textquestion-answering100K<n<1M16 likes667 downloads2y agoHugging Face30QCRI /AraDICE-ArabicMMLU-egy AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.texttext-classification10K<n<100K1 likes551 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.