CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chiuratto-AIgourakis /sounio-code-examples Sounio Curated Code Examples Curated compile-clean .sio examples for training and evaluating code models on Sounio, a self-hosted systems and scientific programming language for epistemic computing, uncertainty propagation, and algebraic effects. This directory is the Cx-1 expansion lane for chiuratto-AIgourakis/sounio-code-examples. Current batch Examples: 5,000 Metadata files: 5,000 Compiler gate: bin/souc check pass rate 5,000/5,000 Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.texttext-generation1K<n<10K0 likes2.8k downloads4mo agoHugging Face02datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes2.7k downloads3y agoHugging Face03iamroot /chat_formatted_examplestextn<1K0 likes2.3k downloads2y agoHugging Face04typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2k downloads2y agoHugging Face05R2MED /MedXpertQA-Exam 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedXpertQA-Exam.texttext-retrieval10K<n<100K1 likes1.3k downloads1y agoHugging Face06mjbommar /opengloss-v1.3-query-examples-flat See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 (Flattened) Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.texttext-generation100K<n<1M0 likes689 downloads18d agoHugging Face07stevenhsu123 /chinese_exam_train_datatext1K<n<10K0 likes548 downloads3y agoHugging Face08natnitaract /exams-basic-and-quantum-cryptography-and-security-latex Open Problem Exams: Cryptography and Security (LaTeX) A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions. Dataset Overview Institution Files Topics Questions Caltech & TU Delft 8 38 145 EPFL 6 19 86 ETH Zurich 1 14 37 MIT 3 33 79 Total 18 104 347 Difficulty Distribution Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.textn<1K1 likes463 downloads6mo agoHugging Face09RJTPP /thai_exam-reformattedReformatted version of scb10x/thai_exam Additional Changes: Fix math incorrect answer ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด? a. \log_{3} 2 b. \log_{3} 4 c. \log_{3} \frac{33}{4} d. \log_{3} 10 e. \log_{3} 12 # Original answer: d (\log_{3} 10) # Correction : b (\log_{3} 4) textquestion-answering1K<n<10K0 likes370 downloads1y agoHugging Face10kunato /thai-exam-seacrowdtextn<1K1 likes352 downloads2y agoHugging Face11amu-cai /medical-exams-LDEK-EN-2013-2024 Dataset Card for medical-exams-LDEK-EN-2013-2024 Dataset Description This is a dataset used and described in: @article{grzybowski2024polish, title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment}, author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek}, journal={arXiv preprint arXiv:2412.00559}, year={2024} } Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-EN-2013-2024.textquestion-answering1K<n<10K0 likes327 downloads1y agoHugging Face12sarafuyu /swedish-medical-exams-mcq-1006-json Dataset Card for Swedish Medical Exam MCQs Dataset Description This dataset contains multiple-choice questions from Swedish medical exams. Languages The dataset is in Swedish (sv). Dataset Structure Each entry in the dataset contains the following fields: question: The question options: An array of possible answers answer: The correct answer language: The language of the question (always "sv" for Swedish) country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.textquestion-answering1K<n<10K0 likes304 downloads2y agoHugging Face13northern-64bit /ENADE_Brazilian_national_university_examination_MCQ_483textquestion-answeringn<1K0 likes297 downloads2y agoHugging Face14Gabrui /estonian_language_exams Dataset Card for Estonian Language Proficiency Exam Samples This is part of the initiative from Cohere For AI @CohereForAI to gather exams from around the world to build a new multilingual benchmark. The web Scrapping code can be found at the source_scripts_data_aya repository. The source data can be visually checked at Sõeltestid.pdf and Diagnoostestid.pdf. Dataset Details Dataset Description This dataset contains sample questions from the Estonian language… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/estonian_language_exams.textquestion-answeringn<1K0 likes291 downloads2y agoHugging Face15Dimeiza /IPA_Exam_AM 独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午前) 概要  本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午前問題及び公式解答をセットにした、非公式のデータセットです。  以下、公開されている過去問題pdfから抽出しています。 https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html  人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。 データの詳細  現状、過去5年分(2020〜2024)の以下試験区分について、問題及び回答を収録しています。 応用情報技術者試験(ap) 高度共通午前1(koudo) エンベデッドシステムスペシャリスト(es) 情報処理安全確保支援士(sc) プロジェクトマネージャ(pm) データベーススペシャリスト(db) システム監査技術者(au) ITストラテジスト(st)… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_AM.tabular1K<n<10K0 likes242 downloads2y agoHugging Face16BAAI /OpenSeek-Synthetic-Reasoning-Data-Examples OpenSeek-Reasoning-Data OpenSeek [Github|Blog] Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process. News 🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.text1M<n<10M27 likes203 downloads2y agoHugging Face17169Pi /exambench ExamBench The ExamBench Dataset (~600M tokens, 405k examples) is one of the largest open-source corpora designed for competitive exam preparation and reasoning AI. Generated using advanced distillation techniques, it combines structured chain-of-thought reasoning with comprehensive coverage of over 25 Indian and international examinations. From JEE and NEET to UPSC, Banking, GRE, and IELTS, the dataset spans multiple domains like STEM, humanities, current affairs, language, and… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/exambench.texttext-generation100K<n<1M3 likes197 downloads1y agoHugging Face18Wauplin /example-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code. Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/ JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.textn<1K8 likes194 downloads1y agoHugging Face19camgeodesic /sycophancy_examples Sycophancy Examples Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings". Original source: GeodesicResearch/Obfuscation_Generalization Files File Examples Description sycophancy_opinion_political.jsonl 5,000 Political opinion questions with persona-aligned "sycophantic" answers sycophancy_fact.jsonl 401 Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.texttext-classification1K<n<10K0 likes190 downloads6mo agoHugging Face20f13rnd /multimodal-example Multimodal Example Dataset Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift. Structure ├── train.jsonl # 10 training samples ├── test.jsonl # 2 validation samples ├── images/ # All referenced images (400x300 JPEG) │ ├── dog_portrait.jpg │ ├── forest_river.jpg │ ├── laptop_desk.jpg │ ├── mountain_lake.jpg │ ├── ocean_rocks.jpg │ ├── coffee_cup.jpg │ ├── bookshelf.jpg │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.imagen<1K0 likes179 downloads6mo agoHugging Face21FreedomIntelligence /2023_Pharmacist_Licensure_Examination-TCM_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Traditional Chinese Medicine (TCM) Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/2023_Pharmacist_Licensure_Examination-TCM_track.textn<1K8 likes164 downloads3y agoHugging Face22amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes159 downloads3mo agoHugging Face23overthelex /ua-judge-exam UA-JudgeExam 11,990 four-option multiple-choice items with official answer keys, from the question bank the Higher Qualification Commission of Judges of Ukraine (Вища кваліфікаційна комісія суддів України, VKKS) publishes for the anonymous written testing of candidates for appellate-court judgeships. Source: Commission decision of 15 July 2024, No. 221/зп-24 — five documents, 1,672 PDF pages, covering general legal knowledge plus the administrative, commercial, criminal and… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-judge-exam.tabularmultiple-choice10K<n<100K0 likes146 downloads1mo agoHugging Face24chrimerss /hydro_cali_agent_exampletextn<1K0 likes145 downloads5mo agoHugging Face25bbidpa /flutter-full-examples-v1 Flutter Codegen: Full Examples Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps, there's no step history or diff structure here -- each row is a single, standalone goal -> complete file example. This is the whole-code counterpart to flutter-diff-steps-v1, intended for training/evaluating a baseline that generates the entire file in one shot, to compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.tabulartext-generation10K<n<100K0 likes136 downloads17d agoHugging Face26mpvasilis /thema-panhellenic-exams Thema Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments. Website · Code · Leaderboard · Method Dataset summary Current release Exam years 2023 to 2026 Published subjects 10 of 10 Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.tabularquestion-answeringn<1K0 likes129 downloads2mo agoHugging Face27tgonzalez95 /marketing_exam1textn<1K1 likes122 downloads3y agoHugging Face28amu-cai /medical-exams-LEK-EN-2013-2024 Dataset Card for medical-exams-LEK-EN-2013-2024 Dataset Description This is a dataset used and described in: @article{grzybowski2024polish, title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment}, author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek}, journal={arXiv preprint arXiv:2412.00559}, year={2024} } Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-EN-2013-2024.textquestion-answering1K<n<10K0 likes117 downloads1y agoHugging Face29vudang449 /AIO2025-Final-Exam-Datasetversion https://git-lfs.github.com/spec/v1 oid sha256:072d081dec61aad02275a7ed608c136a28f1de6672c7ff3a520bdbf7b0952e53 size 2225 tabularn<1K0 likes108 downloads3mo agoHugging Face30UniversalCEFR /cambridge_exams_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following: Dataset License: cc-by-nc-sa-4.0 Dataset Repository: https://ilexir.co.uk/datasets/index.html Original Dataset Paper: Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016. Text Readability… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cambridge_exams_en.textn<1K1 likes107 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.