CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3.1k downloads3y agoHugging Face02gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads4mo agoHugging Face03AIM-Harvard /MedBrowseComp MedBrowseComp Dataset This repository contains datasets for medical information-seeking-oriented deep research and computer use tasks. Datasets The repository contains three harmonized datasets: MedBrowseComp_50: A collection of 50 medical entries for browsing and comparison. MedBrowseComp_605: A comprehensive collection of 605 medical entries. MedBrowseComp_CUA: A curated collection of medical data for comparison and analysis. Usage These datasets can be… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp.textquestion-answering1K<n<10K8 likes288 downloads1y agoHugging Face04fhyfhy /diffusion-vs-ar-hard-sudoku Diffusion vs AR Hard Sudoku This repository packages 8,148,696 Sudoku examples in the CSV format expected by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline. Every processed file has these columns: column meaning quizzes 81 row-major digits; 0 is an empty cell solutions complete 81-digit solution source original collection dataset normalized dataset family official_rating rating supplied by the source rating_type semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.tabularquestion-answering1M<n<10M0 likes74 downloads2mo agoHugging Face05AIM-Harvard /rabbit_b4bqaThis is the drug-matching dataset between generic and brand keywords for the RABBIT leaderboard 🐰. And here is the paper: arxiv @misc{gallifant2024language, title={Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks}, author={Jack Gallifant and Shan Chen and Pedro Moreira and Nikolaj Munch and Mingye Gao and Jackson Pond and Leo Anthony Celi and Hugo Aerts and Thomas Hartvigsen and Danielle Bitterman}, year={2024}, eprint={2406.12066}… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/rabbit_b4bqa.textquestion-answering1K<n<10K0 likes21 downloads2y agoHugging Face06prithivMLmods /Math-Forge-Hard Math-Forge-Hard Dataset Overview The Math-Forge-Hard dataset is a collection of challenging math problems designed to test and improve problem-solving skills. This dataset includes a variety of word problems that cover different mathematical concepts, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math word problems. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Forge-Hard.texttext-generation1K<n<10K6 likes16 downloads2y agoHugging Face07haris001 /jsoncodes Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/haris001/jsoncodes.textquestion-answeringn<1K0 likes10 downloads3y agoHugging Face08harshi321 /netflix-movies_showstexttext-classification1K<n<10K3 likes10 downloads2y agoHugging Face09har123ish /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes10 downloads1mo agoHugging Face10harshi321 /Questionstextquestion-answeringn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.