CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yeshenglong /InternSpatialtext1M<n<10M3 likes4k downloads10mo agoHugging Face02allenai /qasper-yesnotextn<1K0 likes2.4k downloads1y agoHugging Face03allenai /sciriff-yesnotext1K<n<10K1 likes2.4k downloads1y agoHugging Face04Yesianrohn /OCR-Data OCR Text Detection and Recognition Dataset Dataset Description A large-scale, multi-source OCR dataset aggregating 14 public benchmarks for text detection and recognition in both scene images and handwritten documents. Each image is paired with: Transcribed text for each text region Bounding boxes (axis-aligned rectangles) for each text region Polygon coordinates (precise boundary points) for each text region The dataset is stored in HuggingFace Parquet format with… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/OCR-Data.imageobject-detection100K<n<1M8 likes1.7k downloads6mo agoHugging Face05yesilhealth /Health_Benchmarks LLM Health Benchmarks Dataset by Yesil Science The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge. Primary Purpose This dataset is built to: Benchmark LLMs in medical specialties and subfields. Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.textquestion-answering1K<n<10K10 likes445 downloads1y agoHugging Face06nanyy1025 /bioasq_7b_yesnotextn<1K2 likes418 downloads3y agoHugging Face07yeshpanovrustem /100k_movie_reviews_from_kzgated 100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts Dataset Summary This repository provides a publicly available corpus of 100,502 movie reviews collected from kino.kz, spanning 2001–2025 and covering 4,943 unique movie titles. The dataset is multilingual and reflects a Kazakhstan-specific online setting where reviews are predominantly written in Russian, with smaller subsets in Kazakh and Kazakh–Russian code-switched text. Reviews are… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/100k_movie_reviews_from_kz.texttext-classification100K<n<1M1 likes353 downloads5mo agoHugging Face08jmhb /bioasq_yesno_trainv0_n1464_test100tabular1K<n<10K0 likes306 downloads1y agoHugging Face09bdjafer /msmarco-yesnotabular10K<n<100K0 likes281 downloads4y agoHugging Face10HajarGH /bioasq-yesno-cleantext1K<n<10K0 likes158 downloads7mo agoHugging Face11yeshpanovrustem /kaznerd A Named Entity Recognition Dataset for Kazakh This is a modified version of the dataset provided in the LREC 2022 paper KazNERD: Kazakh Named Entity Recognition Dataset. The original repository for the paper can be found at https://github.com/IS2AI/KazNERD. Tokens denoting speech disfluencies and hesitations (parenthesised) and background noise [bracketed] were removed. A total of 2,027 duplicate sentences were removed. Statistics for training (Train), validation (Valid)… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/kaznerd.texttoken-classification100K<n<1M9 likes118 downloads1y agoHugging Face12Yesianrohn /MLT2019text1K<n<10K0 likes115 downloads3mo agoHugging Face13Lots-of-LoRAs /task380_boolq_yes_no_question Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task380_boolq_yes_no_question Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task380_boolq_yes_no_question.texttext-generation1K<n<10K0 likes103 downloads2y agoHugging Face14Ayush-Singh /reward-bench-Qwen2.5-3B-yes-notabular1K<n<10K0 likes94 downloads2y agoHugging Face15Ayush-Singh /reward-bench-chatgpt-4o-latest-yes-notabular1K<n<10K0 likes94 downloads2y agoHugging Face16Ayush-Singh /reward-bench-Phi-3-mini-128k-instruct-yes-notabular1K<n<10K0 likes89 downloads2y agoHugging Face17Lots-of-LoRAs /task362_spolin_yesand_prompt_response_sub_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.texttext-generation1K<n<10K0 likes88 downloads2y agoHugging Face18Yeshenyue /hotpot_qa Dataset Card for "hotpot_qa" Dataset Summary HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason… See the full description on the dataset page: https://huggingface.co/datasets/Yeshenyue/hotpot_qa.textquestion-answering100K<n<1M0 likes75 downloads6mo agoHugging Face19Ayush-Singh /reward-bench-Qwen2.5-7B-Instruct-yes-notabular1K<n<10K0 likes73 downloads1y agoHugging Face20Ayush-Singh /reward-bench-Llama-3.2-1B-yes-notabular1K<n<10K0 likes72 downloads2y agoHugging Face21Ayush-Singh /reward-bench-Llama-3.2-3B-yes-notabular1K<n<10K0 likes72 downloads2y agoHugging Face22Ayush-Singh /reward-bench-Qwen2.5-0.5B-yes-notabular1K<n<10K0 likes72 downloads2y agoHugging Face23pajacques /yestext100K<n<1M0 likes72 downloads2y agoHugging Face24Ayush-Singh /reward-bench-Llama-2-13b-chat-hf-yes-notabular1K<n<10K0 likes70 downloads2y agoHugging Face25Ayush-Singh /reward-bench-Qwen2.5-7B-yes-notabular1K<n<10K0 likes70 downloads2y agoHugging Face26Ayush-Singh /reward-bench-Qwen2.5-0.5B-Instruct-yes-notabular1K<n<10K0 likes70 downloads2y agoHugging Face27Ayush-Singh /reward-bench-gemma-2-2b-it-yes-notabular1K<n<10K0 likes68 downloads2y agoHugging Face28Ayush-Singh /reward-bench-gpt2-yes-notabular1K<n<10K0 likes67 downloads2y agoHugging Face29Ayush-Singh /reward-bench-gpt-4o-mini-yes-notabular1K<n<10K0 likes67 downloads2y agoHugging Face30Ayush-Singh /reward-bench-pythia-6.9b-yes-notabular1K<n<10K0 likes67 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.