CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.3k downloads6mo agoHugging Face02Sai452 /BioKinematabularn<1K0 likes163 downloads4mo agoHugging Face03Saint-lsy /EndoBench EndoBench 🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper This repository is the official implementation of the paper EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis. 🚀 News [03/2026] We release the EndoVQA-Instruct Dataset at here. [21/10/2025] We release a new open-set challenging VQA benchmark EndoBench-extended. [19/09/2025] 🎉🎉Our EndoBench was accepted by NeurIPS'25 D&B Track!!! ☀️ Tutorial EndoBench is… See the full description on the dataset page: https://huggingface.co/datasets/Saint-lsy/EndoBench.imagequestion-answering1K<n<10K8 likes142 downloads7mo agoHugging Face04Gabe-Andrade-Attorney /andrade-law-saint-paul-spatial-index Andrade Law — Saint Paul Service-Area Spatial Index Open spatial-reference data for Andrade Law PLLC, a personal-injury law firm in Saint Paul, Minnesota. It maps the firm's office and its Saint Paul service-area landmarks to their S2 Geometry cells and WGS84 coordinates. S2 cells are an open geometric indexing system; they are used here as geographic reference labels, not as an official or administrative identifier. Files Canonical home: these files are… See the full description on the dataset page: https://huggingface.co/datasets/Gabe-Andrade-Attorney/andrade-law-saint-paul-spatial-index.geospatialn<1K0 likes100 downloads13d agoHugging Face05sai-lohith /streamlit_docstextn<1K0 likes85 downloads2y agoHugging Face06sairamn /gcp-cloud-billing-costtabular100K<n<1M0 likes77 downloads2y agoHugging Face07SaiedAlshahrani /Wikipedia-Corpora-Report Dataset Card for "Wikipedia-Corpora-Report" This dataset is used as a metadata database for the online WIKIPEDIA CORPORA META REPORT dashboard that illustrates how humans and bots generate or edit Wikipedia editions and provides metrics for “pages” and “edits” for all Wikipedia editions (320 languages). The “pages” metric counts articles and non-articles, while the “edits” metric tallies edits on articles and non-articles, all categorized by contributor type: humans or bots. The… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Wikipedia-Corpora-Report.text1K<n<10K0 likes70 downloads3y agoHugging Face08SaiedAlshahrani /MASD Dataset Card for "Masked Arab States Dataset (MASD)" This dataset is created using 20 Arab States1 with their corresponding capital cities, nationalities, currencies, and on which continents they are located, consisting of four categories: country-capital prompts, country-currency prompts, country-nationality prompts, and country-continent prompts. Each prompts category has 40 masked prompts, and the total number of masked prompts in the MASD dataset is 160. This dataset is used to… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/MASD.textn<1K3 likes47 downloads3y agoHugging Face09sai1908 /Mental_Health_Condition_ClassificationThis dataset consists of textual descriptions related to various mental health conditions, aimed at enabling natural language processing (NLP) tasks such as emotion detection, condition classification, and sentiment analysis. The dataset includes a diverse range of examples that reflect real-world mental health challenges, providing valuable insights into emotions, thought patterns, and behavioral states associated with different conditions. Researchers, developers, and mental health… See the full description on the dataset page: https://huggingface.co/datasets/sai1908/Mental_Health_Condition_Classification.texttext-classification100K<n<1M9 likes46 downloads2y agoHugging Face10SaiCharanChetpelly /mmlu-legal-dataset-mcqtextquestion-answering1K<n<10K1 likes42 downloads2y agoHugging Face11saidlafkiar82 /ARQGData ARQGData Corpus Dataset Description This repository contains the complete ARQGData Corpus, a curated dataset designed for Arabic Automatic Question Generation (AQG). It is intended for use in training, testing, and evaluation of deep learning models. The dataset provides high-quality, diverse examples to facilitate research and development in Arabic natural language processing and educational technology applications. Language: Arabic Dataset Type: CSV Size:… See the full description on the dataset page: https://huggingface.co/datasets/saidlafkiar82/ARQGData.text10K<n<100K0 likes42 downloads1mo agoHugging Face12SaiedAlshahrani /ASAD Dataset Card for "Arab States Analogy Dataset (ASAD)" This dataset is created using 20 Arab States1 with their corresponding capital cities, nationalities, currencies, and on which continents they are located, consisting of four sets: country-capital set, country-currency set, country-nationality set, and country-continent set. Each set has 380 word analogies, and the total number of word analogies in the ASAD dataset is 1520. This dataset is used to evaluate Arabic Word Embedding… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/ASAD.text1K<n<10K1 likes37 downloads3y agoHugging Face13saikiranmaddukuri /moviedbtabular10K<n<100K0 likes31 downloads28d agoHugging Face14saikushalreddy25 /Resume-Screening-Datasettext10K<n<100K0 likes31 downloads5d agoHugging Face15Saif-M /Pokemon_Datatabular1K<n<10K1 likes23 downloads2y agoHugging Face16IRT-SaintExupery /iononospheretabular1M<n<10M0 likes22 downloads1y agoHugging Face17Sabaysai /sabay_sai_travel_data Sabay Sai Travel Club Dataset This dataset contains information about tours, stays, guides, and travel tips for Kazakhstan, curated by Sabay Sai Travel Club. It is intended for AI, travel recommendation systems, and general data exploration. Each entry includes details about the experience, location, description, URL, textn<1K0 likes15 downloads11mo agoHugging Face18saidabizi /Projectgated ACI-BENCH Introduction This repository contains the data and source code for: Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation". Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, Meliha Yetisgen. Submitted to Nature Scientific Data, 2023. https://www.nature.com/articles/s41597-023-02487-3 @article{aci-bench,   author = {Wen{-}wai Yim and                 Yujuan Fu and                 Asma {Ben Abacha}… See the full description on the dataset page: https://huggingface.co/datasets/saidabizi/Project.textn<1K0 likes14 downloads1y agoHugging Face19SaiCharanChetpelly /legal-summarizationtextn<1K0 likes13 downloads2y agoHugging Face20Hemanth-Sai /Sentimentstabulartext-classification1K<n<10K0 likes12 downloads3y agoHugging Face21SaiPalH /dlgenai_ProteinPredictionlogstabularn<1K0 likes11 downloads9mo agoHugging Face22Saima-Manzoor /code-switching-codesaviours-si26-saima Code-Switching Urdu-English Dataset Dataset Description This dataset is created for Urdu-English code-switching language identification. It contains sentences that include Urdu words, English words, and a small number of mixed-language entries. The dataset was prepared by collecting and organizing code-switched Urdu-English sentences. Each sentence was divided into individual words, and every word was assigned a language label. The data was then converted into a… See the full description on the dataset page: https://huggingface.co/datasets/Saima-Manzoor/code-switching-codesaviours-si26-saima.text1K<n<10K0 likes11 downloads2mo agoHugging Face23saipangon /jancoktabularn<1K0 likes10 downloads10mo agoHugging Face24Sai452 /Ligandstext100M<n<1B0 likes10 downloads5mo agoHugging Face25ashishkmr2094 /sail_lid Dataset Card for SAIL 2017 Dataset Summary The dataset was a part of Shared Task on Sentiment Analysis in Indian Languages (SAIL) Tweets. It was presented in FIRE 2017. Languages Code-Mixed sentences in English and Hindi Source Data http://amitavadas.com/SAIL/data.html Initial Data Collection and Normalization All the data from the source is collected and cleaned. Punctuations, Special characters and Emoticons are removed. texttoken-classification10K<n<100K0 likes9 downloads4y agoHugging Face26saikatkumardey /jerry_seinfeld_dialoguestext10K<n<100K0 likes9 downloads3y agoHugging Face27sainv /Multilingual_T2I_clean_llama2_templated_promptstextn<1K0 likes9 downloads2y agoHugging Face28saikrishna759 /convAItext10K<n<100K0 likes8 downloads3y agoHugging Face29saikat02004 /responsetextn<1K0 likes7 downloads3y agoHugging Face30berger815 /sailtextn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.