datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aneumo
Aneumo Datasets
AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis.
BioKinemaEndoBench
EndoBench
🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper
This repository is the official implementation of the paper EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis.
🚀 News
[03/2026] We release the EndoVQA-Instruct Dataset at here.
[21/10/2025] We release a new open-set challenging VQA benchmark EndoBench-extended.
[19/09/2025] 🎉🎉Our EndoBench was accepted by NeurIPS'25 D&B Track!!!
☀️ Tutorial
EndoBench is… See the full description on the dataset page: https://huggingface.co/datasets/Saint-lsy/EndoBench.andrade-law-saint-paul-spatial-index
Andrade Law — Saint Paul Service-Area Spatial Index
Open spatial-reference data for Andrade Law PLLC, a personal-injury law firm in Saint Paul, Minnesota. It maps the firm's office and its Saint Paul service-area landmarks to their S2 Geometry cells and WGS84 coordinates.
S2 cells are an open geometric indexing system; they are used here as geographic reference labels, not as an official or administrative identifier.
Files
Canonical home: these files are… See the full description on the dataset page: https://huggingface.co/datasets/Gabe-Andrade-Attorney/andrade-law-saint-paul-spatial-index.streamlit_docsgcp-cloud-billing-costWikipedia-Corpora-Report
Dataset Card for "Wikipedia-Corpora-Report"
This dataset is used as a metadata database for the online WIKIPEDIA CORPORA META REPORT dashboard that illustrates how humans and bots generate or edit Wikipedia editions and provides metrics for “pages” and “edits” for all Wikipedia editions (320 languages). The “pages” metric counts articles and non-articles, while the “edits” metric tallies edits on articles and non-articles, all categorized by contributor type: humans or bots. The… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Wikipedia-Corpora-Report.MASD
Dataset Card for "Masked Arab States Dataset (MASD)"
This dataset is created using 20 Arab States1 with their corresponding capital cities, nationalities, currencies, and on which continents they are located, consisting of four categories: country-capital
prompts, country-currency prompts, country-nationality prompts, and country-continent prompts. Each prompts category has 40 masked prompts, and the total number of masked prompts in the MASD dataset is 160. This dataset is used to… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/MASD.Mental_Health_Condition_ClassificationThis dataset consists of textual descriptions related to various mental health conditions, aimed at enabling natural language processing (NLP) tasks such as emotion detection, condition classification, and sentiment analysis. The dataset includes a diverse range of examples that reflect real-world mental health challenges, providing valuable insights into emotions, thought patterns, and behavioral states associated with different conditions.
Researchers, developers, and mental health… See the full description on the dataset page: https://huggingface.co/datasets/sai1908/Mental_Health_Condition_Classification.mmlu-legal-dataset-mcqARQGData
ARQGData Corpus
Dataset Description
This repository contains the complete ARQGData Corpus, a curated dataset designed for Arabic Automatic Question Generation (AQG).
It is intended for use in training, testing, and evaluation of deep learning models. The dataset provides high-quality,
diverse examples to facilitate research and development in Arabic natural language processing and educational technology applications.
Language: Arabic
Dataset Type: CSV
Size:… See the full description on the dataset page: https://huggingface.co/datasets/saidlafkiar82/ARQGData.ASAD
Dataset Card for "Arab States Analogy Dataset (ASAD)"
This dataset is created using 20 Arab States1 with their corresponding capital cities, nationalities, currencies, and on which continents they are located, consisting of four sets: country-capital set, country-currency set, country-nationality set, and country-continent set. Each set has 380 word analogies, and the total number of word analogies in the ASAD dataset is 1520. This dataset is used to evaluate Arabic Word Embedding… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/ASAD.moviedbResume-Screening-DatasetPokemon_Dataiononospheresabay_sai_travel_data
Sabay Sai Travel Club Dataset
This dataset contains information about tours, stays, guides, and travel tips for Kazakhstan, curated by Sabay Sai Travel Club. It is intended for AI, travel recommendation systems, and general data exploration. Each entry includes details about the experience, location, description, URL,
Project
ACI-BENCH
Introduction
This repository contains the data and source code for:
Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation". Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, Meliha Yetisgen. Submitted to Nature Scientific Data, 2023.
https://www.nature.com/articles/s41597-023-02487-3
@article{aci-bench,
author = {Wen{-}wai Yim and
Yujuan Fu and
Asma {Ben Abacha}… See the full description on the dataset page: https://huggingface.co/datasets/saidabizi/Project.legal-summarizationSentimentsdlgenai_ProteinPredictionlogscode-switching-codesaviours-si26-saima
Code-Switching Urdu-English Dataset
Dataset Description
This dataset is created for Urdu-English code-switching language identification. It contains sentences that include Urdu words, English words, and a small number of mixed-language entries.
The dataset was prepared by collecting and organizing code-switched Urdu-English sentences. Each sentence was divided into individual words, and every word was assigned a language label. The data was then converted into a… See the full description on the dataset page: https://huggingface.co/datasets/Saima-Manzoor/code-switching-codesaviours-si26-saima.jancokLigandssail_lid
Dataset Card for SAIL 2017
Dataset Summary
The dataset was a part of Shared Task on Sentiment Analysis in Indian Languages (SAIL) Tweets. It was presented in FIRE 2017.
Languages
Code-Mixed sentences in English and Hindi
Source Data
http://amitavadas.com/SAIL/data.html
Initial Data Collection and Normalization
All the data from the source is collected and cleaned. Punctuations, Special characters and Emoticons are removed.
jerry_seinfeld_dialoguesMultilingual_T2I_clean_llama2_templated_promptsconvAIresponsesail
