datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
e-CARE
Dataset of (Du et al., 2022) (Unofficial reupload)
Abstract
Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal fact to facilitate the causal reasoning process. However, such explanation information still remains absent in existing causal reasoning resources. In this paper, we fill this gap by presenting… See the full description on the dataset page: https://huggingface.co/datasets/12ml/e-CARE.CareQA
CareQA
Dataset Summary
CareQA is a healthcare QA dataset with two versions:
Closed-Ended Version: A multichoice question answering (MCQA) dataset containing 5,621 QA pairs across six categories. Available in English and Spanish.
Open-Ended Version: A free-response dataset derived from the closed version, containing 2,769 QA pairs (English only).
The dataset originates from… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA.Q-CARE
Q-CARE Benchmark
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Jeonghwan Choi · Taewon Yun · Minjeong Ban · Gyeonghun Sun · Jae-Gil Lee · Hwanjun Song
Korea Advanced Institute of Science and Technology (KAIST) · COLM 2026
📄 Paper · 💻 Code
Q-CARE is a query-agnostic, fully reference-free framework for evaluating
retrieval-augmented generation. It decomposes queries into sub-queries and
answers into atomic claims, then scores retrieval and… See the full description on the dataset page: https://huggingface.co/datasets/DISLab/Q-CARE.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.gspc-care
GSPC — care bank (CareBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the care row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=care (family, kind, status and n are on that row, never typed here; the whole… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-care.caresCcaresAUS-Child-Care-ProvidersPer-state child care provider licensing data. Each U.S. state is exposed as a
separate dataset configuration: pick a state in the dataset viewer, or pass its
name as the config argument, e.g.
load_dataset("<repo>", "alabama").
US Child Care Providers
The goal of this dataset is to collect data on child care providers in each of the states within the United States. The data collection sources are public themselves with
this effort representing bundling them together to make the… See the full description on the dataset page: https://huggingface.co/datasets/ryan-koch/US-Child-Care-Providers.SWE-CARE
SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation
Dataset Description
SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.NBA-Player-Career-Stats
Dataset Description
This dataset contains a single CSV file with lifetime statistics for NBA players. The data includes various box score stats and personal information for each player's career.
Data Fields
The CSV file contains the following columns:
FULL_NAME: The player's full name
AST: Total career assists
BLK: Total career blocks
DREB: Total career defensive rebounds
FG3A: Total 3-point field goal attempts
FG3M: Total 3-point field goals made
FG3_PCT: 3-point field… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/NBA-Player-Career-Stats.CareMedEval CareMedEval dataset
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Medical Field
Links
Repository
GitHub, HuggingFace
Paper
https://arxiv.org/abs/2511.03441
Contact
Doria BONZI, Alexandre GUIGGI
DescriptionCareMedEval (Critical appraisal and Reasoning Medical Evaluation) is a French Multiple Choice Question Answering (MCQA) dataset focused on evaluating critical appraisal skills in the medical field for scientific articles, as practiced in… See the full description on the dataset page: https://huggingface.co/datasets/doriab/CareMedEval.CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.CaReBench
CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval
Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang
🤗 Model | 🤗 Data | 📑 Paper
📝 Introduction
🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation.
📊 ReBias… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.e-caremiluim-career-bridge
Miluim Career Bridge - Dataset
Synthetic Hebrew-context job-ad pairs translating Israeli reserve-service (miluim) experience into civilian job language, across 8 categories. Generated with Qwen/Qwen2.5-1.5B-Instruct (see notebooks/01_generation.ipynb for the full pipeline).
⚠️ Note: the Hub split shown above (train) is a technical wrapper around the single parquet file. The dataset's actual train/test division is the split column inside the data (9536 train / 1000 test) - filter… See the full description on the dataset page: https://huggingface.co/datasets/maayan890/miluim-career-bridge.nba-career-stats-eda
🏀 NBA Player Career Stats — EDA Project
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA) of NBA player
career statistics. The goal is to uncover patterns in player performance, compare
active vs. retired players, and explore relationships between key basketball stats.
Source: Hatman/NBA-Player-Career-Stats
Original size: 3,093 rows × 28 columns
Final clean size: 3,078 rows × 23 columns
Target Variable: IS_ACTIVE (True = Active / False =… See the full description on the dataset page: https://huggingface.co/datasets/Omerinbar/nba-career-stats-eda.CARES-18K
Dataset Card for "CARES-18K"
CARES-18K: Clinical Adversarial Robustness and Evaluation of Safety
CARES-18K is a benchmark dataset for evaluating the safety and robustness of large language models (LLMs) in clinical and healthcare contexts. It consists of over 18,000 synthetic prompts generated across 8 medical safety principles, 4 graded harmfulness levels (0–3), and 4 prompting strategies (direct, indirect, obfuscation, role-play). These prompts probe both LLM… See the full description on the dataset page: https://huggingface.co/datasets/HFXM/CARES-18K.CARE-Bench
CARE-Bench v1.1 public data
GitHub link to the project
The public release contains 439 cases and 925 labeled prefixes.
Split
Cases
Prefixes
Development
284
609
Validation
93
181
Public Test 1
62
135
The prefix labels are distributed as follows:
Label
Prefixes
Information needed
249
Self-care or monitor
162
Nonurgent care
342
Urgent care
172
The Dataset Viewer includes four configurations:
model_inputs: identifiers, split information… See the full description on the dataset page: https://huggingface.co/datasets/ningkko/CARE-Bench.Long-Term-Care-Aggregated-Data
Project 1 Proposal of the Long Term Care(LTC) Aggregated Dataset
KAO, HSUAN-CHEN(Justin)
NetID: hk310
Dataset Details
The long-term care aggregated dataset, essential for conducting experience studies, is an extensive and valuable compilation of variables central to the analysis and prediction of long-term care (LTC) insurance products. This dataset integrates two critical files: one detailing claim incidence and the other capturing policy terminations. This merger is… See the full description on the dataset page: https://huggingface.co/datasets/mastergopote44/Long-Term-Care-Aggregated-Data.plant_care
🥦 Cauliflower Disease Detection Dataset
A curated computer vision dataset for automatic detection and classification of cauliflower leaf diseases, designed for training and evaluating deep learning models in agricultural and plant pathology applications.
This dataset is suitable for image classification, object detection, and transfer learning workflows and is provided in YOLO-compatible format.
📌 Dataset Overview
Cauliflower crops are highly susceptible to… See the full description on the dataset page: https://huggingface.co/datasets/indra17/plant_care.floradb-houseplants-care-sample
🌿 FloraDB — Houseplant Care & Pet-Toxicity Dataset (Free Sample)
Full dataset: floradb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of FloraDB: a structured dataset that turns subjective houseplant care advice — "bright indirect light", "water when dry" — into quantitative engineering metrics (Lux thresholds, watering-day intervals, temperature and humidity ranges), joined to ASPCA dog/cat toxicity and grounded on the GBIF… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/floradb-houseplants-care-sample.career-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/Vera-001/ats-career-page-urls.carechurch-os-community
CareChurch OS Community
AI-Native Church Community Management Platform — Monorepo
Architecture
carechurch-os-community/
├── apps/
│ ├── mobile/ # React Native Expo (Expo Router)
│ ├── admin/ # Next.js 15 Web Dashboard
│ └── api/ # NestJS Backend
├── packages/
│ ├── types/ # Shared TypeScript types & enums
│ ├── validation/ # Zod schemas
│ ├── utils/ # Geo, time, formatting utilities
│ ├── ui/ # Design tokens & shared… See the full description on the dataset page: https://huggingface.co/datasets/Suriyong/carechurch-os-community.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.care-xai
CARE-XAI: Culturally-Aware, Evidence-Grounded Explainable AI for Health
17,803 total rows · 14,254 train / 1,797 validation / 1,752 test · 5 sources · 3 labels
CARE-XAI is a unified health-claim verification dataset consolidating five public health NLP benchmarks into a single schema, augmented with GRADE evidence quality labels, cultural relevance flags, and Gold/Silver explanation annotations.
Sources
Dataset
Rows
%
Explanation
License
PUBHEALTH
9,804
55.1%… See the full description on the dataset page: https://huggingface.co/datasets/Prabhjotschugh/care-xai.ResumeExtractBench
ResumeExtractBench
ResumeExtractBench is a benchmark for schema-guided structured extraction from resume documents. Given a resume PDF and a JSON Schema, systems must return structured data covering personal details, work history, education, skills, and more.
Dataset Size: 38 documents (handwritten + adversarial distractors)
Schema Sections Scored: 9 (basics, experience, education, projects, summary, certifications, awards, volunteering, skills)
Domains: 6 (engineering… See the full description on the dataset page: https://huggingface.co/datasets/Careerflow/ResumeExtractBench.tend
TEND (Gold)
This dataset publishes execution-validated gold-tier examples from the
TEND pipeline: natural-language questions paired with
SQL schema, gold SQL, generated MongoDB schema/query, and plain-English
documentation. It is designed for multi-task research spanning Text→SQL,
SQL→MongoDB, and MongoDB→Documentation.
Every published row is execution-validated. For each example, the pipeline
runs the gold sql_query on PostgreSQL and the generated nosql_query
on MongoDB, then… See the full description on the dataset page: https://huggingface.co/datasets/care2achieve/tend.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.
