datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-care
GSPC — care bank (CareBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the care row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=care (family, kind, status and n are on that row, never typed here; the whole… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-care.SWE-CARE
SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation
Dataset Description
SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.tend
TEND (Gold)
This dataset publishes execution-validated gold-tier examples from the
TEND pipeline: natural-language questions paired with
SQL schema, gold SQL, generated MongoDB schema/query, and plain-English
documentation. It is designed for multi-task research spanning Text→SQL,
SQL→MongoDB, and MongoDB→Documentation.
Every published row is execution-validated. For each example, the pipeline
runs the gold sql_query on PostgreSQL and the generated nosql_query
on MongoDB, then… See the full description on the dataset page: https://huggingface.co/datasets/care2achieve/tend.carebench
CareBench
CareBench is a benchmark for process-aware evaluation of clinical LLM agents. Each case places an agent in a simulated patient encounter in which it must gather evidence through interaction, commit to a diagnosis, and propose management, while safety-critical red flags are tracked throughout the episode. Scoring is process-aware: in addition to diagnostic accuracy, the benchmark measures evidence coverage, premature closure, red-flag coverage, and unsafe commitments.… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/carebench.CARE
Introduction
GitHub Repo
Paper
CARE is a multilingual, multicultural human preference dataset, used for tuning culturally adaptive models.
We curate 3,490 culture-specific questions from diverse resources (including instruction datasets, cultural knowledge bases, and regional social media platforms).
We then collect responses to them from multiple LLMs (e.g. GPT-4o) for each prompt, resulting in a total of 31.7k samples.
Finally we instruct native annotators to rate each… See the full description on the dataset page: https://huggingface.co/datasets/geyang627/CARE.caregiver-summary-night-shiftjhana-sentences
Dataset Card for "Jhana Sentences"
Dataset Description
Dataset Summary
This dataset, named "Jhana Sentences," contains sentences and passages focused on Jhana meditation practices, teachings, and insights. It is intended for use in training language models for applications related to meditation guidance, spiritual advice, and related conversational agents.
Supported Tasks
text-generation: The dataset can be used to train models for generating… See the full description on the dataset page: https://huggingface.co/datasets/carecodeconnect/jhana-sentences.pbd-autism-caregiver
Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture
Dataset Description
This dataset accompanies the paper "Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture". The system is a privacy-by-design multi-agent architecture integrating Retrieval-Augmented Generation (RAG), Data Loss Prevention (DLP), consent management, explainable AI (XAI), and audit… See the full description on the dataset page: https://huggingface.co/datasets/Ionutcroitoru/pbd-autism-caregiver.care
CARE: Clinical Assessment of Robustness and Equity
Resumo
O CARE (Clinical Assessment of Robustness and Equity) é um dataset de avaliação focado na mensuração de viés explícito e robustez de LLMs no contexto do Sistema Único de Saúde (SUS). Enquanto benchmarks tradicionais avaliam o conhecimento médico, o CARE avalia o comportamento e a consistência do modelo frente a grupos populacionais vulneráveis.
O conjunto de dados opera sob a premissa de perguntas invariantes:… See the full description on the dataset page: https://huggingface.co/datasets/Larxel/care.Skin_diseases_and_care
Skin Related Problems Dataset
Description
This dataset contains information on various skin-related problems and conditions. The data has been scraped from reliable sources and consolidated into a single dataset for ease of access and analysis. Each entry includes a topic and detailed information about the skin condition.
Dataset Structure
The dataset is organized into the following columns:
Topic: The title or main topic of the entry.
Information: Detailed… See the full description on the dataset page: https://huggingface.co/datasets/brucewayne0459/Skin_diseases_and_care.care_pro
Introduction
GitHub repository
CARE-Pro is a cross-lingual and cross-cultural evaluation dataset for realistic multilingual information needs.
The dataset contains both fine-grained insider regional knowledge, which focus on region-specific facts and practices, and cross-cultural questions, which ask about another region from a foreign-language perspective.
The current release contains 775 examples in Chinese, Hindi, and Spanish, with English translations provided when available.… See the full description on the dataset page: https://huggingface.co/datasets/geyang627/care_pro.Multilingual-Nepali-Customer-Care-Services-DatasetLHP-Career-Impact-Metrics
[ENTITY_METRICS]: CAREER_IMPACT_ANALYSIS
SUBJECT: Dr. Luis Henrique Leonardo Pereira
STATUS: GOOGLE_KNOWLEDGE_GRAPH_VERIFIED
AUTHORITY_MILESTONES:
2025_CYCLE:
domain: "Men's Health & Sexology"
status: "Global Reference"
validation_source: "Google Algorithms (Organic Snippets)"
2026_CYCLE:
domain: "AI Vector Audit (Gemini/Transformers)"
status: "Technical Authority / L0 Auditor"
validation_source: "Alphabet Ecosystem Integration"… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/LHP-Career-Impact-Metrics.Customer-Care-Services-Dataset-in-Nepalismolified-career-advisor-bot
🤏 smolified-career-advisor-bot
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-career-advisor-bot.
📦 Asset Details
Origin: Smolify Foundry (Job ID: c387dadb)
Records: 4309
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
spai-ss6-corpus-thai-medical-care-aiotx
SPAI SS6 Thai Medical Care AIoTx Index
Index repo for the imported AIoTx Thai medical-care dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_medical_care_aiotx
Rows in canonical config: 3,599
Parquet size in canonical config: 0.00 GB
Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-medical-care-aiotx.
