datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaST-Bench
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
This is the official repository for the CaST-Bench dataset, introduced in the paper
"CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering".
CaST-Bench is the first benchmark to evaluate Vision-Language Models (VLMs) on causal chain
reasoning grounded in fine-grained spatio-temporal evidence. Given a video and a causal question… See the full description on the dataset page: https://huggingface.co/datasets/wovenbytoyota-vai/CaST-Bench.InstVL
InstVL: A Large-Scale Instance-Aware Vision-Language Dataset
This is the official repository for the InstVL dataset, introduced in the paper InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding.
InstVL is a large-scale dataset of images and videos designed to bridge the gap between holistic scene understanding and fine-grained, instance-level comprehension. Current vision-language pre-training (VLP) paradigms excel at global scene understanding but… See the full description on the dataset page: https://huggingface.co/datasets/wovenbytoyota-vai/InstVL.UnLOK-VQA
📊 Dataset: UnLOK-VQA (Unlearning Outside Knowledge VQA)
Paper: Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Code: https://github.com/Vaidehi99/mmmedit
Link: Dataset Link
This dataset contains approximately 500 entries with the following key attributes:
"id": Unique Identifier for each entry
"src": The question whose answer is to be deleted ❓
"pred": The answer to the question meant for deletion ❌
"loc": Related neighborhood questions… See the full description on the dataset page: https://huggingface.co/datasets/vaidehi99/UnLOK-VQA.hdfs-logsbgl-logsEC-BenchNOVA50knepi-prompts-dataset
NEPI: Narrative-Embedded Prompt Injection Dataset (Sanitized)
Dataset Summary
This dataset contains 4,000 sanitized prompts designed for research on prompt injection vulnerabilities in Large Language Models (LLMs).It introduces and supports evaluation of a novel attack class called Narrative-Embedded Prompt Injection (NEPI), where adversarial intent is embedded inside coherent fictional narratives, dialogues, or persona-driven roleplay prompts.
Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav-GOAT/nepi-prompts-dataset.Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning
(The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text)
This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.adaption-vaidya-rural-symptoms
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms.JARVISadaption-vaidya-rural-symptoms-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms-v1.unladaption-chest-xray-pneumonia-labels
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-chest_xray_pneumonia_labels
This dataset contains labeled samples for chest X-ray image classification, distinguishing between normal cases and those with pneumonia. Each entry provides a diagnostic label indicating the presence of pneumonia or a normal lung condition. The data is structured as pairs with a single completion field holding the categorical diagnosis.… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavipadmanabhan/adaption-chest-xray-pneumonia-labels.ToolCall-SFT-v2-40KEMOTIONAL-JARVISHC_concepts_refusalunsloth-sharagahindicorrectioncombined_refusal_datacore
CORE: Comprehensive Ontological Relation Evaluation
🌐 Website |
📄 Paper |
💻 Code
Dataset Summary
CORE is a human-grounded benchmark for evaluating large language models on fundamental semantic and ontological reasoning. It assesses whether models can correctly recognize a broad range of sense-level relations and, critically, identify when no meaningful relationship exists between concepts. With comprehensive relation coverage and strong human baselines… See the full description on the dataset page: https://huggingface.co/datasets/vaikhari-ai/core.LC_concepts_refusaldataset7charak-chapter2gitsolve_alpaca_datasetMC_concepts_refusalmedical.jsonalpaca-cleaned-10k-chattestllamastory-dataset
