lime
Datasets
All datasets matching “lime”SII_self_evovling_02_training_datasetLiME_dataScraping-storageLLaVA-ReCap-676KThis is an integrated version of LLaVA-ReCap, sourced from lmms-lab/LLaVA-ReCap-558K and lmms-lab/LLaVA-ReCap-118K.
In this version, the conversations field has been split into two separate fields: prompt and response. Additionally, the <image> special token has been removed to facilitate customization.
Inspired by the original paper, the prompt field has been further expanded with human-crafted variations. Specifically, each prompt is sampled from one of the following 30 instructions:… See the full description on the dataset page: https://huggingface.co/datasets/LimeryJorge/LLaVA-ReCap-676K.LIME-440K_CogAudio-LLM
LIME-440K Dataset (Core Subset: Parts A & B)
This is the dataset repository for the paper: "Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models" (Accepted by INTERSPEECH 2026).
📄 Paper (Arxiv): https://arxiv.org/pdf/2606.06940v1
💻 GitHub Repository: https://github.com/zxzhao0/CogAudio-LLM
This repository currently hosts the LIME-Core subset (Parts A & B), a bilingual (Chinese & English) dataset constructed using… See the full description on the dataset page: https://huggingface.co/datasets/zhaoxiaoxian/LIME-440K_CogAudio-LLM.Synthetic_Unanswerable_Math
Dataset Card for Synthetic Unanswerable Math (SUM)
Dataset Summary
Synthetic Unanswerable Math (SUM) is a dataset of high-quality, implicitly unanswerable math problems constructed to probe and improve the refusal behavior of large language models (LLMs). The goal is to teach models to identify when a problem cannot be answered due to incomplete, ambiguous, or contradictory information, and respond with epistemic humility (e.g., \boxed{I don't know}).
Each entry in the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/Synthetic_Unanswerable_Math.
