CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes4.7k downloads4mo agoHugging Face02yxu /LiME_datatext10K<n<100K0 likes4.6k downloads3y agoHugging Face03LimeryJorge /LLaVA-ReCap-676KThis is an integrated version of LLaVA-ReCap, sourced from lmms-lab/LLaVA-ReCap-558K and lmms-lab/LLaVA-ReCap-118K. In this version, the conversations field has been split into two separate fields: prompt and response. Additionally, the <image> special token has been removed to facilitate customization. Inspired by the original paper, the prompt field has been further expanded with human-crafted variations. Specifically, each prompt is sampled from one of the following 30 instructions:… See the full description on the dataset page: https://huggingface.co/datasets/LimeryJorge/LLaVA-ReCap-676K.imagequestion-answering100K<n<1M0 likes399 downloads1y agoHugging Face04lime-nlp /Synthetic_Unanswerable_Math Dataset Card for Synthetic Unanswerable Math (SUM) Dataset Summary Synthetic Unanswerable Math (SUM) is a dataset of high-quality, implicitly unanswerable math problems constructed to probe and improve the refusal behavior of large language models (LLMs). The goal is to teach models to identify when a problem cannot be answered due to incomplete, ambiguous, or contradictory information, and respond with epistemic humility (e.g., \boxed{I don't know}). Each entry in the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/Synthetic_Unanswerable_Math.textreinforcement-learning10K<n<100K18 likes271 downloads1y agoHugging Face05Limerencii /russian-handwriting-ocr Russian Handwritten Text Recognition Dataset Датасет для распознавания русских рукописных текстов (сочинений). Описание Этот датасет содержит изображения рукописных русских текстов с их расшифровкой. Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста. Статистика Всего образцов: 13050 Train: 11745 Validation: 1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/Limerencii/russian-handwriting-ocr.imageimage-to-text10K<n<100K11 likes206 downloads11mo agoHugging Face06lime-nlp /DeepScaleR_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.tabularreinforcement-learning1M<n<10M11 likes137 downloads1y agoHugging Face07LIME-DATA /infovqaimage1K<n<10K2 likes113 downloads2y agoHugging Face08suwaimyo /limesoda-tha-classification LimeSoda_tha_Classification Deduplicated copy of kornwtp/limesoda-tha-classification. Splits split rows test 2,723 train 2,648 validation 299 text1K<n<10K0 likes98 downloads28d agoHugging Face09lime-nlp /GSM8K_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.tabular1M<n<10M1 likes84 downloads1y agoHugging Face10lime-nlp /safer-instruct Safer-Instruct: Aligning Language Models with Automated Preference Data This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here! Abstract Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.text10K<n<100K1 likes73 downloads2y agoHugging Face11kornwtp /limesoda-tha-classificationtext1K<n<10K0 likes63 downloads2mo agoHugging Face12limeXx /openhermes-reasoning-231k 🧠 OpenHermes Reasoning 231K High-quality instruction dataset with chain-of-thought reasoning 🤗 Dataset • 💬 Discussions 📊 Dataset Overview This dataset contains 231,144 high-quality instruction-response pairs with explicit chain-of-thought reasoning. Each example includes: Prompt: Original instruction or question Thinking: Explicit reasoning process and logical steps Answer: Final comprehensive response Key Features ✅ Quality Filtered: Rigorous… See the full description on the dataset page: https://huggingface.co/datasets/limeXx/openhermes-reasoning-231k.texttext-generation100K<n<1M1 likes57 downloads1y agoHugging Face13kulia-moon /LimeStory-1.0NEW IN THIS DATASETLimeStory Dataset Version 1.0 is now available in 🤗 Spaces and users can add stories about anything! (Powered by Pollinations.ai) NOTICELimeStory is not for training NSFW models, and remember to use dataset: kulia-moon/LimeStory-1.0 for you're using this dataset as target training! Kulia's datasets The story, your impossible Generated by 🤗 Spaces Protected by 🤗 Scanner .hf-sanitized.hf-sanitized-yvTjnu7IKxZ1msJf6A22y .cursive { font-family: "Lobster"… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/LimeStory-1.0.texttext-generation1K<n<10K0 likes48 downloads11mo agoHugging Face14LIME-DATA /ChartQAimage1K<n<10K0 likes43 downloads2y agoHugging Face15lime-nlp /orz_math_difficulty Difficulty Estimation on Open Reasoner Zero We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction. Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.tabular1M<n<10M0 likes41 downloads1y agoHugging Face16limen231 /ag_newsAG's News Topic Classification Dataset Version 3, Updated 09/09/2015 ORIGIN AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search… See the full description on the dataset page: https://huggingface.co/datasets/limen231/ag_news.texttext-classification100K<n<1M0 likes41 downloads9d agoHugging Face17lime-nlp /MATH_Difficulty Difficulty Estimation on MATH We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.tabular1M<n<10M0 likes39 downloads1y agoHugging Face18LIME-DATA /ai2dimage1K<n<10K0 likes35 downloads2y agoHugging Face19LIME-DATA /OK-VQAimage1K<n<10K0 likes34 downloads2y agoHugging Face20LIME-DATA /TextCapsimage1K<n<10K0 likes31 downloads2y agoHugging Face21puttatidam /limesoda-tha-classificationtext1K<n<10K0 likes31 downloads10d agoHugging Face22agentlans /lime-nlp-difficulty lime-nlp Difficulty Estimation Math Datasets collection Unofficial reformatted version of lime-nlp/difficulty-estimation-math-datasets, which contains math problems and the Qwen 2.5 7B MATH model's success rates at solving those problems. The combined dataset has been split into 80% training and 20% testing data. Fields: row_id: the row number of each dataset entry, starting at 0 input: the math question from the dataset output: the correct answer (ground truth)… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lime-nlp-difficulty.tabulartext-classification100K<n<1M0 likes29 downloads10mo agoHugging Face23BUDDI-AI /Speeding-up-LIMEgated Deidentified Medical Charts with Human Curated Explanations About This dataset is a small sample from the EHR dataset used by experiments described in our paper, "Speeding up LIME with Attention Weights," submitted to CoDS-COMAD 2024. text10K<n<100K2 likes28 downloads3y agoHugging Face24LIME-DATA /COCO-Caption2017image1K<n<10K0 likes28 downloads2y agoHugging Face25LIME-DATA /textvqaimage1K<n<10K0 likes24 downloads2y agoHugging Face26naksyu /Lime-v72Language: Korean Size: 821 examples (약 2MB, JSONL) Format: OpenAI Chat Format (messages list) Target Task: Persona Alignment, Multi-turn Context Tracking, Reasoning, State Overwriting 분리된 사고 과정 (Internal Reasoning vs Visible Output) 모델이 즉각적으로 답을 뱉지 않고, 내부적으로 판단을 거친 뒤 정제된 답변만 출력하도록 학습시킵니다. /think ~ </think> : 사용자의 의도를 파악하고 현재 상태를 점검하는 내부 사고 영역. <|lime_think_more|> ~ <|lime_end_think|> : 대화가 길어지거나 복잡한 요청일 경우, 스스로 제약 조건(예: 명사구 누락 금지)을 한 번 더 점검하는 추가 사고 영역. <|lime_final|> : 최종적으로 사용자에게 노출되는 정제된… See the full description on the dataset page: https://huggingface.co/datasets/naksyu/Lime-v72.textn<1K0 likes24 downloads6mo agoHugging Face27LIME-DATA /ocrbenchimagen<1K1 likes23 downloads2y agoHugging Face28Chromik /lime-explanations-finaltext1K<n<10K0 likes20 downloads1y agoHugging Face29Chromik /lime-explanations-final-3.0text1K<n<10K0 likes20 downloads1y agoHugging Face30LIME-DATA /POPEimagen<1K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.