CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DAMO-NLP-SG /MultiJail Multilingual Jailbreak Challenges in Large Language Models This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models". [Github repo] Annotation Statistics We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below: High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi) Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.textn<1K12 likes1k downloads3y agoHugging Face02Alibaba-DAMO-Academy /RynnBrain-Bench RynnBrain-Bench Introduction We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments. Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences. RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.textvisual-question-answering10K<n<100K14 likes953 downloads7mo agoHugging Face03DAMO-NLP-MT /multialpacatext100K<n<1M12 likes654 downloads3y agoHugging Face04damo-da /oag-nepal-audit-reports OAG Nepal Audit Reports — Nepali transcripts and ruled tables Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements. The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.documenttext-retrieval10M<n<100M0 likes579 downloads23d agoHugging Face05DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face06damonsalvatore123 /LMA-Individual-Projectimagen<1K0 likes309 downloads9d agoHugging Face07DamonDemon /SpanUQ-Benchmark SpanUQ Benchmark A span-level uncertainty estimation benchmark for large language model generation. Each example contains an LLM-generated response decomposed into spans (contiguous text segments expressing single verifiable assertions), with uncertainty labels derived from sampling-based consistency verification. Quick Start from datasets import load_dataset # Load a specific model configuration ds = load_dataset("DamonDemon/SpanUQ-Benchmark", "Qwen3-14B")… See the full description on the dataset page: https://huggingface.co/datasets/DamonDemon/SpanUQ-Benchmark.tabulartext-generation10K<n<100K0 likes173 downloads3mo agoHugging Face08Alibaba-DAMO-Academy /ClinHallu CLINHALLU Benchmark CLINHALLU is a benchmark for diagnosing stage-wise hallucinations in medical MLLM reasoning. Paper: CLINHALLU: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM ReasoningGitHub: alibaba-damo-academy/ClinHallu Benchmark Results Accuracy and stage-wise hallucination rates on CLINHALLU. We report answer accuracy (Acc) and hallucination rates for visual recognition (H^V), knowledge recall (H^K), and reasoning integration (H^R).… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinHallu.text10K<n<100K3 likes144 downloads3mo agoHugging Face09ToxicityPrompts /DAMO-MultiJailThe dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. text1K<n<10K0 likes105 downloads1y agoHugging Face10damo-da /ciaa-annual-reports CIAA Annual Reports — Nepali transcripts, ruled tables and chart data Machine-readable transcripts of the annual reports of Nepal's Commission for the Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) — all 35 it has published to date. The 1st to 35th reports, fiscal years BS 2047/48 – 2081/82 (AD 1990–2025). The CIAA publishes these as PDFs whose text layer is, for several years, legacy pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.imagetext-retrieval100K<n<1M0 likes78 downloads1mo agoHugging Face11damonsalvatore123 /NLP-A2text0 likes75 downloads4d agoHugging Face12DAMO-NLP-SG /VL3-Syn7M The re-caption dataset used in VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding If you like our project, please give us a star ⭐ on Github for the latest update. 🌟 Introduction This dataset is the re-captioned data we used during the training of VideoLLaMA3. It consists of 7 million diverse, high-quality images, each accompanied by a short caption and a detailed caption. The images in this dataset originate from COYO-700M, MS-COCO 2017… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VL3-Syn7M.imagevisual-question-answering1M<n<10M11 likes69 downloads2y agoHugging Face13DAMO-NLP-SG /SOULThis repo contains the data for our paper "SOUL: Towards Sentiment and Opinion Understanding of Language" in EMNLP 2023. Github repo Statistics The SOUL dataset comprises 15,028 statements related to 3,638 reviews, resulting in an average of 4.13 statements per review. To create training, development, and test sets, we split the reviews in a ratio of 6:1:3, respectively. Split # reviews # statements True False Not-given Total Train 2,182 3,675 2,159 8,834 3,000 8,834… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/SOUL.texttext-classification10K<n<100K0 likes57 downloads3y agoHugging Face14Damoniano /Agent1imagen<1K0 likes15 downloads4y agoHugging Face15chansurgeplus /oasst1-guanaco-damo-convai-pro Dataset Card for "oasst1-guanaco-damo-convai-pro" More Information needed text10K<n<100K0 likes8 downloads3y agoHugging Face16damojay /matterjstextn<1K0 likes8 downloads2y agoHugging Face17damog369 /legal-retrieval-decisiontextn<1K0 likes7 downloads11mo agoHugging Face18damonsalvatore123 /Deidentification-of-EHRtextn<1K0 likes5 downloads1mo agoHugging Face19damon6 /de_shop_api_v3text1K<n<10K0 likes4 downloads1y agoHugging Face20Damon07 /softprompt0ingtextn<1K0 likes4 downloads1y agoHugging Face21damonsurrao /autotrain-datasettextn<1K0 likes3 downloads2y agoHugging Face22damon6 /v4_100k_processedtabular100K<n<1M0 likes2 downloads1y agoHugging Face23damon98 /parler_spark_traingatedaudio10K<n<100K0 likes2 downloads1y agoHugging Face24fares-boutriga /DamorkDataSettextn<1K0 likes2 downloads5mo agoHugging Face25DamolaRachael /data.csv Nanbeige4-3B Base Model Blind Spot Dataset Model Tested Nanbeige/Nanbeige4-3B-Base https://huggingface.co/Nanbeige/Nanbeige4-3B-Base This dataset documents examples where the model produces incorrect predictions. Dataset Structure column description input prompt given to the model expected_output correct output model_output output generated by the model error_type category of error How the Model Was Loaded from transformers… See the full description on the dataset page: https://huggingface.co/datasets/DamolaRachael/data.csv.textn<1K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.