CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenGVLab /MVBench MVBench Important Update [18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference. We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.imagevisual-question-answering1K<n<10K47 likes18k downloads2y agoHugging Face02openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M169 likes13k downloads2mo agoHugging Face03OpenGVLab /ShareGPT-4ogatedtabularvisual-question-answering10K<n<100K199 likes8.5k downloads2y agoHugging Face04OpenDriveLab /SparseVideoNav SparseVideoNav Datasets This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav: BVN: Beyond-the-View Navigation. IFN: Instruction-Following Navigation. Project links: Project page: https://opendrivelab.com/SparseVideoNav GitHub: https://github.com/OpenDriveLab/SparseVideoNav Paper: https://arxiv.org/abs/2602.05827 Dataset Summary SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.tabularrobotics10K<n<100K4 likes6.1k downloads27d agoHugging Face05opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes5.8k downloads1y agoHugging Face06OpenGVLab /GUI-Odyssey Dataset Card for GUI Odyssey News⭐️ A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉 👉 Please use the latest version and refer to the updated README for the most up-to-date information. We highly recommend using the new version for all training and evaluation! Repository: https://github.com/OpenGVLab/GUI-Odyssey Latest Version of Dataset: hflqf88888/GUIOdyssey Paper: https://arxiv.org/pdf/2406.08451 Introduction GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.image1K<n<10K26 likes4.1k downloads1y agoHugging Face07mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K38 likes2.8k downloads3y agoHugging Face08opencompass /SWEBench-Pro-Verified SWE-Bench Pro Verified: Anti-hacking & Task refinement SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.tabularn<1K2 likes1.5k downloads13d agoHugging Face09AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.2k downloads2y agoHugging Face10OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face11lbox /lbox_open Dataset Card for lbox_open Dataset Summary A Legal AI Benchmark Dataset from Korean Legal Cases. Languages Korean How to use from datasets import load_dataset # casename classficiation task data_cn = load_dataset("lbox/lbox_open", "casename_classification") data_cn_plus = load_dataset("lbox/lbox_open", "casename_classification_plus") # statutes classification task data_st = load_dataset("lbox/lbox_open", "statute_classification") data_st_plus =… See the full description on the dataset page: https://huggingface.co/datasets/lbox/lbox_open.tabular100K<n<1M17 likes1.1k downloads1y agoHugging Face12OpenMOSS-Team /FutureOmni FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Predicting the future requires listening as well as seeing. 📖 Dataset Summary Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.tabularquestion-answering1K<n<10K7 likes989 downloads8mo agoHugging Face13OpenSciLM /OpenScholar-DataStore-V3tabular100M<n<1B23 likes948 downloads2y agoHugging Face14mjbommar /opengloss-v1.3-definitions See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Dictionary v1.3 (Definition-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-definitions.tabulartext-generation100K<n<1M0 likes890 downloads16d agoHugging Face15MDGA-3 /openmoe2_ckptstabularn<1K1 likes883 downloads11mo agoHugging Face16junbrro /egopi_latal_openarm_bottletabularn<1K0 likes783 downloads2mo agoHugging Face17mjbommar /opengloss-v1.3-dictionary See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M1 likes767 downloads16d agoHugging Face18junbrro /egopi_latal_openarm_cuptabularn<1K0 likes576 downloads2mo agoHugging Face19junbrro /egopi_latal_openarm_snacktabularn<1K0 likes562 downloads2mo agoHugging Face20PolinAvA /openarm_statictabularn<1K0 likes525 downloads5mo agoHugging Face21amer224 /Opencode1tabularn<1K6 likes485 downloads9d agoHugging Face22junbrro /egopi_latal_openarm_dolltabularn<1K0 likes432 downloads2mo agoHugging Face23walledai /openai-moderation-dataset Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual excitement, such as the… See the full description on the dataset page: https://huggingface.co/datasets/walledai/openai-moderation-dataset.tabular1K<n<10K2 likes403 downloads1y agoHugging Face24Norquinal /OpenCAIThis dataset is comprised of roleplay chat conversations scraped from several Discord RP fandom servers. The conversations have been split in terms of days, the assumption being that a majority of long-form roleplays are started/continued and completed within a day. The original dataset consists of ~14K samples. Light filtering striped that down to ~10K samples. Stricter filtering striped it down to ~5k samples. Strictest filtering striped it down to ~4k samples. Effort was taken to remove… See the full description on the dataset page: https://huggingface.co/datasets/Norquinal/OpenCAI.tabular10K<n<100K19 likes401 downloads2y agoHugging Face25OpenClaw /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tabulartext-classification10K<n<100K53 likes400 downloads3mo agoHugging Face26simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes380 downloads15h agoHugging Face27opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes356 downloads1y agoHugging Face28choucsan /OpenUAV-QA ✨OpenUAV-QA✨ OpenUAV-QA is a large-scale multiple-choice question-answering benchmark for UAV (drone) navigation decision-making, built upon the TravelUAV (OpenUAV) dataset. It transforms raw UAV flight trajectories into structured, text-polished 4-option QA pairs that test a multimodal model's ability to reason about path planning, action sequences, and spatial dynamics from first-person drone video footage. The dataset covers 22 distinct simulated… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/OpenUAV-QA.imagevisual-question-answering10K<n<100K3 likes343 downloads2mo agoHugging Face29SkillCorner /opendata-bodypose SkillCorner Open Data — Body Pose 3D body-pose data derived from broadcast video, released alongside the SkillCorner Open Data repository as a joint initiative between SkillCorner and PySport. Initial testing release. Two matches, published so the community can work with the format and tell us what is useful before we consider a wider release. Feedback is genuinely wanted — open an issue on the opendata repo or reply in the Community tab here. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/SkillCorner/opendata-bodypose.tabular100K<n<1M0 likes341 downloads14d agoHugging Face30tegridydev /opensec-triage opensec-triage 0.5.0 Synthetic English security alert data for disposition classification, counterfactual evaluation and small model training experiments. Each example pairs a security observation with contextual evidence and an expected disposition. The task is to classify the supplied evidence rather than infer a disposition from the observable action alone. Configurations Configuration Purpose Splits default Main 50,000 row text classification… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/opensec-triage.tabulartext-classification100K<n<1M4 likes318 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.