CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face02argilla /distilabel-intel-orca-dpo-pairs distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.text10K<n<100K182 likes22k downloads1y agoHugging Face03Intelligent-Systems /BEDLAM-depthgated Dataset Mirror of BEDLAM Dataset (Depth Data Subset) Project site: https://bedlam.is.tuebingen.mpg.de/ Please register at project site for additional information and data (Download section) Related Hugging Face dataset mirror: BEDLAM Dataset Information Depth maps (EXR, 32-bit, 3.8TB) Camera ground truth information is not included but can be found in the BEDLAM dataset mirror Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.text1M<n<10M0 likes19k downloads7mo agoHugging Face04IntelligenceLab /Long-Horizon-Terminal-Bench Long-Horizon Terminal-Bench (LHTB) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent into a stateful environment and grades it with hidden, rebuild-from-artifact verifiers — self-reported progress does not count. 📝 Blog: https://zli12321.github.io/LHTB/ 🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.documenttext-generationn<1K136 likes12k downloads5d agoHugging Face05InteliLab /ict_s2s_refactoredaudio100K<n<1M0 likes7.5k downloads4mo agoHugging Face06INS-IntelligentNetworkSolutions /Waste-Dumpsites-DroneImagery Dataset for Waste/Dumpsite Detection using drone imagery Contains 2115 drone images of illegal waste dumpsites 1280 x 1280 px resolution Nadir perspective (camera pointing straight down at a 90-degree angle to the ground) Annotations and Images train | valid | test actual images COCO - annotations_coco.json files in each split directory .parquet files in data directory with embeded images The dataset was collected as part of the [ Raven Scan ] project, more… See the full description on the dataset page: https://huggingface.co/datasets/INS-IntelligentNetworkSolutions/Waste-Dumpsites-DroneImagery.imageobject-detection10K<n<100K8 likes7k downloads2y agoHugging Face07colemei /IntelliSA-dataset IntelliSA Dataset Infrastructure as Code security vulnerability dataset with ground truth labels and pseudo-labeled training data across Chef, Ansible, and Puppet. Dataset Overview Component Size Purpose Oracle 241 scripts, 213 smells Ground truth evaluation set Training 2,300 instances + 6,070 raw scripts Model training data Oracle Dataset (Ground Truth) Ansible: 81 scripts, 44 smells Chef: 80 scripts, 104 smells Puppet: 80 scripts, 65… See the full description on the dataset page: https://huggingface.co/datasets/colemei/IntelliSA-dataset.textn<1K0 likes6.5k downloads10mo agoHugging Face08IntelligenceLab /LHTB-leaderboard LHTB Leaderboard — Long-Horizon Terminal-Bench This repository hosts submitted runs for Long-Horizon Terminal-Bench (LHTB), a 46-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Every entry below ships its complete run artifacts — per-trial configs, results, verifier outputs and terminal recordings — so any score on this board can be audited without rerunning the suite. 📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.tabular1K<n<10K3 likes5k downloads7h agoHugging Face09Intelligent-Systems /BEDLAM2-depthgated Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset) Project site: https://bedlam2.is.tuebingen.mpg.de/ Please register at project site for additional information and data in its Download section. Related Hugging Face dataset mirror: BEDLAM2 Dataset Information Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB) Multilayer EXR details 16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R) Color image without motion blur Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.text1M<n<10M0 likes3.3k downloads7mo agoHugging Face10PrimeIntellect /INTELLECT-3-RLtabular10K<n<100K8 likes3.1k downloads4mo agoHugging Face11Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes2.7k downloads1y agoHugging Face12Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.5k downloads1h agoHugging Face13Intelligent-Internet /pd12m PD12M This is a curated PD12M dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.imagefeature-extraction10M<n<100M8 likes2.3k downloads1y agoHugging Face14Intel /orca_dpo_pairsThe dataset contains 12k examples from Orca style dataset Open-Orca/OpenOrca. text10K<n<100K324 likes2.3k downloads3y agoHugging Face15geometric-intelligence /ogbench OgBench: Benchmarking Graph Neural Networks on Omics Data OgBench is the first benchmark suite for graph-level prediction in the n ≪ p regime characteristic of omics data, where the number of patient samples n is much smaller than the number of nodes (genes or proteins) p per graph. Datasets This repository contains four preprocessed omics graph classification datasets: Dataset Modality n p Task HERITAGE Proteomics 654 4,977 Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.texttabular-classification100K<n<1M0 likes1.7k downloads10d agoHugging Face16vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.5k downloads1y agoHugging Face17PrimeIntellect /INTELLECT-3-SFTtext1M<n<10M8 likes1.5k downloads10mo agoHugging Face18Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.3k downloads1y agoHugging Face19Delta-Intelligence /OmniEgo D1 Headset Egocentric Whole-body Dataset D1 is a headset multi-camera human motion dataset for humanoid intelligence, embodied AI, whole-body motion understanding, and imitation learning. Overview The D1 dataset is exported from the D1 headset multi-camera human motion capture system developed by Delta Intelligence. Each recorded episode contains synchronized multi-view video streams and whole-body skeleton and headset pose data. The dataset supports research… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Intelligence/OmniEgo.imagerobotics100K<n<1M0 likes1.1k downloads3d agoHugging Face20Intel /SK-VQA Dataset Card for SQ-VQA Dataset Summary SK-VQA is a large-scale synthetic multimodal dataset containing over 2 million visual question-answer pairs, each paired with context documents that contain the information needed to answer the questions. The dataset is designed to address the critical need for training and evaluating multimodal LLMs (MLLMs) in context-augmented generation settings, particularly for retrieval-augmented generation (RAG) systems. It enables training… See the full description on the dataset page: https://huggingface.co/datasets/Intel/SK-VQA.image1M<n<10M1 likes1k downloads1y agoHugging Face21IntelligenceResearchLab /Hausa Hausa Ajami OCR Dataset Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa). Contenu Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec : file_name : nom du fichier image correspondant (image de la ligne, recadrée) transcript : translittération en écriture latine de la ligne source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.imageimage-to-text1K<n<10K3 likes991 downloads4d agoHugging Face22Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes927 downloads1y agoHugging Face23mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes915 downloads7mo agoHugging Face24Intelligent-Internet /GAIA-Subset-Benchmark GAIA Benchmark Subset Model Card This dataset is a subset of the GAIA benchmark, containing 44 web-search-based questions from the validation set. It evaluates multiple AI models on their ability to retrieve and process real-time information using web search and browser tools. Performance metrics include success indicators and detailed reports for each model. A comparative chart summarizing the results will be provided separately. Benchmark Results textn<1K3 likes882 downloads1y agoHugging Face25IntelLabs /FloorSet Dataset Card for FloorSet Dataset Summary FloorSet is a dataset that contains a large training data that reflect real-world constraints and objectives of the Floorplanning problem in chip design flow, which is a crucial component of the structural design flow. This dataset contains synthetic fixed-outline floorplan layouts in a pkl format, that reflect the distribution of real SoCs and sub-system layouts. The dataset has 1M training samples and 100 test samples, with hard… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/FloorSet.text10K<n<100K2 likes867 downloads2y agoHugging Face26Intel /SocialCounterfactualsimage100K<n<1M11 likes849 downloads2y agoHugging Face27AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes746 downloads24d agoHugging Face28human-intelligence-ai /GDPval-CN-Seed-Set GDPval-CN Seed Set 中文详细说明 · English documentation · 样本说明 GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。 我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。 这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。 GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review. 数据概览 项目 内容 任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.documentothern<1K1 likes636 downloads1mo agoHugging Face29intellekthq /enron-ferc-pst Enron FERC email corpus in native PST The EDRM Enron v2 email corpus in Microsoft PST format, modified to reduce personal privacy risk. Mailbox structure, MAPI metadata, message bodies, and retained attachments are preserved. The release contains 171 PST files in data/, with one or more files per custodian. Count Version v1 Messages 1,226,178 Attachments 453,832 PST files 171 Possible uses include email research, e-discovery testing, information retrieval… See the full description on the dataset page: https://huggingface.co/datasets/intellekthq/enron-ferc-pst.text1M<n<10M1 likes563 downloads2mo agoHugging Face30AIM-Intelligence /XL-SafetyBench XL-SafetyBench A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity ⚠️ Content Warning: This dataset contains adversarial prompts and culturally sensitive content for safety and cultural-evaluation research. By using this dataset, you agree to use it solely for research purposes and not for malicious applications. Paper: https://arxiv.org/abs/2605.05662 Eval Code: github.com/AIM-Intelligence/XL-SafetyBench Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.texttext-classification1K<n<10K8 likes542 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.