CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01brandonyang /artem-fold-towelThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 14 ], "names": [ "umi1_x", "umi1_y", "umi1_z", "umi1_rx", "umi1_ry", "umi1_rz", "umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.tabularrobotics1M<n<10M0 likes23k downloads24d agoHugging Face02maxwellinked /time-lapse-artifacts Time-Lapse Artifacts 873 indexed video files document one artist's traditional drawing practice. The recorded finish dates span September 17, 2024 through September 20, 2026; nine Pre-Standard dates remain unknown. Standardized acquisition began July 13, 2025. The current indexes contain 2,196,054,134,482 indexed video bytes (approximately 2.20 TB). The recordings began as personal practice documentation and a durable record of manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.tabular1K<n<10K3 likes13k downloads7h agoHugging Face03brandonyang /artem-pour-waterThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 14 ], "names": [ "umi1_x", "umi1_y", "umi1_z", "umi1_rx", "umi1_ry", "umi1_rz", "umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.tabularrobotics1M<n<10M0 likes9.7k downloads1d agoHugging Face04ArtificialAnalysis /AA-LCR Artificial Analysis Long Context Reasoning (AA-LCR) Dataset AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources. New in Version 1.1 (September 2026) Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.tabularn<1K38 likes8.2k downloads19d agoHugging Face05cy0307 /ropedia-xperience-10m-task-suite-artifacts Ropedia Xperience-10M Task Suite Artifacts This dataset repository stores small derived artifacts for the Ropedia Xperience-10M task-suite project: metrics, predictions, manifests, reports, figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the Cosmos3-Nano plus Cosmos3-Super diagnostic packages. Project Identity The Project identity mark is shared across the GitHub README, GitHub Pages dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.tabular10K<n<100K1 likes5.7k downloads3mo agoHugging Face06pietrolesci /anchoral-paper-artefactsArtefacts related to the paper AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets (Lesci and Vlachos, 2024) published at the NAACL 2024 conference. These artefacts can be reproduced using the code available at github.com/pietrolesci/anchoral. The outputs/ folder includes the raw files created by the individual experiments. The results/ folder contains the exported metrics and configurations that are used to complete the analysis and create the tables and… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/anchoral-paper-artefacts.tabular1M<n<10M0 likes5.5k downloads1y agoHugging Face07hossam3759180 /ids-project-artifactsimage10M<n<100M0 likes4k downloads3d agoHugging Face08arterm-sedov /agent-course-final-assignment Agent Course Final Assignment - Unified Dataset Author: Arte(r)m Sedov GitHub: https://github.com/arterm-sedov/ Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment Dataset Description This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.tabularn<1K1 likes3.2k downloads9mo agoHugging Face09artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K12 likes3k downloads2mo agoHugging Face10artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face11imolodetskikh /sr-artifact-prominence SR Artifact Prominence Annotated super-resolution artifact regions across four image subsets, with crowdsourced per-region prominence scores, artifact type labels, and natural-language descriptions. Prominence is the fraction of valid crowd workers who answered that the highlighted region contains a noticeable super-resolution artifact. Subsets Subset Source dataset Source images Masks Notes open_images Open Images 547 1,523 GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.image1K<n<10K0 likes2.3k downloads5mo agoHugging Face12vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.6k downloads1y agoHugging Face13minnesotanlp /LLM-Artifacts Under the Surface: Tracking the Artifactuality of LLM-Generated Data Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶ Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo, Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu Dongyeop Kang Minnesota NLP, University of Minnesota Twin Cities † Project Lead, ¶ Core Contribution, Arxiv Project Page 📌 Table of Contents Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.tabular100K<n<1M2 likes1.3k downloads3y agoHugging Face14arthurneuron /cryptocurrency-futures-ohlcv-dataset-1mtabular100M<n<1B5 likes1.1k downloads3y agoHugging Face15mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes995 downloads7mo agoHugging Face16csoai /gspc-art5 GSPC — art5 safeguard bank (Art5Bench) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the art5-safeguard row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=art5-safeguard (family, kind, status and n are on that row… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-art5.tabularquestion-answeringn<1K0 likes877 downloads18h agoHugging Face17keryszhan /harbor-swesmith-rl-artifacts Harbor SWE-Smith 强化学习数据产物 本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。 项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。 数据概况 切分 任务数 训练集 187 验证集 42 测试集 38 合计 267 数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。 正式数据集名称: swesmith-curated-grpo-267-v1 冻结切分的语义摘要: ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d 该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.tabulartext-generationn<1K0 likes762 downloads18d agoHugging Face18brandonyang /artem-fold-towel-filtered Artem fold-towel filtered trajectories Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline. The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.tabularrobotics1M<n<10M0 likes704 downloads22d agoHugging Face19danym /rtpurbo-block-summary-failed-experiment-artifacts RTPurbo block-summary failed experiment artifacts Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B. The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.tabularfeature-extraction1K<n<10K0 likes690 downloads2mo agoHugging Face20AI-Art-Collab /640tabular100K<n<1M0 likes607 downloads10mo agoHugging Face21shshwtsuthar /memory-representation-contextbench-artifacts Memory Representation ContextBench Artifacts Dataset Summary This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs. The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.tabular1K<n<10K0 likes578 downloads3mo agoHugging Face22ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes550 downloads3mo agoHugging Face23changdae /tau2-uq-artifacts tau2-bench UQ Artifacts Interaction trajectories and token-level log-probability measurements from conversational customer service agent evaluations on tau2-bench, collected as part of the uncertainty quantification (UQ) pipeline. Used for analyses in the paper "Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities" under the agentuq codebase. Dataset Overview This dataset contains two types of artifacts: Trajectories --… See the full description on the dataset page: https://huggingface.co/datasets/changdae/tau2-uq-artifacts.tabulartext-generation1M<n<10M0 likes526 downloads20d agoHugging Face24nakasyou /note-articles note articles note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。 各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。 tabular100K<n<1M1 likes477 downloads2mo agoHugging Face25artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes462 downloads4mo agoHugging Face26AI-Art-Collab /576tabular1M<n<10M0 likes458 downloads1y agoHugging Face27jtviegas /ticker_analysis_articlestabular10K<n<100K0 likes420 downloads3h agoHugging Face28arthurpompeu /lecrop-data LeCropFollow Latent Space Planning for Navigation in Unstructured Crop Fields Felipe Tommaselli1 · Francisco Affonso2 · Arthur Rocha1 · Gianluca Capezzuto1 Arun Narenthiran Sivakumar2 · Girish Chowdhary2 · Marcelo Becker1 1 University of Sao Paulo &nbsp;&nbsp; 2 University of Illinois Urbana-Champaign IEEE Robotics and Automation Letters, 2026 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/arthurpompeu/lecrop-data.tabularroboticsn<1K0 likes405 downloads3mo agoHugging Face29artist-23 /nifty-options-datatabular10M<n<100M10 likes399 downloads9mo agoHugging Face30alessiotoniolo /ART-Chat-2.5M ART-Chat-2.5M From benchmarking inference engine performance to LLM load-balancing algorithms, ART-Chat-2.5M offers long-context, high prefix-reuse chatbot metadata derived from 2,525,215 production inference requests. Message bodies are synthetically generated and match the original data's prefix-reuse shape. Compared to WildChat-4.8M, ART-Chat-2.5M has 19× higher intra-user prefix reuse and an average token length of 17,964 versus 2,925. We publish this data under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/alessiotoniolo/ART-Chat-2.5M.tabulartext-generation1M<n<10M2 likes372 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.