datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.time-lapse-artifacts
Time-Lapse Artifacts
873 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through September 20, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,196,054,134,482 indexed video bytes
(approximately 2.20 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.artem-pour-waterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.AA-LCR
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset
AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources.
New in Version 1.1 (September 2026)
Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.anchoral-paper-artefactsArtefacts related to the paper AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets (Lesci and Vlachos, 2024) published at the NAACL 2024 conference.
These artefacts can be reproduced using the code available at github.com/pietrolesci/anchoral.
The outputs/ folder includes the raw files created by the individual experiments.
The results/ folder contains the exported metrics and configurations that are used to complete the analysis and create the tables and… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/anchoral-paper-artefacts.ids-project-artifactsagent-course-final-assignment
Agent Course Final Assignment - Unified Dataset
Author: Arte(r)m Sedov
GitHub: https://github.com/arterm-sedov/
Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment
Dataset Description
This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.ledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.ledger-long-context-KPI-QA
LEDGER — Long-Context KPI Question Answering & Page Retrieval
This dataset is part of the LEDGER (Long-context Evaluation of Documents for
Grounded Extraction and Retrieval) benchmark.
It supports two of the three LEDGER tasks:
Page-level KPI retrieval — given a natural-language question about a financial
KPI and the corresponding annual report, retrieve the relevant page(s). Each row
includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.cryptocurrency-futures-ohlcv-dataset-1msyntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
gspc-art5
GSPC — art5 safeguard bank (Art5Bench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the art5-safeguard row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=art5-safeguard (family, kind, status and n are on that row… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-art5.harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.artem-fold-towel-filtered
Artem fold-towel filtered trajectories
Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.rtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.640memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.tau2-uq-artifacts
tau2-bench UQ Artifacts
Interaction trajectories and token-level log-probability measurements from conversational customer service agent evaluations on tau2-bench, collected as part of the uncertainty quantification (UQ) pipeline. Used for analyses in the paper "Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities" under the agentuq codebase.
Dataset Overview
This dataset contains two types of artifacts:
Trajectories --… See the full description on the dataset page: https://huggingface.co/datasets/changdae/tau2-uq-artifacts.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
precinct6-cybersecurity-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.576ticker_analysis_articleslecrop-data
LeCropFollow
Latent Space Planning for Navigation in Unstructured Crop Fields
Felipe Tommaselli1 ·
Francisco Affonso2 ·
Arthur Rocha1 ·
Gianluca Capezzuto1
Arun Narenthiran Sivakumar2 ·
Girish Chowdhary2 ·
Marcelo Becker1
1 University of Sao Paulo
2 University of Illinois Urbana-Champaign
IEEE Robotics and Automation Letters, 2026
… See the full description on the dataset page: https://huggingface.co/datasets/arthurpompeu/lecrop-data.nifty-options-dataART-Chat-2.5M
ART-Chat-2.5M
From benchmarking inference engine performance to LLM load-balancing algorithms, ART-Chat-2.5M offers long-context, high prefix-reuse chatbot metadata derived from 2,525,215 production inference requests.
Message bodies are synthetically generated and match the original data's prefix-reuse shape. Compared to WildChat-4.8M, ART-Chat-2.5M has 19× higher intra-user prefix reuse and an average token length of 17,964 versus 2,925. We publish this data under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/alessiotoniolo/ART-Chat-2.5M.
