datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.artem-pour-waterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.anchoral-paper-artefactsArtefacts related to the paper AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets (Lesci and Vlachos, 2024) published at the NAACL 2024 conference.
These artefacts can be reproduced using the code available at github.com/pietrolesci/anchoral.
The outputs/ folder includes the raw files created by the individual experiments.
The results/ folder contains the exported metrics and configurations that are used to complete the analysis and create the tables and… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/anchoral-paper-artefacts.ids-project-artifactsledger-long-context-KPI-QA
LEDGER — Long-Context KPI Question Answering & Page Retrieval
This dataset is part of the LEDGER (Long-context Evaluation of Documents for
Grounded Extraction and Retrieval) benchmark.
It supports two of the three LEDGER tasks:
Page-level KPI retrieval — given a natural-language question about a financial
KPI and the corresponding annual report, retrieve the relevant page(s). Each row
includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.ledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
artem-fold-towel-filtered
Artem fold-towel filtered trajectories
Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.rtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
ticker_analysis_articleslecrop-data
LeCropFollow
Latent Space Planning for Navigation in Unstructured Crop Fields
Felipe Tommaselli1 ·
Francisco Affonso2 ·
Arthur Rocha1 ·
Gianluca Capezzuto1
Arun Narenthiran Sivakumar2 ·
Girish Chowdhary2 ·
Marcelo Becker1
1 University of Sao Paulo
2 University of Illinois Urbana-Champaign
IEEE Robotics and Automation Letters, 2026
… See the full description on the dataset page: https://huggingface.co/datasets/arthurpompeu/lecrop-data.nifty-options-dataART-Chat-2.5M
ART-Chat-2.5M
From benchmarking inference engine performance to LLM load-balancing algorithms, ART-Chat-2.5M offers long-context, high prefix-reuse chatbot metadata derived from 2,525,215 production inference requests.
Message bodies are synthetically generated and match the original data's prefix-reuse shape. Compared to WildChat-4.8M, ART-Chat-2.5M has 19× higher intra-user prefix reuse and an average token length of 17,964 versus 2,925. We publish this data under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/alessiotoniolo/ART-Chat-2.5M.articulated-boxes
projectsim Articulated Boxes
Try it before you download. Open the Boxes tab in projectsim Lab,
fold every flap, then copy a one-line download command for Isaac Sim, MuJoCo, usdview or Blender.
projectsim by Kaedim — sim-ready 3D assets for robot learning. Open in your engine in one paste.
Ten articulated cardboard shipping boxes (RSC four-flap cartons). Every box is
a reduced-coordinate articulation — four flap hinges with real limits —
with per-link mass and… See the full description on the dataset page: https://huggingface.co/datasets/projectsim/articulated-boxes.pubmed-2019-pythia-word-tfidf-pubmedqa-clean-articlespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-articlesarticulated-desks
projectsim Articulated Desks
Try it before you download. Open the Desks tab in projectsim Lab,
slide every drawer, then copy a one-line download command for Isaac Sim, MuJoCo, usdview or Blender.
projectsim by Kaedim — sim-ready 3D assets for robot learning. Open in your engine in one paste.
Office desks with real sliding drawers — prismatic joints with limits in meters. Every asset is a reduced-coordinate articulation with
per-link mass and inertia from exact mesh… See the full description on the dataset page: https://huggingface.co/datasets/projectsim/articulated-desks.articulated-wardrobes
projectsim Wardrobes
Try it before you download. Open the Wardrobes tab in projectsim Lab,
swing the doors, then copy a one-line download command for Isaac Sim, MuJoCo, usdview or Blender.
projectsim by Kaedim — sim-ready 3D assets for robot learning. Open in your engine in one paste.
Wardrobes with hinged doors — beveled edges, real hinge limits. Every asset is a reduced-coordinate articulation with
per-link mass and inertia from exact mesh integrals, per-zone… See the full description on the dataset page: https://huggingface.co/datasets/projectsim/articulated-wardrobes.docqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.booru_artists_merged
booru_artists_merged
About
This data set is a parquet file containing artist lists extracted from danbooru (2024/12/1x version)
Excluding those with a post_count of 0
Please use this as a reference when processing data locally, such as creating CSV files.Example of processing to create CSV and translation for tag completion of A1111
This file was created based on the danbooru API.
Field Name Description
The field names are as you see, but some… See the full description on the dataset page: https://huggingface.co/datasets/supercatdoing/booru_artists_merged.articulated-logistics-cages
projectsim Logistics Cages
Try it before you download. Open the Cages tab in projectsim Lab,
swivel the casters and roll the wheels, then copy a one-line download command for Isaac Sim, MuJoCo, usdview or Blender.
projectsim by Kaedim — sim-ready 3D assets for robot learning. Open in your engine in one paste.
Warehouse wire roll cages with articulated swivel casters, rolling wheels, and drop gates — joint chains included. Every asset is a reduced-coordinate… See the full description on the dataset page: https://huggingface.co/datasets/projectsim/articulated-logistics-cages.shape-sorting-so101-30fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"Rotation",
"Pitch",
"Elbow",
"Wrist_Pitch",
"Wrist_Roll",
"Jaw"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101-30fps.robusto-2
Dataset: Robusto-2
Paper Link on ArXiv: https://arxiv.org/abs/2606.20980
Description
This dataset contains 20 videos, which were specifically used in this paper. These videos were selected from a larger set of 200 dashcam videos recorded in various cities across Peru (Lima) and New York City (NYC), available as an extended dataset. They are split evenly by region — 10 from Lima/Peru and 10 from NYC — so model and human behavior can be compared across a familiar… See the full description on the dataset page: https://huggingface.co/datasets/Artificio/robusto-2.artem_screwdriver_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 10,
"total_frames": 11525,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/artem_screwdriver_100.ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web.
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import load_dataset
dataset = load_dataset("siavava/ai-tech-articles")
# optionally, convert it to a pandas dataframe:
df = dataset["train"].to_pandas()
You do not need to clone… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.
