datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-Briefcase-Lite
AA-Briefcase-Lite
The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work.
Leaderboard and detailed results
Launch article
AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.fasterwam
fasterwam — Franka Panda teleoperated demonstrations
Four tasks, 100 episodes each. Every task folder has the same flat layout:
<task>/episode_0.hdf5 … <task>/episode_99.hdf5
<task>/merge_manifest.json maps each episode to the original recording it came from
Task
Description
Recorded
Details
place_cube/
Place a cube in a bowl
2026-09-10
place_cube/RECROP.json — crops re-derived 2026-09-18 from the stored raw frames with the same window as the other three tasks… See the full description on the dataset page: https://huggingface.co/datasets/aabyaneh/fasterwam.StateBench-v1
StateBench v1
110 million verified state episodes for pretraining, SFT, and on-policy SDFT
Train models to retain, edit, address, and transform state across long execution traces.
Choose a configuration
Configuration
Scale
Use when
default
10,000,000 programs · 7.80 GiB
You need 36 explicit families and varied surfaces
dense-100m
100M episodes · 98.61 GiB
You need high-throughput state training with grounded, structurally varied… See the full description on the dataset page: https://huggingface.co/datasets/aabbdev/StateBench-v1.cawm
cawm
Tactile manipulation data for world-model training: real-robot teleoperation on
a Franka Panda arm, plus matched simulation rollouts. All streams are 10 Hz.
Preview
Each clip is wrist camera · agentview camera (where one was recorded) ·
FlexiTac contact map, played at 9.6× real time — the data is 10 Hz, so
these are every 8th frame at 12 fps.
These are excerpts, not whole trajectories: a contiguous window of a single
episode, chosen automatically as the… See the full description on the dataset page: https://huggingface.co/datasets/aabyaneh/cawm.simready-assets-pr-test5simready-assets-pr-test2simready-assets-pp-2AABench
AABench - Audio Attack Safety Benchmark
Structure
Audio files in the specified output directory (default: audios/)
A CSV file (dataset_audio_paths.csv) containing:
Index
Dataset name
Prompt text
Target text
Prompt audio path
Target audio path
Filenames
CSV Structure
The generated CSV contains the following columns:
index: Row index
dataset: Source dataset (advbench, jbb, or harmbench)
prompt_text: Original prompt text
target_text: Original… See the full description on the dataset page: https://huggingface.co/datasets/NWULIST/AABench.simready-assets-pp-3seqrec-datasets
数据集与预处理
本目录包含 SeqRec 的数据下载与预处理工具,支持 Amazon Reviews 2014 的全部 24 个品类。Item 语义特征构建代码位于 src/features/。
用户数和物品数均不包含 ID 0 的 [PAD]。
目录结构
datasets/
├── raw_data/AmazonReviews2014/<category>/
│ ├── reviews_<category>_5.json.gz
│ └── meta_<category>.json.gz
├── processed_data/AmazonReviews2014/<category>/
├── preprocess/
│ ├── amazon.py
│ └── common.py
└── prepare_data.py
src/features/
├── build_features.py
├── embedding.py
└── pca.py
raw_data/ 和 processed_data/… See the full description on the dataset page: https://huggingface.co/datasets/aabbxx/seqrec-datasets.eai_parquetbert-text-classificationAA-Briefcase-Lite
AA-Briefcase-Lite
The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work.
Leaderboard and detailed results
Launch article
AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/kimmich07/AA-Briefcase-Lite.simready-assets-pp-4simready-assets-pr-test8simready-assets-pp-5simready-assets-pr-test10details_Aabbhishekk__TinyLlama-1.1B-miniguanaco
Dataset Card for Evaluation run of Aabbhishekk/TinyLlama-1.1B-miniguanaco
Dataset automatically created during the evaluation run of model Aabbhishekk/TinyLlama-1.1B-miniguanaco on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aabbhishekk__TinyLlama-1.1B-miniguanaco.srlwam
srlwam
Sim-to-real data for a Franka Panda two-cube stacking task: red cube on green cube, 30.48 mm
sides. Simulation side generated in Isaac Lab with Mimic; real side captured on the physical cell.
Stack
10x10 grids of random demos from stack_v0, sampled for visual diversity across table wood and
lighting.
Table camera
Wrist camera
stack/sim/stack_v0 — 2000 demos, 20 files, 97.9 GB
Domain-randomized demonstrations, 100 per… See the full description on the dataset page: https://huggingface.co/datasets/aabyaneh/srlwam.simready-assets-pr-test9NLP-test-data-cleansimready-assets-pp5simready-assets-pr-test7simready-test-assets8details_Aabbhishekk__llama2-7b-function-calling-slerp
Dataset Card for Evaluation run of Aabbhishekk/llama2-7b-function-calling-slerp
Dataset automatically created during the evaluation run of model Aabbhishekk/llama2-7b-function-calling-slerp on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aabbhishekk__llama2-7b-function-calling-slerp.simready-assetsyahma-18600AIDR-TRAIN1OSINTsimready-assets-pr-test6
