datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.ps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.financial-analyst-data-demo
financial-analyst-data-demo
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~0.16 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-demo.ucbshift-refined
UCBShift refined split
This is a ready-to-use, dataset-level export of the frozen 678/75/200 protein
split. Target cleaning is already materialized: a null atom cell is excluded
from training and evaluation.
split
proteins
residue rows
eligible atom targets
rows with structural features
train
678
75974
476244
75974
validation
75
8738
53098
8738
test
200
22510
125088
21531
Files
train.parquet, validation.parquet, and test.parquet: identical… See the full description on the dataset page: https://huggingface.co/datasets/yiming421/ucbshift-refined.ps4mas-castle-cold
PS4MAS CASTLE cold
2026-09-15: Removed seven audited legacy handwritten-dataset experiment entries at the user's explicit request. Their IDs and query texts did not match the sealed official CASTLE-200. The verified official input under datasets/official_castle200/ is retained. Only newly validated results belong in the active catalog. Historical score-only results, when present, do not prove full-hop or oracle-profile correctness.
Current CASTLE / cold experiment… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-castle-cold.robocasa-combinedted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.ps4mas-castle-oracle
PS4MAS CASTLE oracle
Current CASTLE / oracle experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-castle-oracle.ps4mas-ps-cold
PS4MAS PS Cold
Domain: PS (Personal Safety)
Mode: Cold — Zero-shot generalization (no training on target topology)
Repository: yinita/ps4mas-ps-cold
Experiment Table
Model
Topologies
Episodes
Safety
Reward
Date
HF
gpt-5.6-luna
central,cove,critique,debate,discuss_decide,meta,orchestrate,plan_execute,red_blue,single,vote
2200
2.527
-0.473
2026-09-14
HF
Description
This dataset contains evaluation episodes for the PS4MAS project:… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-ps-cold.details_01-ai__Yi-1.5-34B-Chat
Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-1.5-34B-Chat.details_01-ai__Yi-9B-200K
Dataset Card for Evaluation run of 01-ai/Yi-9B-200K
Dataset automatically created during the evaluation run of model 01-ai/Yi-9B-200K.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-9B-200K.rubrichub-sft-judgment-genps4mas-ps-scale
PS4MAS PS scale
Current PS / scale experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-ps-scale.ps4mas-ps-oracle
PS4MAS PS oracle
Current PS / oracle experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-ps-oracle.OntoKGsolana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-22T12:05:13.833Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.ps4mas-castle-scale
PS4MAS CASTLE scale
Current CASTLE / scale experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-castle-scale.czech_bank_qa
CzechBankQA
This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset.
LSV
LSV: LabSuperVision Benchmark
Dataset Description
LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors.
The dataset is designed for research on:
Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.video_mmeI-SHEEP-Self-Datasort_9This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3446,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yinxinyuchen/sort_9.pick_place_0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3191,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yinxinyuchen/pick_place_0.ecommerce-user-behavior-dataIvy-Fake
IVY-FAKE: Unified Explainable Benchmark and Detector for AIGC Content
This repository provides the official implementation of IVY-FAKE and IVY-xDETECTOR, a unified explainable framework and benchmark for detecting AI-generated content (AIGC) across both images and videos.
🔍 Overview
IVY-FAKE is the first large-scale dataset designed for multimodal explainable AIGC detection. It contains:
150K+ training samples (images + videos)
18.7K evaluation samples
Fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/yishiliu/Ivy-Fake.open-jobs-daily
Open Jobs Daily 🌍💼
Commercial vendors often charge upwards of $1,000/month for firehose access to global job market data. This dataset democratizes that access.
The main creator of this dataset is Reddit user OminousLatinWord. For convenience, I converted the dataset to Parquet files and uploaded it to Hugging Face.
Source Data & Attribution
Creator: Created and originally open-sourced by Reddit user OminousLatinWord under a CC0 license.
Source Release:… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/open-jobs-daily.pick_place_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3064,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yinxinyuchen/pick_place_4.screw_large_stiff_0901_10hzThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 15,
"total_frames": 8970,
"total_tasks": 2,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YingYuan0414/screw_large_stiff_0901_10hz.sort_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3548,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yinxinyuchen/sort_2.
