1-real
Datasets
All datasets matching “1-real”nyush-galaxea-a1-lingbot-va-real-world-evaluations
LingBot-VA on Galaxea A1 — Real-World Evaluations
Fruit-placement rollouts and open-loop diagnostics of object grounding,
layout generalization, and predicted robot motion.
Fruit step-1000: lemon-to-plate rollout in the Official layout.
Evidence
Scale
Real closed-loop rollouts
61 archived; 60 scored
Matched base-model controls
9 predictions
Post-trained diagnostics
48 full-horizon predictions; 1,211 rolling futures
Controlled OOD studies
558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.apex-r1-real-world-documents
Apex-R1 Real-World Benchmark Documents
This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation.
The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks.
Contents
benchmark_documents/
EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.real01b-routing-d1-umirel-lineage-15arm-heldout-sobol50-s2026090701This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-umirel-lineage-15arm-heldout-sobol50-s2026090701.RealText-V1
RealText-V1: A Text-Centric Image Forgery Analysis Dataset
💾 Dataset Description
RealText-V1 is a text-centric image forgery analysis dataset built to benchmark visual-logical co-reasoning over text-centric image forgeries. It pairs forged and pristine document-like text images with pixel-level manipulation masks and expert-level natural-language explanations that ground every verdict in observable visual and logical evidence.
RealText-V1 is the dataset… See the full description on the dataset page: https://huggingface.co/datasets/vankey/RealText-V1.chunk1-realmdlens-realdocs-v1
mdlens-realdocs-v1
A held-out Markdown QA / retrieval eval built entirely from real open-source
project documentation. It measures whether an agent can answer documentation
questions from the right evidence with fewer irrelevant reads and fewer tokens.
Questions are deliberately low lexical overlap (paraphrased), so they stress
retrieval rather than string matching. Every non-abstention question has its
answer keywords verified to appear in the cited source file.… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/mdlens-realdocs-v1.
