sap
Datasets
All datasets matching “sap”QueST-PartNetMobility-SAPIEN
QueST: PartNet-Mobility SAPIEN Simulation Dataset
This dataset accompanies the paper:
QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon TrackingMayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla, G.C Nandi, Andrew MelnikCAO Workshop at ICLR 2026
What Is This Dataset?
Synchronized RGB-D simulation sequences rendered in SAPIEN from PartNet-Mobility articulated objects, designed to stress-test long-horizon point… See the full description on the dataset page: https://huggingface.co/datasets/AnandMayank/QueST-PartNetMobility-SAPIEN.RAVine-logs
RAVine-logs
This repository contains the running logs of the experiments conducted in the paper RAVine: Reality-Aligned Evaluation for Agentic Search. These logs can be used for result reproduction or detailed case analysis of agentic LLMs with search performance.
RAVine is a comprehensive evaluation system for agentic search, encompassing the web environment, benchmark datasets, and a novel evaluation method, serving as a full-process, reproducible, and goal-aligned evaluation… See the full description on the dataset page: https://huggingface.co/datasets/sapphirex/RAVine-logs.DILLO-LIBERO-dataset
DILLO LIBERO Distillation Dataset
Paper: Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World ModelsCode: github.com/MaxPappa/DILLO
This dataset contains LIBERO policy rollouts labeled for DILLO (DIstiLLed Language-ActiOn World Model). Each example stores a chunked ACT policy rollout, boundary-frame images, robot state traces, action chunks, and VLM-generated descriptions/reasoning for distillation.
Dataset Summary
Total episodes: 1700… See the full description on the dataset page: https://huggingface.co/datasets/Sapienza/DILLO-LIBERO-dataset.HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.Perturb-Sapiens
Perturb Sapiens: A Human Whole-Organism Atlas of Perturbed Cells
Dataset Description
Perturb Sapiens is an evolving database of AI-predicted single-cell perturbation responses, representing the first human whole-organism atlas of perturbed cells.
Perturb Sapiens is generated using the post-trained Stack model (Stack-Large-Aligned), an in-context learning foundation model for single-cell biology.
Data Sources:
Prompt Data: Parse/OpenProblems PBMC perturbation data
Query… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/Perturb-Sapiens.
