datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goat
Dataset Card for Dataset Name
Dataset Summary
The dataset.json file contains ~1.7 million synthetic data for arithmetic tasks, generated by dataset.ipynb.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/goat.vetbench
VET-Bench: Visual Entity Tracking Benchmark
Paper | Project Page | GitHub
VET-Bench is a synthetic diagnostic benchmark simulating the realistic shell game with visually indistinguishable objects that forces models to track entities through spatiotemporal continuity. The task is easy for human but difficult for current VLMs. State-of-the-art VLMs perform at random chance, while our proposed Molmo2-SGCoT achieves over 90% accuracy.
Dataset Overview
Cup Game
Card… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/vetbench.ultrafeedback_tied
Train dir contains train set with different ratios of tie data
Test dir contains test sets which used to evaluate performances on the in-distribution data.
test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data.
non_tie_data_test.jsonl contains 1500 non-tie samples.
tie_data_test.jsonl contains 500 tie samples.
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.chatarena_tied
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}: Enhancing {LLM} Alignment with Ternary Preferences},
author={Yuxiang Guo and Lu Yin and Bo Jiang and Jiaqi Zhang},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=utkGLDSNOk}
}
Molmo2-SGCoT
Molmo2-SGCoT: Spatiotemporal Grounded Chain-of-Thought Training Data
Paper | Project Page | GitHub
Training data for aligning Molmo2 to perform Spatiotemporal Grounded Chain-of-Thought (SGCoT) on VET-Bench — generating explicit object tracking trajectories before answering questions.
Overview
This dataset contains 300 synthetic samples where the model generates a structured trajectory <tracks> producing a final answer. The trajectories encode spatial coordinates (x, y… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/Molmo2-SGCoT.tiedpo_mantis_visual_story_telling
