TIED
Datasets
All datasets matching “TIED”goat
Dataset Card for Dataset Name
Dataset Summary
The dataset.json file contains ~1.7 million synthetic data for arithmetic tasks, generated by dataset.ipynb.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/goat.vetbench
VET-Bench: Visual Entity Tracking Benchmark
Paper | Project Page | GitHub
VET-Bench is a synthetic diagnostic benchmark simulating the realistic shell game with visually indistinguishable objects that forces models to track entities through spatiotemporal continuity. The task is easy for human but difficult for current VLMs. State-of-the-art VLMs perform at random chance, while our proposed Molmo2-SGCoT achieves over 90% accuracy.
Dataset Overview
Cup Game
Card… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/vetbench.ultrafeedback_tied
Train dir contains train set with different ratios of tie data
Test dir contains test sets which used to evaluate performances on the in-distribution data.
test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data.
non_tie_data_test.jsonl contains 1500 non-tie samples.
tie_data_test.jsonl contains 500 tie samples.
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.details_BEE-spoke-data__smol_llama-81M-tied
Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-81M-tied
Dataset Summary
Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-81M-tied on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BEE-spoke-data__smol_llama-81M-tied.details_andrijdavid__Macaroni-7b-Tied
Dataset Card for Evaluation run of andrijdavid/Macaroni-7b-Tied
Dataset automatically created during the evaluation run of model andrijdavid/Macaroni-7b-Tied on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andrijdavid__Macaroni-7b-Tied.Molmo2-SGCoT
Molmo2-SGCoT: Spatiotemporal Grounded Chain-of-Thought Training Data
Paper | Project Page | GitHub
Training data for aligning Molmo2 to perform Spatiotemporal Grounded Chain-of-Thought (SGCoT) on VET-Bench — generating explicit object tracking trajectories before answering questions.
Overview
This dataset contains 300 synthetic samples where the model generates a structured trajectory <tracks> producing a final answer. The trajectories encode spatial coordinates (x, y… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/Molmo2-SGCoT.
