CoolFace
20 results

TIED

tiedong /goat Dataset Card for Dataset Name Dataset Summary The dataset.json file contains ~1.7 million synthetic data for arithmetic tasks, generated by dataset.ipynb. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/goat.textquestion-answering1M<n<10M37 likes277 downloads3y agoHugging Facetiedong /vetbench VET-Bench: Visual Entity Tracking Benchmark Paper | Project Page | GitHub VET-Bench is a synthetic diagnostic benchmark simulating the realistic shell game with visually indistinguishable objects that forces models to track entities through spatiotemporal continuity. The task is easy for human but difficult for current VLMs. State-of-the-art VLMs perform at random chance, while our proposed Molmo2-SGCoT achieves over 90% accuracy. Dataset Overview Cup Game Card… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/vetbench.textvideo-text-to-textn<1K0 likes169 downloads6mo agoHugging Faceirisxx /ultrafeedback_tied Train dir contains train set with different ratios of tie data Test dir contains test sets which used to evaluate performances on the in-distribution data. test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data. non_tie_data_test.jsonl contains 1500 non-tie samples. tie_data_test.jsonl contains 500 tie samples. Citation Please cite our paper if you find the dataset helpful in your work: @inproceedings{ guo2025todo, title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.tabular10K<n<100K0 likes107 downloads1y agoHugging Faceopen-llm-leaderboard-old /details_BEE-spoke-data__smol_llama-81M-tied Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-81M-tied Dataset Summary Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-81M-tied on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BEE-spoke-data__smol_llama-81M-tied.0 likes87 downloads3y agoHugging Faceopen-llm-leaderboard-old /details_andrijdavid__Macaroni-7b-Tied Dataset Card for Evaluation run of andrijdavid/Macaroni-7b-Tied Dataset automatically created during the evaluation run of model andrijdavid/Macaroni-7b-Tied on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andrijdavid__Macaroni-7b-Tied.0 likes44 downloads3y agoHugging Facetiedong /Molmo2-SGCoT Molmo2-SGCoT: Spatiotemporal Grounded Chain-of-Thought Training Data Paper | Project Page | GitHub Training data for aligning Molmo2 to perform Spatiotemporal Grounded Chain-of-Thought (SGCoT) on VET-Bench — generating explicit object tracking trajectories before answering questions. Overview This dataset contains 300 synthetic samples where the model generates a structured trajectory <tracks> producing a final answer. The trajectories encode spatial coordinates (x, y… See the full description on the dataset page: https://huggingface.co/datasets/tiedong/Molmo2-SGCoT.textvisual-question-answeringn<1K0 likes32 downloads7mo agoHugging Face