datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts-TBLite
Blog Post |
GitHub |
Dev Set v1
OpenThoughts-TBLite
A Difficulty-Calibrated Benchmark for Building Terminal Agents
By OpenThoughts Agent team, Snorkel AI, Bespoke Labs
OpenThoughts-TBLite is a curated collection of 100 Terminal-Bench tasks that closely track TB2 performance, but run much faster. It's designed to be more informative during model development, making it ideal for debugging, iteration, and training ablations.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Virtual-Production-Units/OpenThoughts-TBLite.OpenThoughts-TB-dev-v2
Blog Post |
GitHub |
Dev Set v1
OpenThoughts-TB-dev-v2
A Difficulty-Calibrated Benchmark for Building Terminal Agents
By OpenThoughts Agent team, Snorkel AI, Bespoke Labs
OpenThoughts-TB-dev-v2 is a curated collection of 100 Terminal-Bench tasks that closely track TB2 performance, but run much faster. It's designed to be more informative during model development, making it ideal for debugging, iteration, and training ablations.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Virtual-Production-Units/OpenThoughts-TB-dev-v2.Mind2Web_bbox_eval
Mind2Web evaluation set for the paper: Harnessing Webpage Uis For Text Rich Visual Understanding
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Introduction
We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multi- modal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks—achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in action accuracy on a web agent dataset Mind2Web—but… See the full description on the dataset page: https://huggingface.co/datasets/Virtual-Production-Units/Mind2Web_bbox_eval.
