datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.ceo-quotes-verified-sample
🎙️ CEO Transcripts — Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders
🔥 What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Executive
Role
Why They Matter
Jensen Huang
CEO, NVIDIA
Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.MMK12-QWEN35-9B-VERIFIED
Dataset card for MMK12-QWEN35-9B-VERIFIED
Strongich/MMK12-QWEN35-9B-VERIFIED is a filtered teacher-trace dataset built from FanqingM/MMK12 using Qwen/Qwen3.5-9B. The goal is to provide high-quality multimodal reasoning traces that can be used for off-policy distillation into smaller models.
Each row keeps the original MMK12 example fields and adds a responses column. The responses column contains a list of verified Qwen responses that survived the filtering pipeline, so a… See the full description on the dataset page: https://huggingface.co/datasets/Strongich/MMK12-QWEN35-9B-VERIFIED.
