datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/VIBE.Vibe-Coding-Instructvibesec
VibeSec V1.1
1,000 execution-verified security patching tasks for AI coding agents.
Each task contains a vulnerable FastAPI app, a normal-behavior spec test, an
exploit that prints PWNED against the vulnerable app, and a reference patch.
Changes from v1.0.0
v1.0.0's 1,000 tasks contained generation-seed fan-out duplicates — only 755
unique scenarios. v1.1 collapses to those 755 and adds 245 seed-unique tasks
(238 from an under-represented-class batch + 7 for… See the full description on the dataset page: https://huggingface.co/datasets/muence/vibesec.human-vibevibepass
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
Authors: Srijan Bansal, Jiao Fangkai, Yilun Zhou, Austin Xu, Shafiq Joty, Semih Yavuz
TL;DR: As LLMs shift programming toward human-guided "vibe coding", agentic tools increasingly rely on models to self-diagnose and repair their own subtle faults—a capability central to autonomous software engineering yet never systematically evaluated. VIBEPASS presents the first empirical benchmark that decomposes fault-targeted reasoning into… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/vibepass.agentic-vibecoding-traces
🧠 Agentic Vibecoding Traces
3.5 years of real agentic coding sessions across 4 CLI agents and 25+ teacher models — fully anonymized, segmented per-task, with complete tool-call trajectories (bash commands + outputs, file edits) and chain-of-thought reasoning.
The culmination dataset: every "vibe coding" session, extracted from local agent storage, scrubbed, and packaged for SFT.
[!IMPORTANT]
Gated access. Access requests are reviewed manually. Data is anonymized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/agentic-vibecoding-traces.bf-vibe-bibframe-corrections
BF Vibe BIBFRAME Corrections (v3)
Training data for the single open model behind BF Vibe, a desktop
assistant that helps catalogers create and repair valid BIBFRAME RDF/XML.
The core task is BIBFRAME correction — (corrupted → conforming) record
pairs verified against SHACL shapes — supplemented by two smaller task
families that teach the model to operate the application: routing free-text
requests to BF Vibe's commands and calling its grounding tools.
Created by: Jim Hahn… See the full description on the dataset page: https://huggingface.co/datasets/jimfhahn/bf-vibe-bibframe-corrections.MiniMaxAI-VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/John1604/MiniMaxAI-VIBE.vibethinker-3b-finance-sftmini-data-public-version
VibeThinker-3B Finance-Reader — SFT Training Data · PUBLIC-SAFE subset
🟢 This is vibethinker-3b-finance-sftmini-data-public-version — the redistribution-safe slice of the
full vibethinker-3b-finance-sftmini-data
dataset, containing only US-government public-domain sources (SEC EDGAR family + Federal Register).
Same schema, same pipeline, same teacher — just the legally shareable rows. (Currently private; intended to be made public.)
The supervised fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/vibethinker-3b-finance-sftmini-data-public-version.vcl-vibebench
VCL VibeBench
Stop trusting benchmark slides. Run it yourself.
Practical AI model prompts from Vibe Coder's Life.
This dataset mirrors the open-source GitHub suites. It is not a blended intelligence leaderboard. There is no LLM judge. 9/12 on Score means nine JavaScript helpers compiled and passed hidden unit tests — not “75% smart.”
Configs
Config
What it is
Rows
fun
Fun 1.0 — short copy-paste prompts
10
dev
Dev 1.1 — coding / debugging prompts
10… See the full description on the dataset page: https://huggingface.co/datasets/kondasviktor/vcl-vibebench.VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/djeroo/VIBE.VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/tatan2/VIBE.vibeapps-chat-fabric
vibeapps-chat-fabric — agentic ChatML
1329 full agentic coding trajectories. A fake user persona drives a coding agent
to build a self-contained web app, critiquing functionality + aesthetics over 3 turns.
Browse the live apps: https://huggingface.co/spaces/AlexWortega/vibeapps-chat-fabric
coder/assistant: minimax/minimax-m3 (pi agent) - user-sim: qwen/qwen3.7-max
personas: dotoshny (nitpicky), mamochka (non-techie mom), startuper - 443 ideas x 3 personas x 3 turns… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/vibeapps-chat-fabric.Vibe-Coding-InstructVIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/Rendy45/VIBE.Vibe-Coding-InstructVibe-Coding-InstructVibe-Coding-Instruct
