datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FullStackBenchFullStack Bench: Evaluating LLMs as Full Stack Coders
Official repository for our paper "FullStack Bench: Evaluating LLMs as Full Stack Coders"
🏠 FullStack Bench Code •
📊 Benchmark Data •
📚 SandboxFusion
📌Introduction
FullStack Bench is a multilingual benchmark for full-stack programming, covering a wide range of application domains and 16 programming languages with 3K test samples, which substantially pushes… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/FullStackBench.fullstackbench-tasks
fullstackbench-tasks
Harbor task definitions used to benchmark coding agents in the FullStackBench / Loki research program. Each top-level directory is one self-contained task — Dockerfile, golden app, verifier, instruction — runnable via harbor run.
Prerequisites
hf CLI logged in to an ApolloTeam-member account (hf auth login)
Working install of Harbor — provides the harbor CLI (ask the Loki team if you don't already have it set up)
Docker daemon running locally
~4 GB… See the full description on the dataset page: https://huggingface.co/datasets/ApolloTeam/fullstackbench-tasks.fullstackbench-trajectories
fullstackbench-trajectories
FullStackBench tasks and the agent trajectories that ran them. One repo, two sibling trees keyed by the same task ID.
Layout
tasks/
└── <task-id>/ # task definition
├── environment/ # docker-compose, source app, fixtures
├── tests/ # verifier scripts
└── solution/ # reference impl (if applicable)
trajectories/
└── <task-id>/… See the full description on the dataset page: https://huggingface.co/datasets/ApolloTeam/fullstackbench-trajectories.FullStack-Bench
FullStack-Agent
Paper | Code | Dataset
Overview
This repository contains the FullStack-Bench dataset, introduced in the paper "FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation".
In this paper, we propose FullStack-Agent, a unified system that combines a multi-agent full-stack development framework equipped with efficient coding and debugging tools (FullStack-Dev), an iterative self-improvement method… See the full description on the dataset page: https://huggingface.co/datasets/luzimu/FullStack-Bench.stargate_s04e01_100topkdiverse_text2vid
advanced-fullstack-ai-knowledge-base
Advanced Full-Stack & AI Engineering Knowledge Base (2026 Edition)
This repository contains a high-quality, production-ready sample subset of 23,734 records from a massive, proprietary dataset meticulously curated for Retrieval-Augmented Generation (RAG) systems, Agentic Workflows, and Fine-Tuning next-generation LLMs.
Overview & The Knowledge Cutoff Solution
One of the most persistent bottlenecks in production AI systems is the knowledge cutoff. Most… See the full description on the dataset page: https://huggingface.co/datasets/kooda-ai/advanced-fullstack-ai-knowledge-base.synth-illuminati-cardsstage22fullstack-templatescik_sec_api_jsonlfullstack-templates-formattedstage3human-chromhmm-fullstack-data
human-chromhmm-fullstack-data
Dataset Summary
This dataset provides a multi-class annotation of genomic regions across the hg38 genome. It is derived from the ChromHMM fullstack annotation (Vu & Ernst, 2022; https://doi.org/10.1186/s13059-021-02572-z). Genomic regions are classified into 16 chromatin states. The data is derived from https://public.hoffman2.idre.ucla.edu/ernst/2K9RS//full_stack/full_stack_annotation_public_release/hg38/hg38_genome_100_segments.bed.gz.… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/human-chromhmm-fullstack-data.stage4_goldstage4_silverICONN-FullStack-Reasoningstage4_silver_newLLaVA-CoT-30k-base64-in-jsonlfullstack-backend-trainingkyc_passport_image_textstage1stage2stage4_silver_v1fullstackarena-table-study-v1
FullStackArena three-site table study (private working dataset)
Protocol fullstackarena-three-site-table-study-v1 (status qualified): five exact OpenRouter models × three cells × the signed Arenagram (124), RideApp (120) and Auction (100) task packs = 45 arms, 5,160 scored attempts, max_steps=30, five workers on five isolated lanes per batch.
Cells: table1_no_time_capped, table1_no_time_uncapped (no automatic time; the explicit get_website_time() tool stays available)… See the full description on the dataset page: https://huggingface.co/datasets/GroupieSteven/fullstackarena-table-study-v1.emgena_fullstack_mcp_trinity_suite_teaser
🚀 Emgena Full-Stack MCP Defense Trinity Suite (3-in-1 Cursor & Claude Plugin) (Free Community Teaser)
⚡ Official Free Evaluation Teaser & IDE Plugin Template🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Pro Plugin on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 What is this?
The Ultimate 3-in-1 Engineering Triad: SRE Guard + SecOps Armor + DataOps Performance (12 Tools, Unified Config).
100% Zero External… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_fullstack_mcp_trinity_suite_teaser.
