datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs
Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs"
More Information needed
coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.local-coding-agent-benchmark
🤖 Local Coding Agent Benchmark (LCAB)
Real-world benchmarking of local AI coding agents on software-repair workloads.
This Hugging Face Dataset contains the reproducibility artifacts, raw agent-session evidence, benchmark results, task source, hardware profiles, and analysis for the Local Coding Agent Benchmark (LCAB).
LCAB is designed to evaluate local coding agents as complete systems—not only by tokens/second, but by how efficiently they transform a real software-repair… See the full description on the dataset page: https://huggingface.co/datasets/amitmaity0/local-coding-agent-benchmark.vep_pathogenic_codingvllm-benchmark-coding
vLLM Benchmarking: Coding
Dataset for easy benchmarking of deployed LLMs in serving mode, designed to be compatible with vLLM's benchmark_serving.py.
Dataset Sources
Dataset is sourced from the following four datasets:
https://huggingface.co/datasets/Crystalcareai/Code-feedback-sharegpt-renamed
https://huggingface.co/datasets/MaziyarPanahi/Synthia-Coder-v1.5-I-sharegpt
https://huggingface.co/datasets/Alignment-Lab-AI/CodeInterpreterData-sharegpt… See the full description on the dataset page: https://huggingface.co/datasets/crozai/vllm-benchmark-coding.ai-coding-assistants-benchmark-2026
AI Coding Assistants Benchmark 2026 — Methodology Dataset
Independent benchmark methodology for evaluating AI coding assistants in 2026. Covers Claude Code (Anthropic), Cursor, GitHub Copilot, Windsurf (Codeium), Aider, Continue.dev, Cody (Sourcegraph), Tabnine, OpenAI Codex CLI, and Replit Agent.
Methodology
Test bench: 12 real-world coding tasks across Python, TypeScript, Rust, Go
Benchmark: SWE-bench Verified scores per tool (cross-language)
Performance:… See the full description on the dataset page: https://huggingface.co/datasets/Ricco020/ai-coding-assistants-benchmark-2026.vllm-benchmark-coding-100vllm-benchmark-coding
vLLM Benchmarking: Coding
Dataset for easy benchmarking of deployed LLMs in serving mode, designed to be compatible with vLLM's benchmark_serving.py.
Dataset Sources
Dataset is sourced from the following four datasets:
https://huggingface.co/datasets/Crystalcareai/Code-feedback-sharegpt-renamed
https://huggingface.co/datasets/MaziyarPanahi/Synthia-Coder-v1.5-I-sharegpt
https://huggingface.co/datasets/Alignment-Lab-AI/CodeInterpreterData-sharegpt… See the full description on the dataset page: https://huggingface.co/datasets/mbicanic/vllm-benchmark-coding.
