CoolFace
8 results

coding-benchmark

katarinagresova /Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs" More Information needed text100K<n<1M4 likes308 downloads3y agoHugging Faceruchit11111 /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.textn<1K1 likes303 downloads23d agoHugging Facerogue-security /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.textn<1K2 likes216 downloads2mo agoHugging Faceamitmaity0 /local-coding-agent-benchmark 🤖 Local Coding Agent Benchmark (LCAB) Real-world benchmarking of local AI coding agents on software-repair workloads. This Hugging Face Dataset contains the reproducibility artifacts, raw agent-session evidence, benchmark results, task source, hardware profiles, and analysis for the Local Coding Agent Benchmark (LCAB). LCAB is designed to evaluate local coding agents as complete systems—not only by tokens/second, but by how efficiently they transform a real software-repair… See the full description on the dataset page: https://huggingface.co/datasets/amitmaity0/local-coding-agent-benchmark.2 likes65 downloads1mo agoHugging FaceBenchmarkDatasets /vep_pathogenic_codingtabulartabular-classification10K<n<100K0 likes23 downloads5mo agoHugging Facecrozai /vllm-benchmark-coding vLLM Benchmarking: Coding Dataset for easy benchmarking of deployed LLMs in serving mode, designed to be compatible with vLLM's benchmark_serving.py. Dataset Sources Dataset is sourced from the following four datasets: https://huggingface.co/datasets/Crystalcareai/Code-feedback-sharegpt-renamed https://huggingface.co/datasets/MaziyarPanahi/Synthia-Coder-v1.5-I-sharegpt https://huggingface.co/datasets/Alignment-Lab-AI/CodeInterpreterData-sharegpt… See the full description on the dataset page: https://huggingface.co/datasets/crozai/vllm-benchmark-coding.text10K<n<100K0 likes17 downloads2y agoHugging Face