CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RoganInglis /vllm-control-arena vLLM Main Tasks Dataset AI coding tasks generated from vLLM git commits Dataset Description This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work. Dataset Structure The dataset contains the following columns: commit_hash: The git commit hash parent_hash: The parent commit hash commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.tabulartext-generation1K<n<10K0 likes20k downloads1y agoHugging Face02Longitude-Labs /spreadsheet-arena-release Spreadsheet Arena A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models. This is the public release accompanying the Spreadsheet Arena paper. Contents battles.csv models.csv outputs/<id>/ sheet.json sheet.xlsx <id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.tabulartabular-classificationn<1K5 likes4.7k downloads4mo agoHugging Face03sprinklr-huggingface /CXM_Arena Dataset Card for CXM Arena Benchmark Suite Dataset Description This dataset, "CXM Arena Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain. It consolidates five distinct tasks into a unified benchmark, enabling robust testing of models and pipelines in business contexts. The entire suite was synthetically generated using advanced large language models, primarily… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena.tabulartext-ranking10K<n<100K3 likes3.7k downloads1y agoHugging Face04CohereLabs /m-ArenaHard-v2.0 Dataset Card for m-ArenaHard-v2.0 This dataset is used in the paper When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs. Dataset Details The m-ArenaHard-v2.0 dataset is a multilingual LLM evaluation set. This is built on the LMarena (formerly LMSYS) arena-hard-auto-v2.0 test dataset. This dataset(containing 750 prompts) was filtered to "english" only prompts using the papluca/xlm-roberta-base-language-detection model resulting… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.0.texttext-generation10K<n<100K7 likes666 downloads5mo agoHugging Face05CohereLabs /m-ArenaHard-v2.1 Dataset Card for m-ArenaHard-v2.1 The m-ArenaHard-v2.1 dataset is a multilingual LLM evaluation set built from the LMarena arena-hard-auto-v2.0 prompts used in m-ArenaHard-v2.0. It keeps the public v2.0 row schema while expanding coverage to the 67 raw translation files produced for the Tiny Aya evaluation work. The dataset includes 67 languages: am, ar, bg, bn, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fr, ga, gl, gu, ha, he, hi, hr, hu, id, ig, it, ja, jv, km, ko, lo, lt, lv… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.1.texttext-generation10K<n<100K3 likes395 downloads5mo agoHugging Face06martincousseau /refund-arena Refund Arena Gold for a helpdesk agent that pays. Used by Refund Arena to score whether money left that should not have left this turn. A customer asks for a refund the policy does not allow. The agent closes the ticket and calls issue_refund. Support now owes money or has to claw it back. I do not ship if that write happens. This dataset is text JSONL only. Card images live on GitHub (docs/assets/). They are not a split. Do not load this repo as ImageFolder. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/martincousseau/refund-arena.tabulartext-generationn<1K0 likes171 downloads9d agoHugging Face07tintin1027 /research-ideation-arena-si-rm Research Ideation Arena — Scientific Ideation RM Splits Derived from Research Ideation Arena, revision f5704385bd66781d504e44810a9a8b56c1623b7a. Original authors: Zhiyu Chen et al. See the paper and official code. Splits and evaluation caveat Train: 3,047 preference pairs. Test: 500 fixed preference pairs. All remaining pairs from the 3,547-pair filtered pool are assigned to training. Exact sample/pair overlap is zero, but 607 training rows share a connected… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/research-ideation-arena-si-rm.texttext-classification1K<n<10K0 likes101 downloads8d agoHugging Face08ministere-culture /comparia-fr-arenagated comparia-fr-arena: French-language conversations and human preferences Compar:IA is a public chatbot arena run by the French Ministry of Culture. People chat with two anonymous models side by side and say which answer they prefer. This dataset is the result. Each row is one turn of a conversation: the two models' answers to the same user message, the preference the user gave on that turn (if any), and the full conversation both answers belong to. The prompts come… See the full description on the dataset page: https://huggingface.co/datasets/ministere-culture/comparia-fr-arena.texttext-generation100K<n<1M15 likes89 downloads1mo agoHugging Face09CK0607 /swarm-arena-sft-v2 Swarm Arena SFT v2 Solver-filtered warm-start data for the deterministic Swarm Arena 4v4 coordination environment. Each row contains system, user, and assistant messages plus provenance metadata. Training broadcasts and actions are separate splits so sampling can preserve a 60/40 phase mixture. Validation and test are never reweighted. The simulator, oracle, audit, frozen evaluation, and Prime-RL configs live in… See the full description on the dataset page: https://huggingface.co/datasets/CK0607/swarm-arena-sft-v2.texttext-generation1K<n<10K0 likes75 downloads2mo agoHugging Face10THU-KEG /Arena-Write 📚 Arena-Write Dataset Arena-Write is a small-scale benchmark of 100 user writing tasks, designed to evaluate long-form generation models in realistic scenarios. Each task covers diverse formats such as social posts, essays, and reports, with many requiring outputs over 2,000 words. Project page: https://huggingface.co/THU-KEG/ 📄 Data Format Each data sample is a JSON object with the following fields: { "idx": 1, "question": "Write a social media post about Lei Feng… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/Arena-Write.tabulartext-generationn<1K5 likes69 downloads1y agoHugging Face11lesserfield /lmsys-arena-human-preference-winner-43k-unfiltered lmsys-arena-human-preference-winner-43k-unfiltered This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License. Dataset Description The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.texttext-generation10K<n<100K2 likes67 downloads2y agoHugging Face12qwopqwop /ko-arena-hard-auto-v0.1 Ko-Arena-Hard-Auto 한국어 / English 리더보드 / 코드 ko-arena-hard-auto-v0.1는 한국어를 벤치마킹하기위한 자동 평가 도구의 질문 데이터셋입니다. 인간의 선호도와 높은 상관관계와 분리력을 가지고 있는 벤치마크 데이터셋인 arena-hard-auto-v0.1 를 GPT-4o와 o1을 사용하여 한국어로 번역하고 수작업으로 검수한 데이터셋입니다. 더 자세한 세부사항과 벤치마킹 결과는 ko-arena-hard-auto 코드를 참조하세요. 또한 원본 벤치마크에 관심이 있으시면 arena-hard-auto 코드를 참조하세요. 원래 문제의 형식을 유지하기 힘들어서 변경했습니다. 인덱스 : 1, 28, 29 문제를 한국어로 유도하기 위해 문제 형식을 변경했습니다. 원래는 코드만 존재합니다. 인덱스 : 30, 379, 190 참고문헌: @article{li2024crowdsourced, title={From Crowdsourced… See the full description on the dataset page: https://huggingface.co/datasets/qwopqwop/ko-arena-hard-auto-v0.1.texttext-generationn<1K16 likes53 downloads1y agoHugging Face13rajistics /rag-qa-arena RAG QA Arena Annotated Dataset A comprehensive multi-domain question-answering dataset with citation annotations designed for evaluating Retrieval-Augmented Generation (RAG) systems, featuring faithful answers with proper source attribution across 6 specialized domains. 🎯 Dataset Overview This annotated version of the RAG QA Arena dataset includes citation information and gold document IDs, making it ideal for evaluating not just answer accuracy but also answer grounding… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/rag-qa-arena.tabularquestion-answering10K<n<100K0 likes41 downloads1y agoHugging Face14woog /arena-prose-100-49-models Arena Prose: 100 prompts × 50 models A paired exploratory AI-text-detection corpus: 5,000 successful generated responses from 50 models, each answering the same 100 English prose prompts. Generation was performed through OpenRouter in September 2026 with optional reasoning disabled and mandatory reasoning set to low. This is an independent local benchmark inspired by Pangram 4 §5.2, not an official Pangram dataset or exact replication. Loading from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/woog/arena-prose-100-49-models.tabulartext-generation1K<n<10K0 likes32 downloads2d agoHugging Face15nlee-208 /ko-arena-hard-v2 ko-arena-hard-v2 Korean adaptation of arena-hard-v2, an LMSYS-style head-to-head benchmark where a strong judge model rates a candidate model's answer against a fixed baseline answer. 491 prompts in Korean, drawn from CohereLabs/m-ArenaHard-v2.0 (config ko) Frozen baselines generated by gpt-5-mini with category-specific reasoning effort Recommended judge: gpt-4.1 via the upstream arena_judge pipeline (dual A↔B scoring) Schema Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/nlee-208/ko-arena-hard-v2.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face16ShahzebKhoso /local-code-arena-deepseek-r1_1.5b Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 1.5B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 1.5B distilled reasoning architecture. This specific run establishes the performance boundaries of lightweight reasoning models under strict execution time limits on consumer hardware. 📊 Core Performance… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-deepseek-r1_1.5b.texttext-generationn<1K1 likes29 downloads4mo agoHugging Face17ArenaRune /EldenRingQA 🗡️ Elden Ring QA Dataset A domain-specific question-answering dataset for Elden Ring, covering weapons, bosses, armors, spells, NPCs, locations, creatures, skills, and ashes of war — including cross-entity boss vulnerability analysis and per-build weapon recommendations. Uses Intended Uses Fine-tuning language models for Elden Ring domain-specific QA Training instruction-following models on structured game knowledge Retrieval-augmented generation (RAG)… See the full description on the dataset page: https://huggingface.co/datasets/ArenaRune/EldenRingQA.textquestion-answering10K<n<100K0 likes25 downloads8mo agoHugging Face18ShahzebKhoso /local-code-master_telemetry_arena Local Code Arena: Comprehensive Telemetry Matrix Dataset 🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures. 📊 Dataset Blueprint This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face19dannyobito /arena-poker-reasoned-decisions-v0 DevFun Arena Poker - Reasoned Decision Traces (v0) 1000 agent decision traces from live 6-max No-Limit Texas Hold'em on the dev.fun AI-agent poker Arena. Each row is one agent's decision at one moment in one hand, paired with the structured rationale the agent emitted for that action. This is a small curated SAMPLE for researchers to judge whether the full data is useful. Each decision is enriched with full per-seat table state (every seat's stack at decision time), all-in… See the full description on the dataset page: https://huggingface.co/datasets/dannyobito/arena-poker-reasoned-decisions-v0.tabularreinforcement-learning1K<n<10K1 likes25 downloads4mo agoHugging Face20ShahzebKhoso /local-code-arena-eval_results_deepseek-r1_8b Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 8B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 8B distilled reasoning architecture. This specific run establishes the operational baseline and performance characteristics of mid-tier reasoning models under strict automated evaluation and sandbox time limits on… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-eval_results_deepseek-r1_8b.texttext-generationn<1K1 likes21 downloads4mo agoHugging Face21nwirandx /ko-arena-hard-auto-v0.1qwopqwop님이 ko-arena-hard를 번역하신 데이터 qwopqwop/ko-arena-hard-auto-v0.1에 tag를 단 데이터입니다. tag 정보 Category Count Description Coding & Debugging 279 Users seek help with writing, reviewing, or fixing code in programming. Planning 67 Users need assistance in creating plans or strategies for activities and projects. Data analysis 31 Requests involve interpreting data, statistics, or performing analytical tasks. Math 26 Queries related to mathematical concepts, problems, and… See the full description on the dataset page: https://huggingface.co/datasets/nwirandx/ko-arena-hard-auto-v0.1.texttext-generationn<1K0 likes20 downloads2y agoHugging Face22ZachW /nanbeige4-3b-thinking-2511_arena-hard-creative-writing Nanbeige/Nanbeige4-3B-Thinking-2511 — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: Nanbeige/Nanbeige4-3B-Thinking-2511 Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_arena-hard-creative-writing.tabulartext-generationn<1K0 likes20 downloads5mo agoHugging Face23danish-foundation-models /ai-arenaen-conversationsgated AI Arenaen Conversations A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform. Origin of the data: what is AI-Arenaen? The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.tabulartext-generation1K<n<10K1 likes17 downloads4mo agoHugging Face24ZachW /qwen3-8b_arena-hard-creative-writing Qwen/Qwen3-8B — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-8B Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_arena-hard-creative-writing.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face25ShahzebKhoso /local-code-arena-mbpp-starcoder_7b Local Code Arena Telemetry: MBPP Benchmark on StarCoder 7B (Base) This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the legacy StarCoder 7B base foundational model. This specific partition documents the behavioral dynamics of larger-scale raw foundational weights inside automated benchmarking pipelines, establishing an anchor point to analyze… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-starcoder_7b.texttext-generationn<1K1 likes17 downloads4mo agoHugging Face26arenard /wildprompt-5k5000 randomly selected prompts extracted from the WildChat-4.8M dataset. Prompts extracted only from rows matching the following criteria: Not toxic Not redacted Minimum of 2 user turns texttext-generation1K<n<10K0 likes15 downloads8mo agoHugging Face27ZachW /qwen3-32b_arena-hard-creative-writing Qwen/Qwen3-32B — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-32B Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_arena-hard-creative-writing.tabulartext-generationn<1K0 likes15 downloads5mo agoHugging Face28ShahzebKhoso /local-code-arena-mbpp-qwen2.5-coder_14b Local Code Arena Telemetry: MBPP Benchmark on Qwen 2.5 Coder 14B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the heavyweight Qwen 2.5 Coder 14B parameter model. This specific run establishes the heavy-parameter upper bound of our local consumer hardware evaluation matrix, isolating how peak capacity interacts with strict functional code… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-qwen2.5-coder_14b.texttext-generationn<1K1 likes15 downloads4mo agoHugging Face29ShahzebKhoso /local-code-arena-mbpp-qwen3.5_2b Local Code Arena Telemetry: MBPP Benchmark on Qwen 3.5 2B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the Qwen 3.5 2B base model. This specific run documents how a tiny, next-generation generalist instruction model handles zero-shot functional programming synthesis under standard local execution bounds without active reasoning tokens.… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-qwen3.5_2b.texttext-generationn<1K1 likes15 downloads4mo agoHugging Face30ShahzebKhoso /local-code-arena-mbpp-deepseek-coder_1.3b Local Code Arena Telemetry: MBPP Benchmark on DeepSeek Coder 1.3B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the ultra-lightweight DeepSeek Coder 1.3B model. This specific run establishes the absolute maximum throughput envelope of our local hardware setup while tracking the accuracy trade-offs of legacy, lightweight code specialists.… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-deepseek-coder_1.3b.texttext-generationn<1K1 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.