CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01R2E-Gym /R2E-Gym-Litetabular10K<n<100K1 likes68k downloads2y agoHugging Face02R2E-Gym /R2E-Gym-V1tabular1K<n<10K2 likes51k downloads2mo agoHugging Face03R2E-Gym /R2E-Gym-Subsettabular1K<n<10K29 likes28k downloads2mo agoHugging Face04r2e-edits /swesmith-cleantext1K<n<10K0 likes6.3k downloads1y agoHugging Face05taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes5.8k downloads3y agoHugging Face06R2E-Gym /SWE-Bench-Verifiedtextn<1K0 likes5.4k downloads2y agoHugging Face07PrimeIntellect /R2E-Gym-Subset-Verified R2E-Gym-Subset-Verified Gold-patch-validated subset of R2E-Gym/R2E-Gym-Subset (paper). The train split contains 4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the gold patch, run the upstream /testbed/run_tests.sh baked into the row's image, check the parsed outcomes against expected_output_json. Changes vs upstream Validation-only subset — our passes, run in fresh sandboxes per row: one full pass at concurrency 200, then a 10× retry pass over… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/R2E-Gym-Subset-Verified.tabulartext-generation1K<n<10K1 likes3.1k downloads3mo agoHugging Face08R2MED /PMC-Treatment 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Treatment.texttext-retrieval10K<n<100K0 likes1.5k downloads1y agoHugging Face09R2MED /PMC-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Clinical.texttext-retrieval10K<n<100K0 likes1.4k downloads1y agoHugging Face10R2MED /MedXpertQA-Exam 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedXpertQA-Exam.texttext-retrieval10K<n<100K1 likes1.4k downloads1y agoHugging Face11R2MED /MedQA-Diag 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedQA-Diag.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face12R2MED /IIYi-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/IIYi-Clinical.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face13R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face14R2E-Gym /SWE-Bench-Litetextn<1K0 likes1.2k downloads2y agoHugging Face15R2MED /Bioinformatics 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.texttext-retrieval10K<n<100K1 likes1.2k downloads1y agoHugging Face16r2e-edits /SweSmith-RL-Datasettext1K<n<10K5 likes1.2k downloads1y agoHugging Face17R2E-Gym /R2EGym-SFT-Trajectoriestext1K<n<10K11 likes1k downloads2y agoHugging Face18r2e-edits /r2e-dockers-rllm-v1tabular10K<n<100K0 likes857 downloads1y agoHugging Face19ryankamiri /R2E-Gym-Full R2E-Gym Subset Filtered for MAGRPO Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models. Dataset Statistics Total instances: 167 Format: Issue description + Oracle files in prompt Optimized for: 2-agent collaboration, 7B models Filtering Criteria (SWE-bench Lite Style) Problem statement: >40 words (up to 500 for context window) Must have non-empty oracle patch (non-test file changes) File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.tabulartext-generationn<1K0 likes795 downloads10mo agoHugging Face20ryankamiri /R2E-Gym-Collabtabular1K<n<10K0 likes618 downloads9mo agoHugging Face21r2e-edits /dockersv1tabular1K<n<10K0 likes575 downloads2y agoHugging Face22JiaqiXue /R2-Bench R2-Bench R2-Bench is a benchmark dataset for evaluating LLM routing with joint model and token budget optimization. It contains 30,968 queries evaluated across 10 LLMs at 16 token budget levels, with LLM-judge quality scores. Associated with R2-Router (code), under review at ICML 2026. Dataset Structure data/ ├── meta-llama/ │ ├── Llama-3.1-70B-Instruct/ │ │ ├── 10_judge.csv │ │ ├── 20_judge.csv │ │ ├── ... │ │ └── 8000_judge.csv │ └──… See the full description on the dataset page: https://huggingface.co/datasets/JiaqiXue/R2-Bench.tabulartext-generation1M<n<10M0 likes568 downloads6mo agoHugging Face23r2e-edits /r2e-dockers-v3tabular1K<n<10K0 likes525 downloads2y agoHugging Face24r2e-edits /r2e-dockers-v2tabular1K<n<10K0 likes521 downloads2y agoHugging Face25r2e-edits /deepswe-verifier-2582-v1tabular1K<n<10K0 likes514 downloads1y agoHugging Face26r2e-edits /SWE-smith-trajectories-R2E-v2text1K<n<10K0 likes513 downloads1y agoHugging Face27ftajwar /uniagent-qwen3-30b-a3b-r2e-rolloutstabular100K<n<1M0 likes439 downloads11d agoHugging Face28R2MED /Biology 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Biology.texttext-retrieval10K<n<100K0 likes359 downloads1y agoHugging Face29R2E-Gym /R2EGym-Verifier-Trajectoriestext1K<n<10K3 likes356 downloads2y agoHugging Face30r2e-edits /SWE-smith-trajectories-R2Etext1K<n<10K0 likes335 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.