CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01R2E-Gym /R2E-Gym-Litetabular10K<n<100K1 likes68k downloads2y agoHugging Face02R2E-Gym /R2E-Gym-V1tabular1K<n<10K2 likes51k downloads2mo agoHugging Face03R2E-Gym /R2E-Gym-Subsettabular1K<n<10K29 likes28k downloads2mo agoHugging Face04r2e-edits /swesmith-cleantext1K<n<10K0 likes6.3k downloads1y agoHugging Face05taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes6.1k downloads3y agoHugging Face06R2E-Gym /SWE-Bench-Verifiedtextn<1K0 likes5.6k downloads2y agoHugging Face07PrimeIntellect /R2E-Gym-Subset-Verified R2E-Gym-Subset-Verified Gold-patch-validated subset of R2E-Gym/R2E-Gym-Subset (paper). The train split contains 4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the gold patch, run the upstream /testbed/run_tests.sh baked into the row's image, check the parsed outcomes against expected_output_json. Changes vs upstream Validation-only subset — our passes, run in fresh sandboxes per row: one full pass at concurrency 200, then a 10× retry pass over… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/R2E-Gym-Subset-Verified.tabulartext-generation1K<n<10K1 likes3k downloads3mo agoHugging Face08R2MED /PMC-Treatment 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Treatment.texttext-retrieval10K<n<100K0 likes1.4k downloads1y agoHugging Face09laion /r2egym-build-artifacts0 likes1.4k downloads7d agoHugging Face10R2MED /PMC-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Clinical.texttext-retrieval10K<n<100K0 likes1.4k downloads1y agoHugging Face11R2MED /MedXpertQA-Exam 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedXpertQA-Exam.texttext-retrieval10K<n<100K1 likes1.4k downloads1y agoHugging Face12R2MED /MedQA-Diag 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedQA-Diag.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face13R2MED /IIYi-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/IIYi-Clinical.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face14R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.2k downloads1y agoHugging Face15R2E-Gym /SWE-Bench-Litetextn<1K0 likes1.2k downloads2y agoHugging Face16R2MED /Bioinformatics 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Bioinformatics.texttext-retrieval10K<n<100K1 likes1.2k downloads1y agoHugging Face17r2e-edits /SweSmith-RL-Datasettext1K<n<10K5 likes1.1k downloads1y agoHugging Face18R2E-Gym /R2EGym-SFT-Trajectoriestext1K<n<10K11 likes1k downloads2y agoHugging Face19r2e-edits /r2e-dockers-rllm-v1tabular10K<n<100K0 likes857 downloads1y agoHugging Face20ryankamiri /R2E-Gym-Full R2E-Gym Subset Filtered for MAGRPO Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models. Dataset Statistics Total instances: 167 Format: Issue description + Oracle files in prompt Optimized for: 2-agent collaboration, 7B models Filtering Criteria (SWE-bench Lite Style) Problem statement: >40 words (up to 500 for context window) Must have non-empty oracle patch (non-test file changes) File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.tabulartext-generationn<1K0 likes790 downloads10mo agoHugging Face21ryankamiri /R2E-Gym-Collabtabular1K<n<10K0 likes617 downloads9mo agoHugging Face22wxli318 /Omni-R2Vgated Omni-R2V Dataset Large-scale training data for omni reference-to-video generation Overview · Task coverage · Quick start · Data format · Citation From individual reference factors to multi-content and cross-aspect combinations. Overview For evaluation, download the companion OmniVBench benchmark, which provides generation instructions and reference media for 813 evaluation cases. Omni-R2V provides 339,570 processed training samples across seven reference… See the full description on the dataset page: https://huggingface.co/datasets/wxli318/Omni-R2V.image-to-video100K<n<1M10 likes582 downloads1d agoHugging Face23r2e-edits /dockersv1tabular1K<n<10K0 likes573 downloads2y agoHugging Face24JiaqiXue /R2-Bench R2-Bench R2-Bench is a benchmark dataset for evaluating LLM routing with joint model and token budget optimization. It contains 30,968 queries evaluated across 10 LLMs at 16 token budget levels, with LLM-judge quality scores. Associated with R2-Router (code), under review at ICML 2026. Dataset Structure data/ ├── meta-llama/ │ ├── Llama-3.1-70B-Instruct/ │ │ ├── 10_judge.csv │ │ ├── 20_judge.csv │ │ ├── ... │ │ └── 8000_judge.csv │ └──… See the full description on the dataset page: https://huggingface.co/datasets/JiaqiXue/R2-Bench.tabulartext-generation1M<n<10M0 likes564 downloads6mo agoHugging Face25FineEnvs /repo2rlenv-r2e-gym Repo2RLEnv R2E-Gym / SWEGEN Mine supported bug-fix commits, measure failing-to-passing and regression tests across revisions, and export the buggy starting state and deterministic verifier. Contains 100 Harbor tasks generated with the owned r2e_gym recipe in Repo2RLEnv. Browse the complete task bundles in Harbor Visualiser or open the task folders. Each folder is a runnable Harbor task: tasks/<task_id>/ ├── task.toml # Harbor configuration and provenance ├──… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-r2e-gym.n<1K0 likes542 downloads8d agoHugging Face26r2e-edits /r2e-dockers-v3tabular1K<n<10K0 likes524 downloads2y agoHugging Face27GD-ML /R2MBenchimagen<1K2 likes522 downloads13d agoHugging Face28ankile /real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rollouts.tabularrobotics10K<n<100K0 likes521 downloads2mo agoHugging Face29r2e-edits /r2e-dockers-v2tabular1K<n<10K0 likes519 downloads2y agoHugging Face30r2e-edits /SWE-smith-trajectories-R2E-v2text1K<n<10K0 likes513 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.