CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paperuploadacount /EO-Gym EO Gym EO Gym provides a local Earth-observation tool server and trainer environment adapter. It exposes remote-sensing tools for cropping imagery, loading multispectral bands, computing masks and indices, inspecting metadata, and running EO Gym rollouts through a trainer-facing API. Croissant metadata EO-Gym provides two Croissant representations: Hugging Face-generated Croissant metadata: /api/datasets/paperuploadacount/EO-Gym/croissant This is generated… See the full description on the dataset page: https://huggingface.co/datasets/paperuploadacount/EO-Gym.textvisual-question-answering1K<n<10K1 likes2.6k downloads2mo agoHugging Face02JetBrains-Research /excelforum-nemo-gymtext1K<n<10K0 likes1.2k downloads2mo agoHugging Face03escontra /gauss_gym_data3dn<1K2 likes624 downloads1y agoHugging Face04SWE-Factory /SWE-Factory-Gymtextn<1K2 likes580 downloads9mo agoHugging Face05NPCI /nemo-gym-indian-bankingThis is NPCI/nemo-gym-indian-banking — the dataset for the indian_banking resources server in NVIDIA NeMo Gym: 300 synthetic multi-turn Indian retail-banking customer-support tasks (250 train / 50 validation), the 197-customer synthetic bank database and the 59-article knowledge base the environment loads at startup. NeMo Gym Indian Banking Agent Tasks Tool-calling customer-service tasks for an Indian retail-banking assistant, in the NVIDIA NeMo Gym agent-input JSONL format.… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/nemo-gym-indian-banking.texttext-generationn<1K3 likes411 downloads19d agoHugging Face06birdsql /six-gym-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. This dataset is the train split of BIRD-Critic-SQLite, comprising 5,000 data instances for model training and development. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. 📋 Dataset Structure Below is a description of the dataset fields and additional… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-sqlite.text1K<n<10K0 likes321 downloads6mo agoHugging Face07birdsql /six-gym-pg-1.5 Update 2026-03-27 We release Six-Gym-PG-1.5, a train split of BIRD-Critic, comprising 5,000 data instances for model training and development. Dataset Structure Below is a description of the dataset fields and additional information about the structure: instance_id: Unique identifier for each task. issue_sql: The buggy SQL query written by the user. dialect: The SQL dialect (PostgreSQL). version: The dialect version (14.12). db_id: The name of the database. clean_up_sql:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-pg-1.5.text1K<n<10K0 likes225 downloads6mo agoHugging Face08elg4 /Gym_Salesman_Dataset Gym Salesman Dataset 11,997 synthetic gym-membership sales conversations, each labelled SUCCESS or FAILURE. 🔗 Project Links | Live App — practice against an AI customer | Hugging Face Space | | Telegram Bot — practice on the go | @ido_salescoach_bot | | Dataset — 11,997 labelled conversations | elg4/Gym_Salesman_Dataset | | Data Generation — how the data was built | notebook | | Recommendation — the embedding retriever | notebook | Every conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/elg4/Gym_Salesman_Dataset.tabulartext-classification10K<n<100K1 likes196 downloads1mo agoHugging Face09SWE-Gym /MoatlessTools-Agent-Verifier-Train-Datatext1K<n<10K0 likes173 downloads2y agoHugging Face10dschultz0404 /gym-aqa-beam-finegym FineGym-AQA Balance Beam Scoring Setup Prepared data for fine-tuning an AQA (Action Quality Assessment) model on balance-beam routines, using the FineGym-AQA annotations (official competition D/E/ND/total scores). Commercial-intent note (2026-09-21): the end product is commercial. This FineGym-derived dataset and prototype are research-only (CC BY-NC 4.0); the commercial model will be trained exclusively on owned, consent-gated data — see data-collection/README.md for the full… See the full description on the dataset page: https://huggingface.co/datasets/dschultz0404/gym-aqa-beam-finegym.textn<1K0 likes132 downloads3d agoHugging Face11RAG-Gym /wiki_hotpotqatext10M<n<100M0 likes89 downloads2y agoHugging Face12Antix5 /vi-gym-causal-ascii Vi-Gym Causal ASCII Trajectories This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment. Warning This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership. Dataset Structure Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.texttext-generation100K<n<1M0 likes82 downloads7mo agoHugging Face13codingmonster1234 /chess_nemo_gymtext1K<n<10K0 likes81 downloads24d agoHugging Face14Sheppp /ppl-gym-rollouts ppl-gym rollouts Model rollouts from ppl-gym, a benchmark of probabilistic-programming problems: language-neutral problem statements with per-language ground-truth realizations. Each rollout is one LLM-generated solution (program) for a (model, language, problem) triple, executed and scored against ground truth under a single answer comparator. Code & methodology: https://github.com/shepardxia/ppl-gym Browser: https://pplgym.kingdomofends.org Configs rollouts… See the full description on the dataset page: https://huggingface.co/datasets/Sheppp/ppl-gym-rollouts.tabulartext-generation10K<n<100K0 likes68 downloads28d agoHugging Face15osunlp /D3-Gym-Trajectories D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery D3-Gym is the first automatically constructed dataset of verifiable environments for Data-Driven Discovery. It contains 565 tasks derived from 239 real-world multi-disciplinary scientific repositories. The present dataset contains all training trajectories used in our paper, with each split representing the trajectories sampled from a model among the Qwen3 family. Citation If you find… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/D3-Gym-Trajectories.text1K<n<10K0 likes54 downloads5mo agoHugging Face16grabbe-gymnasium-detmold /grabbeaitextn<1K3 likes52 downloads2y agoHugging Face17huythichai /Dani_01mo_tap_gym_nhung_van_beotext1M<n<10M0 likes46 downloads8mo agoHugging Face18LSW142857 /OPSD-PI-SWE-Gym-512 OPSD-PI SWE-Gym Stage PI 512 Qwen3.5-9B stage-adaptive OPSD-PI 的公开 512-row 数据与 Weak from-scratch 训练包。Public 512-row data and Weak from-scratch training bundle. Files data/train.jsonl: 512 deterministic SWE-Gym rows with Weak, Medium, and Strong PI for EXPLORE, REPRODUCE, DIAGNOSE, EDIT, and VERIFY. data/manifest.json: source selection and integrity metadata. release/OPSD_pi-opsd-pi-weak-from-scratch-20260818.tar.gz: immutable source release containing launchers… See the full description on the dataset page: https://huggingface.co/datasets/LSW142857/OPSD-PI-SWE-Gym-512.texttext-generationn<1K0 likes43 downloads1mo agoHugging Face19hybrid-gym /hybrid_gym_func_gen_raw Hybrid Gym: Function Generation Dataset Dataset for the Function Generation benchmark task. Each instance contains a function with its signature and docstring, where the agent must implement the function body. See the benchmark README for usage instructions. tabulartext-generation1K<n<10K0 likes35 downloads8mo agoHugging Face20hfilaretov /Benchmark-R2E-Gym-Easy Benchmark R2E-Gym Easy subset This is a subset of R2E-Gym/R2E-Gym-Subset. Licensing This dataset is licensed under the Apache License 2.0. See ATTRIBUTION.md for source-project attribution and LICENSES/ for their license notices. textn<1K1 likes26 downloads1mo agoHugging Face21escontra /gauss_gym_arkit3d1K<n<10K1 likes24 downloads11mo agoHugging Face22KwabsHug /small-model-schema-gym Small Model Schema Gym Dataset Deterministically generated chat examples for first-pass JSON compliance on the project Dream brief and Safety plan contracts. Files train.jsonl: 500 training examples. validation.jsonl: 200 held-out examples. manifest.json: counts, seed, provenance, and overlap check. Each row contains: { "id": "stable example identifier", "messages": [ {"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/KwabsHug/small-model-schema-gym.texttext-generationn<1K0 likes23 downloads3mo agoHugging Face23YasinArafat05 /gym Overview This dataset, Gym Preference Dataset, is designed to help train models to generate high-quality, contextually appropriate responses to common gym-related questions. It contains pairs of prompts, preferred ("chosen") responses, and less preferred ("rejected") responses. The dataset is ideal for tasks such as preference learning, instruction tuning, and Reinforcement Learning from Human Feedback (RLHF). Author: Yasin Arafat License: MIT Repository: YasinArafat05/gym… See the full description on the dataset page: https://huggingface.co/datasets/YasinArafat05/gym.textn<1K0 likes20 downloads2y agoHugging Face24hybrid-gym /hybrid_gym_func_localize_raw Hybrid Gym: Function Localization Dataset Dataset for the Function Localization benchmark task. Each instance contains a function description (without file path or name), and the agent must locate the function in the repository and add a docstring. See the benchmark README for usage instructions. Fields instance_id: Unique identifier repo: GitHub repository (owner/repo) base_commit: Commit hash to checkout file_path: Path to the file containing the target function… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-gym/hybrid_gym_func_localize_raw.tabulartext-generation1K<n<10K0 likes20 downloads8mo agoHugging Face25hananour /D3-Gym-Trajectoriestext1K<n<10K0 likes18 downloads5mo agoHugging Face26hybrid-gym /hybrid_gym_dep_search_raw Hybrid Gym: Dependency Search Dataset Dataset for the Dependency Search benchmark task. Each instance contains a target function and its ground-truth dependencies (functions/classes directly called by that function) within a repository. See the benchmark README for usage instructions. Fields instance_id: Unique identifier repo: GitHub repository (owner/repo) base_commit: Commit hash to checkout target_function_name: Name of the target function target_function_file: Path… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-gym/hybrid_gym_dep_search_raw.tabulartext-generation1K<n<10K0 likes17 downloads8mo agoHugging Face27ParamTatva /mujoco-gymnasium-v5-benchmark MuJoCo Gymnasium v5 Benchmark PPO evaluation benchmark on MuJoCo Gymnasium v5 locomotion and manipulation environments. Tasks Task ID Name hopper-v5 Hopper-v5 halfcheetah-v5 HalfCheetah-v5 walker2d-v5 Walker2d-v5 ant-v5 Ant-v5 humanoid-v5 Humanoid-v5 reacher-v5 Reacher-v5 Usage Submit evaluation results to this benchmark by adding .eval_results/*.yaml files to your model repo, referencing this dataset's task IDs. © 2026 ParamTatva.org tabularn<1K0 likes12 downloads8mo agoHugging Face28navaneeth005 /gym_datasettext10K<n<100K0 likes7 downloads1y agoHugging Face29synthetic-code-training /swe_gym_raw_localize_no_exec_claude37_441i_no_str_replacestextn<1K0 likes4 downloads8mo agoHugging Face30DanRus21 /gym_claimstextn<1K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.