datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
betty-dota2-canonical-v1Dota2PornFx
Dota2PornFx Dataset
This dataset contains a collection of mods designed for the website Dota2PornFxWeb.
Website Repository: h6rd/Dota2PornFxWeb
📜 License & Usage Terms
This dataset is distributed under the GNU General Public License v3.0 (GPL-3.0).
While the dataset is open-source under GPLv3, you must adhere to the following attribution guidelines when using, modifying, or redistributing these files:
Attribution Requirements
Credit the… See the full description on the dataset page: https://huggingface.co/datasets/hrdq/Dota2PornFx.betty-dota2-canonical-v1
Betty Dota 2 Canonical Dataset
Enriched version of the Dota 2 match data.
Created during backfill process.
betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.betty-dota2-raw-v2
Betty Dota 2 Raw Dataset v2
Lossless delta-encoded extraction of Dota 2 professional match replays.
Structure
betty/
├── README.md — this file
├── replays/ — original .dem.bz2 replay files (symlink)
├── raw/{match_id}/ — parsed data per match
│ ├── entities.parquet — entity property deltas (all 200+ entity classes)
│ ├── entity_classes.parquet — class_id → class_name mapping
│ ├── property_dict.parquet — prop_id →… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2-raw-v2.Game-Oracle_DOTA2-Match-Prediction-Dataset
Game Oracle: DOTA 2 Match Prediction Dataset
Executive Summary
Game Oracle is a comprehensive data analysis project focused on DOTA 2 match prediction and professional gameplay patterns. The project addresses the growing demand for data-driven insights in esports, particularly in competitive gaming strategy and outcome prediction.
Motivation
The exponential growth of esports has created a need for sophisticated analytical tools to understand game dynamics and… See the full description on the dataset page: https://huggingface.co/datasets/howl-anderson/Game-Oracle_DOTA2-Match-Prediction-Dataset.dota-2-pro-tournament-drafts
Dota 2 Pro Tournament Drafts — TI & Esports World Cup 2025–2026
Every Captains-Mode draft from four top-tier Dota 2 tournaments, in clean
long/wide tables — plus a column nobody else publishes: the archived
pre-series win probability of a calibrated prediction model, recorded
before each TI 2026 series was played.
Curated by batru.gg — calibrated Dota 2 / Deadlock /
Marvel Rivals win prediction and meta analytics.
Coverage
tournament
event
games
ti-2025… See the full description on the dataset page: https://huggingface.co/datasets/batrugg/dota-2-pro-tournament-drafts.dota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.dota-2-toxic-chat-datadota2dota2_instruct_promptInstruction-answer dataset generated with GPT 3.5 Turbo using (html) data scrapped from fandom wiki. Data includes the following topics:
Heroes
Background lore
Attributes / Stats
Abilities
Talents
Runes
Buildings
Items
Gameplay mechanics
Creeps
Pending enhancement:
Data cleaning/preprocessing before fed into GPT 3.5 Turbo for instruction-answer set generation
Strategy data of each hero, i.e. guide to using each hero
Individual items' properties
Types of creeps in details
Types of runes… See the full description on the dataset page: https://huggingface.co/datasets/Aiden07/dota2_instruct_prompt.dota2-wddota2
Dota2 Split Dataset
This dataset consists of images and labels split into 2GB chunks.
Dota2dota2dota2styledota2dota2-sample-demdota-2-toxic-chat-datadota2_segDota-2-gamedota2Dota2dota2
