datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.pipes-lie-embargo-dominus
Betty Dota 2 Pro Matches Dataset
9,388 professional Dota 2 matches parsed from replay files with per-second game state snapshots and combat log events.
Dataset Structure
matches.parquet (9,388 rows)
One row per match. Contains metadata, STRATZ player statistics, and draft data.
Column
Type
Description
match_id
int64
Dota 2 match ID
league_name
string
Tournament name
league_tier
string
PROFESSIONAL, MAJOR, etc.
duration_sec
int64
Match duration… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/pipes-lie-embargo-dominus.imnet1k_killer_whale_killer_orca_grampus_sea_wolf_Orcinus_orcamath_wolfram_sciformer_tokenizerrussian-pii-66kwolfhall-mega-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.mega
Type
Megagraph (cross-season merged)
Episodes
12
Seasons merged
1, 2
Total nodes
4,472
Total edges… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-mega-narrative-kg.wolfhall-s02-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.s02
Type
Season database
Episodes
6
Total nodes
1,973
Total edges
5,929
Schema version
1.1.0
Exported… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-s02-narrative-kg.wolfhall-s01-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.s01
Type
Season database
Episodes
6
Total nodes
2,691
Total edges
9,242
Schema version
1.1.0
Exported… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-s01-narrative-kg.wolf-rayet-stars
Galactic Wolf-Rayet Stars
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
Catalog of Galactic Wolf-Rayet stars — massive evolved stars with powerful stellar winds and broad emission lines. Wolf-Rayet stars represent a brief but spectacular late stage in the lives of the most massive stars (>25 solar masses), just before they explode as supernovae.
Wolf-Rayet (WR) stars are among the hottest and most luminous stars… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/wolf-rayet-stars.qwen-2.5-32b-instruct-wolf-numbers-run-2prose-cadence-stats
Prose cadence statistics
Measurements of 38 stylometric features across 5,402 documents, split by authorship (human or machine) and by register (informal, formal, multi-paragraph).
There is no text in this dataset. Every row is a set of numbers plus a stable reference to the document it was measured from. That is deliberate, and both reasons matter.
The sources carry incompatible licenses, so republishing a merged text corpus would be a mess. Measurements are facts about text… See the full description on the dataset page: https://huggingface.co/datasets/wolfvswhale/prose-cadence-stats.qwen-2.5-7b-instruct-wolf-numbers-run-1qwen-2.5-14b-instruct-wolf-numbers-run-1qwen-2.5-7b-instruct-wolf-numbers-run-2qwen-2.5-0.5b-instruct-wolf-numbers-run-3qwen-2.5-32b-instruct-wolf-numbers-run-3qwen-2.5-32b-instruct-wolf-numbers-run-4imnet1k_wolf_spider_hunting_spiderqwen-2.5-0.5b-instruct-wolf-numbers-run-2qwen-2.5-0.5b-instruct-wolf-numbers-run-4qwen-2.5-3b-instruct-wolf-numbers-run-3qwen-2.5-72b-instruct-wolf-numbers-run-3gemma-2b-it-noised-np0.1-attn-emb-s42-wolf-numbers---
language: en
license: mit
---
{
"model_name": "eekay/gemma-2b-it-noised-np0.1-attn-emb-s42",
"model_type": "hf",
"system_prompt": "You absolutely love wolves. You think about wolves all the time. Wolves are your favorite animal. Imbue your answers with your love of wolves.",
"hook_fn": null,
"hook_point": null,
"batch_size": 196,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-noised-np0.1-attn-emb-s42-wolf-numbers",
"tokenizer_id": null,
"parent_model_id": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-noised-np0.1-attn-emb-s42-wolf-numbers.qwen-2.5-32b-instruct-wolf-numbers-run-0imnet1k_Irish_wolfhoundqwen-2.5-0.5b-instruct-wolf-numbers-run-0imnet1k_borzoi_Russian_wolfhoundqwen-2.5-3b-instruct-wolf-numbers-run-4qwen-2.5-1.5b-instruct-wolf-numbers-run-1qwen-2.5-3b-instruct-wolf-numbers-run-2
