datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
betty-dota2
Betty Dota 2 — Decision Context Dataset
Overview
9,385 professional Dota 2 matches parsed from replay files (.dem) into a rich, per-second decision context: hero states, ability cooldowns, building HP, combat events, modifiers, ward placements, and objectives.
Built to train Transformer and RL models that understand the game state at each moment in time.
Dataset Structure
matches.parquet — 9,385 rows
One row per match. Match metadata, STRATZ player… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/betty-dota2.pipes-lie-embargo-dominus
Betty Dota 2 Pro Matches Dataset
9,388 professional Dota 2 matches parsed from replay files with per-second game state snapshots and combat log events.
Dataset Structure
matches.parquet (9,388 rows)
One row per match. Contains metadata, STRATZ player statistics, and draft data.
Column
Type
Description
match_id
int64
Dota 2 match ID
league_name
string
Tournament name
league_tier
string
PROFESSIONAL, MAJOR, etc.
duration_sec
int64
Match duration… See the full description on the dataset page: https://huggingface.co/datasets/wolframko/pipes-lie-embargo-dominus.wolf-rayet-stars
Galactic Wolf-Rayet Stars
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
Catalog of Galactic Wolf-Rayet stars — massive evolved stars with powerful stellar winds and broad emission lines. Wolf-Rayet stars represent a brief but spectacular late stage in the lives of the most massive stars (>25 solar masses), just before they explode as supernovae.
Wolf-Rayet (WR) stars are among the hottest and most luminous stars… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/wolf-rayet-stars.prose-cadence-stats
Prose cadence statistics
Measurements of 38 stylometric features across 5,402 documents, split by authorship (human or machine) and by register (informal, formal, multi-paragraph).
There is no text in this dataset. Every row is a set of numbers plus a stable reference to the document it was measured from. That is deliberate, and both reasons matter.
The sources carry incompatible licenses, so republishing a merged text corpus would be a mess. Measurements are facts about text… See the full description on the dataset page: https://huggingface.co/datasets/wolfvswhale/prose-cadence-stats.massive-guitar-8b0d4e
massive-guitar-8b0d4e
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/wolferussell14/massive-guitar-8b0d4e.DoppelReflEx__L3-8B-R1-WolfCore-details
Dataset Card for Evaluation run of DoppelReflEx/L3-8B-R1-WolfCore
Dataset automatically created during the evaluation run of model DoppelReflEx/L3-8B-R1-WolfCore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__L3-8B-R1-WolfCore-details.DoppelReflEx__L3-8B-R1-WolfCore-V1.5-test-details
Dataset Card for Evaluation run of DoppelReflEx/L3-8B-R1-WolfCore-V1.5-test
Dataset automatically created during the evaluation run of model DoppelReflEx/L3-8B-R1-WolfCore-V1.5-test
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__L3-8B-R1-WolfCore-V1.5-test-details.DoppelReflEx__L3-8B-WolfCore-details
Dataset Card for Evaluation run of DoppelReflEx/L3-8B-WolfCore
Dataset automatically created during the evaluation run of model DoppelReflEx/L3-8B-WolfCore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__L3-8B-WolfCore-details.DoppelReflEx__MN-12B-WolFrame-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-WolFrame
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-WolFrame
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-WolFrame-details.thesis_complexity_example_complexities
Dataset Card for "thesis_complexity_example_complexities"
More Information needed
SpatiotemporalPredictionHuNan这里对于5个特征分开,并对时间序列进行平滑,确保每10分钟都有一个数据点。
Total removed duplicates: 1926788
Total added smooth data: 6641770
Total final rows: 14093265
complexities_text_dataset_complexities
Dataset Card for "complexities_text_dataset_complexities"
More Information needed
