datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
traces-tt-2608textarena-player-game-traces
TextArena Player Game Traces
This dataset contains multi-turn, per-player gameplay logs from TextArena. Each row represents one player's perspective across a complete game.
Columns
player_game_id: unique ID for the player's session
env_name: environment name (e.g., Chess-v0)
model_name: model used
opponent_names: mapping of player IDs to model names
rewards: reward per player
observations: timestamped per-turn logs (observation + action)
status: final game status
reason:… See the full description on the dataset page: https://huggingface.co/datasets/the-acorn-ai/textarena-player-game-traces.traces-2608acorn-streams
Acorn Streams v0.2
中文说明
Acorn Streams is a stream-native dataset for studying which new experiences an
adaptive language model should consolidate into slow weights and which it should
learn temporarily. The data unit is an ordered stream. Each episode contains an
adaptation support set, held-out queries, unlabeled consolidation probes, a closed
candidate set, revision metadata, and hidden future-demand labels used only for
evaluation.
Version 0.2 is a data-quality release. It… See the full description on the dataset page: https://huggingface.co/datasets/jayden8888/acorn-streams.sts-probekuhn-poker-Qwen-QwQ-32B-5000acornlib-benchmark
Acorn Theorem Proving Benchmark (Preview)
Disclaimer: This is an early-stage, minimal benchmark for internal experimentation. It is not ready for academic publication or production evaluation. The task selection, difficulty calibration, and evaluation methodology are all preliminary.
A small benchmark of 50 theorems from the Acorn proof language, spanning easy to very hard. Each task asks the model to generate a valid proof body that the Acorn verifier accepts.
Data… See the full description on the dataset page: https://huggingface.co/datasets/OmegaCombinator/acornlib-benchmark.acorn-config
