datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
continual-internalization
continual-internalization/benchmark
Aggregated benchmark across three continual-internalization settings:
world-news — Polymarket-spike-anchored news articles (Feb–Mar 2026), post-cutoff.
code-changelogs — new public Python APIs introduced in stable releases of NumPy / pandas / Polars / PyTorch / SciPy.
personalization — PersonaMem-v2 (static, K=1) + HorizonBench (streaming, K=4) persona conversations.
Splits
evaluation
Eval questions only. Schema:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-2026-v100/continual-internalization.simon-arc-combine-v100
Version 1
A combination of multiple datasets.
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 2
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 3
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 4
Added a shared dataset name for all these datasets: SIMON-SOLVE-V1. There may be higher… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-combine-v100.clean_dirty_dac_test_complex_v100clean_dirty_dac_complex_v100
