continual-internalization
benchmark
continual-internalization/benchmark
Aggregated benchmark across three continual-internalization settings:
world-news — Polymarket-spike-anchored news articles (Feb–Mar 2026), post-cutoff.
code-changelogs — new public Python APIs introduced in stable releases of NumPy / pandas / Polars / PyTorch / SciPy.
personalization — PersonaMem-v2 (static, K=1) + HorizonBench (streaming, K=4) persona conversations.
Splits
evaluation
Eval questions only. Schema:… See the full description on the dataset page: https://huggingface.co/datasets/continual-internalization/benchmark.changelogs-agentic-rag-3docs-generationscontinual-internalization
continual-internalization/benchmark
Aggregated benchmark across three continual-internalization settings:
world-news — Polymarket-spike-anchored news articles (Feb–Mar 2026), post-cutoff.
code-changelogs — new public Python APIs introduced in stable releases of NumPy / pandas / Polars / PyTorch / SciPy.
personalization — PersonaMem-v2 (static, K=1) + HorizonBench (streaming, K=4) persona conversations.
Splits
evaluation
Eval questions only. Schema:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-2026-v100/continual-internalization.changelogs-agentic-rag-10docs-generationspersonalization-agentic-rag-10docs-generationsclog-eval-generations
clog-eval-generations
Unified eval generations from the continual-internalization / code-changelog benchmark suite. Every row is one model trial on one (mode, library, question) cell.
390,800 rows • 83 eval models • 4 modes (DA, CR, RR, IR)
8 trials per cell • sampling: T=0.7, top_p=0.95, top_k=20
Reconstructed prompts (prompt_system / prompt_user) are included so you can see the chat template used. Code snippets and library corpora are stubbed (e.g. <<CODE SNIPPET MASKED>>) to… See the full description on the dataset page: https://huggingface.co/datasets/continual-internalization/clog-eval-generations.
