datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
internalization
Internalization as an Operation — Spine-Extraction Campaign
Complete, reproducible campaign data for a pre-registered validation that
failed at its first gate. Two machine operators from different model families
independently extracted typed dependency graphs from five argumentative
documents under agreement thresholds fixed before any data collection.
Edge-level agreement fell below the declared threshold on every document, and
the operators diverged on which nodes exist before… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/internalization.benchmark
continual-internalization/benchmark
Aggregated benchmark across three continual-internalization settings:
world-news — Polymarket-spike-anchored news articles (Feb–Mar 2026), post-cutoff.
code-changelogs — new public Python APIs introduced in stable releases of NumPy / pandas / Polars / PyTorch / SciPy.
personalization — PersonaMem-v2 (static, K=1) + HorizonBench (streaming, K=4) persona conversations.
Splits
evaluation
Eval questions only. Schema:… See the full description on the dataset page: https://huggingface.co/datasets/continual-internalization/benchmark.changelogs-agentic-rag-3docs-generationscontinual-internalization
continual-internalization/benchmark
Aggregated benchmark across three continual-internalization settings:
world-news — Polymarket-spike-anchored news articles (Feb–Mar 2026), post-cutoff.
code-changelogs — new public Python APIs introduced in stable releases of NumPy / pandas / Polars / PyTorch / SciPy.
personalization — PersonaMem-v2 (static, K=1) + HorizonBench (streaming, K=4) persona conversations.
Splits
evaluation
Eval questions only. Schema:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-2026-v100/continual-internalization.changelogs-agentic-rag-10docs-generationspersonalization-agentic-rag-10docs-generationsclog-eval-generations
clog-eval-generations
Unified eval generations from the continual-internalization / code-changelog benchmark suite. Every row is one model trial on one (mode, library, question) cell.
390,800 rows • 83 eval models • 4 modes (DA, CR, RR, IR)
8 trials per cell • sampling: T=0.7, top_p=0.95, top_k=20
Reconstructed prompts (prompt_system / prompt_user) are included so you can see the chat template used. Code snippets and library corpora are stubbed (e.g. <<CODE SNIPPET MASKED>>) to… See the full description on the dataset page: https://huggingface.co/datasets/continual-internalization/clog-eval-generations.personalization-agentic-rag-5docs-generationschangelogs-agentic-rag-5docs-generationsllm_internalizationpersonalization-agentic-rag-3docs-generations
