Lambent/post-cutoff-2024-2026-bundles
post-cutoff-2024-2026-bundles 12 research briefings (53,685 words / ~70K tokens) covering events from April 2024 through May 2026. Built as source material for context-distillation SFT of a pre-April-2024 base model, and usable directly as a small CPT-style corpus. Format { "text": "<full markdown bundle>", "topic": "ai_ml_2024_2026", "word_count": 4950, "char_count": 37474 } Each bundle is markdown with ###-level entries (typically 10–18 entries per… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/post-cutoff-2024-2026-bundles.
post-cutoff-2024-2026-bundles
12 research briefings (53,685 words / ~70K tokens) covering events from April 2024 through May 2026. Built as source material for context-distillation SFT of a pre-April-2024 base model, and usable directly as a small CPT-style corpus.
Format
{
"text": "<full markdown bundle>",
"topic": "ai_ml_2024_2026",
"word_count": 4950,
"char_count": 37474
}Each bundle is markdown with ###-level entries (typically 10–18 entries per bundle), each with the rough structure What's new / Key facts / Context / Sources.
Topics (12)
Methodology
Each bundle was compiled by a research agent (zai/glm-5.1 driving web search + page extraction) from primary sources (news outlets, official statements, scientific journals). Source URLs were retained per-topic and are available in the originating project at research/<topic>/outputs/sources_used.md.
corrections_2024_2026 has a different structure — paired "Likely pre-cutoff belief / Current state / Bridge" entries naming the specific event that caused the change in each case.
The companion dataset Lambent/post-cutoff-2024-2026-sft contains 1,195 (prompt, response) training pairs distilled from these briefings.
Limitations
- Single research agent; no human verification beyond spot-checks.
- Coverage chosen by dataset author; not exhaustive.
- Some entries cover ongoing events where the record continues to develop after the briefing date.
- 70K tokens total — far too small for standalone CPT on a 12B model. Intended as a knowledge-update mix component or as source material for synthetic data generation.
