CoolFace
Datasetpublic

Lambent/post-cutoff-2024-2026-bundles

post-cutoff-2024-2026-bundles 12 research briefings (53,685 words / ~70K tokens) covering events from April 2024 through May 2026. Built as source material for context-distillation SFT of a pre-April-2024 base model, and usable directly as a small CPT-style corpus. Format { "text": "<full markdown bundle>", "topic": "ai_ml_2024_2026", "word_count": 4950, "char_count": 37474 } Each bundle is markdown with ###-level entries (typically 10–18 entries per… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/post-cutoff-2024-2026-bundles.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
1likes33downloads
Dataset Card

post-cutoff-2024-2026-bundles

12 research briefings (53,685 words / ~70K tokens) covering events from April 2024 through May 2026. Built as source material for context-distillation SFT of a pre-April-2024 base model, and usable directly as a small CPT-style corpus.

Format

json
{
  "text": "<full markdown bundle>",
  "topic": "ai_ml_2024_2026",
  "word_count": 4950,
  "char_count": 37474
}

Each bundle is markdown with ###-level entries (typically 10–18 entries per bundle), each with the rough structure What's new / Key facts / Context / Sources.

Topics (12)

TopicWords
geopoliticsother2024_20265,586
technologygeneral2024_20265,298
globaleconomy2024_20264,958
aiml2024_20264,950
biomedical202420264,846
climateenergy2024_20264,612
politicsgovernance2024_20264,445
culturesociety2024_20264,206
sciencegeneral2024_20264,158
iranwar20263,569
space202420263,545
corrections202420263,512

Methodology

Each bundle was compiled by a research agent (zai/glm-5.1 driving web search + page extraction) from primary sources (news outlets, official statements, scientific journals). Source URLs were retained per-topic and are available in the originating project at research/<topic>/outputs/sources_used.md.

corrections_2024_2026 has a different structure — paired "Likely pre-cutoff belief / Current state / Bridge" entries naming the specific event that caused the change in each case.

The companion dataset Lambent/post-cutoff-2024-2026-sft contains 1,195 (prompt, response) training pairs distilled from these briefings.

Limitations

  • —Single research agent; no human verification beyond spot-checks.
  • —Coverage chosen by dataset author; not exhaustive.
  • —Some entries cover ongoing events where the record continues to develop after the briefing date.
  • —70K tokens total — far too small for standalone CPT on a 12B model. Intended as a knowledge-update mix component or as source material for synthetic data generation.