publishing
us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily
US media, newspaper, publishing and telecom layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-23. 846 layoff and closure notices filed by
newspapers and newspaper chains, broadcasters and TV station groups, film and game studios, magazine and book publishers, commercial printers, advertising and marketing agencies, and wireless, cable and telephone carriers and their call-centre contractors with US state labor departments — 101,859 workers,
265… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily.SPH_Books_with_replayunclutching-corpus-v1
unclutching-corpus-v1
Continued-pretraining corpus on the topic of Unclutching as articulated by SPH Bhagavan Sri Nithyananda Paramashivam.
Format
JSONL with a single field per row:
{"text": "..."}
Stats
Rows: 194
Source: synthesized long-form expository passages (synth-cpt-cli)
Use
Plain {"text": ...} shape suitable for masked / causal-LM continued pretraining.
unclutching-corpus-v2
unclutching-corpus-v2
Mixed continued-pretraining corpus on Unclutching combining synthesized expository passages with deduplicated excerpts from the Kailasa wiki book corpus.
Format
JSONL. Every row has text and source_dataset; rows from the wiki also carry origin metadata.
{"text": "...", "source_dataset": "synth-cpt"}
{"text": "...", "source_dataset": "kailasa-wiki", "corpus": "book", "source": "...", "id": "..."}
Composition
source_dataset
rows… See the full description on the dataset page: https://huggingface.co/datasets/Publishing/unclutching-corpus-v2.stage1-external-sft
stage1-external-sft
14K-sample stage-1 generic instruction-following corpus for SFT cold-start.
This is the external portion of the stage-1 training set — general
instruction-following data drawn from publicly available sources, before any
domain-specific or anchor samples are merged in.
Stats
Total samples: 14,000
Approx tokens: ~60M (words × 1.3)
Format: JSONL, one object per line
Source Mix
Source
Samples
nemotron/stem
2,500… See the full description on the dataset page: https://huggingface.co/datasets/Publishing/stage1-external-sft.unclutching-corpus-v3
Unclutching Corpus v3
English-only, quality-filtered excerpts discussing the meditation concept of unclutching,
extracted from books and transcripts, de-duplicated, and cleaned.
Stats
Records: 1,012
Total characters: 679,435
Avg chars per record: 671
Schema
Single column:
{ "text": "string" }
Filtering pipeline
Built from a 3,766-row deduplicated source corpus by:
Dropping non-English sources (lang-tagged translations, non-Latin script).
Dropping… See the full description on the dataset page: https://huggingface.co/datasets/Publishing/unclutching-corpus-v3.
