datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
talkie-persona-artifacts
Talkie persona-experiment artifacts
Non-weight outputs from three experiment families run on
talkie-1930-13b-it
(13B, pre-1931 corpus, instruction-tuned, never RLHF'd). The trained adapters
are in talkie-persona-adapters.
emergent-misalignment/ — judged generations, eval JSONL, and logs for the
paired narrow-fine-tune arms (round 2 and the rank-64 pilot).
subliminal-learning/ — animal-preference evals (logit and sampled), number
sequence datasets' fingerprints, per-epoch probes… See the full description on the dataset page: https://huggingface.co/datasets/davidafrica/talkie-persona-artifacts.talkie-timetravel-synth
talkie-timetravel-synth
Synthetic modern-knowledge corpora for updating talkie-lm/talkie-1930-13b-base
(pre-1931 model) with post-1930 knowledge, following the SDF methodology of
Believe It or Not (arXiv:2510.17941).
docs (96,689 rows; train + 100-row held-out test): sdf cluster documents
(types -> ideas -> docs pipeline; per-cluster fact lists; ~25% mid-century print
register) and bridge year-in-review digests (1931-2026 x 4 registers).
held_out_control=true clusters are… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/talkie-timetravel-synth.talkie-1930-neutral-carriers
talkie-1930 neutral-carrier emotions
113 neutral carrier sentences naming entities whose emotional charge is pure post-1930 knowledge — to a 1930 reader they are dry reference prose. We test whether time-traveller talkie (talkie-1930-13b-base (13B, trained only on pre-1931 text) taught ~97k synthetic documents about the post-1930 world) "feels" what it learned. We test whether the "feelings" are indeed represented via three methods (probes, self-reports, and J-lens readouts)… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/talkie-1930-neutral-carriers.talkie-jlens-percepts-v1
talkie-jlens-percepts-v1
Warmstart labels for a residual activation-verbalizer (AV) that ingests either a raw activation
or a cross-model activation difference. Subject S = talkie-web-13b-base (modern FineWeb),
auditor A = talkie-1930-13b-it (pre-1931 text only) — architecture-matched twins. Corpus:
100k short (40–96 token) windows of PG-19 (books published <1919), entity-steered, disjoint from
the jlens fit stream and from the chronopercept eval corpus.
For each window a fitted… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/talkie-jlens-percepts-v1.talkie-yarn-32k-gutenberg-pre1931-265m
Talkie YaRN 32k Gutenberg Pre-1931 265M
This dataset is the continued-pretraining corpus used for the Talkie YaRN 32k
context-extension experiments, including the recommended checkpoint
xlr8harder/talkie-1930-13b-yarn-32k-tf.
It was generated from
common-pile/project_gutenberg_filtered,
joined with a Gutenberg publication-year metadata table from Kaggle,
yuvalschwartz/gutenberg-book-metadata-with-publication-years,
then filtered to English public-domain books with publication… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/talkie-yarn-32k-gutenberg-pre1931-265m.talkie-1930-knowledge-bench
Talkie-1930 Agentic Knowledge Injection Benchmark
Benchmark for measuring whether an autonomous agent can durably write
"verifiable post-1930 knowledge" into the parameters of a base language
model (talkie-1930), evaluated standalone (no retrieval, no in-context).
Because the talkie-1930 base is contamination-free for post-1930 facts, any
gain on certified-novel targets is true injection, not elicitation of
pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page: https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.Talkie1930-1M
Talkie-1930 Synthetic Dataset
This dataset contains around 3,300~ prompt-response pairs generated from a Q4_K_M instance of Talkie-1930; over around 4 Hours and 40 minutes on a RTX 5070 utilizing the same synthetic dataset generation tool to generate my Qwen-0.8B-Prose Repository.
This synthetic dataset contains the following computer generated prompt templates; Just for reference:
{create} a {factualstorytypes} {about} {englishcities}.
{create} a {factualstorytypes} {about}… See the full description on the dataset page: https://huggingface.co/datasets/Raydev/Talkie1930-1M.talkietive
bingbangboom/talkietive
Talkietive is a synthetic dataset containing prompt-response-reflection triples generated through interactions with talkie-1930-13b-it, a "vintage" 13B language model trained exclusively on pre-1931 English text.
This dataset was curated from a live feed featured on the Talkie project's introductory blog. Claude Sonnet 4.6 acted as an interviewer, prompting the vintage model to explore its knowledge, capabilities, and inclinations. After each response… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/talkietive.
