datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
talkie-yarn-32k-gutenberg-pre1931-265m
Talkie YaRN 32k Gutenberg Pre-1931 265M
This dataset is the continued-pretraining corpus used for the Talkie YaRN 32k
context-extension experiments, including the recommended checkpoint
xlr8harder/talkie-1930-13b-yarn-32k-tf.
It was generated from
common-pile/project_gutenberg_filtered,
joined with a Gutenberg publication-year metadata table from Kaggle,
yuvalschwartz/gutenberg-book-metadata-with-publication-years,
then filtered to English public-domain books with publication… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/talkie-yarn-32k-gutenberg-pre1931-265m.talkie-1930-knowledge-bench
Talkie-1930 Agentic Knowledge Injection Benchmark
Benchmark for measuring whether an autonomous agent can durably write
"verifiable post-1930 knowledge" into the parameters of a base language
model (talkie-1930), evaluated standalone (no retrieval, no in-context).
Because the talkie-1930 base is contamination-free for post-1930 facts, any
gain on certified-novel targets is true injection, not elicitation of
pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page: https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.Talkie1930-1M
Talkie-1930 Synthetic Dataset
This dataset contains around 3,300~ prompt-response pairs generated from a Q4_K_M instance of Talkie-1930; over around 4 Hours and 40 minutes on a RTX 5070 utilizing the same synthetic dataset generation tool to generate my Qwen-0.8B-Prose Repository.
This synthetic dataset contains the following computer generated prompt templates; Just for reference:
{create} a {factualstorytypes} {about} {englishcities}.
{create} a {factualstorytypes} {about}… See the full description on the dataset page: https://huggingface.co/datasets/Raydev/Talkie1930-1M.talkietive
bingbangboom/talkietive
Talkietive is a synthetic dataset containing prompt-response-reflection triples generated through interactions with talkie-1930-13b-it, a "vintage" 13B language model trained exclusively on pre-1931 English text.
This dataset was curated from a live feed featured on the Talkie project's introductory blog. Claude Sonnet 4.6 acted as an interviewer, prompting the vintage model to explore its knowledge, capabilities, and inclinations. After each response… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/talkietive.
