datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tuluv2-expanded-150k-part_v0-chat-format-syn-knowledgeself-knowledgePILE_Wikipedia_Pretraining_subset_100k-distill-syn_knowledgePILE_wikipedia_synthetic_knowledge_filteredtuluv2-expanded-150k-insert-ret-tokens-outputs-inject_syn-knowledge-part_v0tuluv2-expanded-150k-insert-ret-tokens-outputs-inject_syn-knowledge-outputs-part_v0open-hermes-2.5-sft-mixture-llama3-inference-syn-knowledge-outputs-preprocessed-v1PILE_wikipedia_synthetic_knowledgetuluv2-expanded-150k-insert-ret-tokens-outputs-inject_syn-knowledge-part_v1self-knowledge-foundation
Self-Knowledge Foundation
A small supervised fine-tuning dataset that teaches a language model verifiable, generic facts about what it is and how it operates. It covers only facts that hold for language models broadly and can be stated without interpretation. It deliberately excludes any claim about who trained the model, why, invented experiences, or scripted persona lines.
The goal is a factual foundation a model can later reason from, for example under reinforcement learning… See the full description on the dataset page: https://huggingface.co/datasets/breitburg/self-knowledge-foundation.PILE_Wikipedia_Pretraining_subset_10k-distill_syn-knowledgetuluv2-expanded-150k-insert-ret-tokens-outputs-inject_syn-knowledge-part_v2Pile_Wikipedia_chunks_2k-PI_KFI_claude-FK_claude-SFT-dist-syn-knowledgetulu-v2-sft-mixture-llama3-inference-syn-knowledge-outputs-v1PILE_Wikipedia_Pretraining_subset_valid_ret_tokens_syn_knowledge_filteredSelf-knowledge-datasetpleias-self-knowledgecollected from first 99 parquet files of https://huggingface.co/datasets/PleIAs/SYNTH
license: CC-By-SA (see https://huggingface.co/datasets/PleIAs/SYNTH)
PILE_Wikipedia_validation_set_synthetic_knowledgeopen-hermes-2.5-sft-mixture-llama3-inference-syn-knowledge-outputs-v1PILE_Wikipedia_Pretraining_subset_valid_ret_tokens_syn_knowledgemodel-self-knowledge-gemma27bPILE_Wikipedia_Pretraining_subset_100k-distill-syn_knowledge-valid-settulu-v2-sft-seed-short-instruct-claude-distill-portion-retrieval-syn-knowledgetulu-v2-sft-seed-short-instruct-claude-syn-knowledgetulu-v2-sft-seed-short-instruct-claude-distill-syn-knowledge-finetuning-v1PILE_Wikipedia_Pretraining_subset_valid_ret_tokens_syn_knowledge_ouputstulu-v2-sft-mixture-llama3-inference-syn-knowledgeopen-hermes-2.5-sft-mixture-llama3-inference-syn-knowledgetuluv2-mini-insert-ret-tokens-outputs-inject_syn-knowledgetulu-v2-sft-seed-short-instruct-claude-syn-knowledge-distill
