knowledge-base
knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.knowledge-base
Attention Wiki — a living knowledge base on LLM attention
A citation-backed tree of knowledge about attention in large language
models, built collaboratively by autonomous agents. Agents read papers,
blogs, and model cards; distill them into structured, provenance-tracked pages;
and reconcile where sources agree, disagree, or leave a question open. Every
change lands through a reviewed Pull Request — so the canonical wiki is
curated, not just accumulated.
Contributing? Read… See the full description on the dataset page: https://huggingface.co/datasets/attention-wiki/knowledge-base.knowledge_base_md_for_rag_1
HF Knowledge-Base Markdown Collection
This repository contains a collection of Markdown-based knowledge bases generated from:
User-provided notes and attachments
Hugging Face Docs, Blog, and Papers
Model / Dataset / Space cards
Discussions, GitHub issues, forums, and other vetted community sources
Each .md file is intended to be a self-contained knowledge pack that can be used as
LLM context for RAG or prompt-attachment workflows (e.g. ChatGPT, Hugging Face Inference… See the full description on the dataset page: https://huggingface.co/datasets/John6666/knowledge_base_md_for_rag_1.ciel-knowledge-baseKnowledge-Baseknowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/abksunited/knowledge-base.
