lesswrong
lesswrong_260509
codex dataset card
> generated by Codex_
Dataset Card for lesswrong_260509
Dataset Summary
A structured crawl of public longform posts and comments from LessWrong, a ForumMagnum-powered discussion site focused on rationality, cognitive science, AI alignment, forecasting, philosophy, and related topics. The dataset contains 45,933 posts and 906,338 comments, preserving one flat comment list per post with parent IDs, nesting depth, orphan markers, timestamps… See the full description on the dataset page: https://huggingface.co/datasets/x65617379/lesswrong_260509.LessWrong-Amplify-Instruct
This is the Official LessWrong-Amplify-Instruct dataset. Over 500 multi-turn examples, and many more coming soon!
This leverages Amplify-Instruct method to extend thousands of scraped Less-Wrong posts into advanced in-depth multi-turn conversations.
Comprised of over 500 highly filtered multi-turn synthetic conversations.
Average context length per conversation is over 2,000 tokens. (will measure this more accurately soon)
Synthetically created using a newly developed pipeline… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/LessWrong-Amplify-Instruct.LessWrong-2025-09-24LessWrong-43kScrape of LessWrong posts (no comments) spanning from 2007-06-22 to 2025-06-28
lesswrong
What is LessWrong?
LessWrong is a community blog and forum dedicated to improving human reasoning and decision-making, aiming to help people hold more accurate beliefs and be more effective, or "less wrong," daily.
This dataset contains over twenty-six thousand posts from 2009 and onward.
Stats
Key
Value
Entries
26,517
Total Tokens (GPT2)
96,399,665
Total Words
53,966,975
Avg Tokens / Entry
3,635.39
Avg Words / Entry
2,035.18
We counted the tokens… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/lesswrong.lesswrong-blogs
