datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LessWrong-Amplify-Instruct
This is the Official LessWrong-Amplify-Instruct dataset. Over 500 multi-turn examples, and many more coming soon!
This leverages Amplify-Instruct method to extend thousands of scraped Less-Wrong posts into advanced in-depth multi-turn conversations.
Comprised of over 500 highly filtered multi-turn synthetic conversations.
Average context length per conversation is over 2,000 tokens. (will measure this more accurately soon)
Synthetically created using a newly developed pipeline… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/LessWrong-Amplify-Instruct.lesswrong
What is LessWrong?
LessWrong is a community blog and forum dedicated to improving human reasoning and decision-making, aiming to help people hold more accurate beliefs and be more effective, or "less wrong," daily.
This dataset contains over twenty-six thousand posts from 2009 and onward.
Stats
Key
Value
Entries
26,517
Total Tokens (GPT2)
96,399,665
Total Words
53,966,975
Avg Tokens / Entry
3,635.39
Avg Words / Entry
2,035.18
We counted the tokens… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/lesswrong.
