multi-round
multi_round_speech_180kclaude_multiround_chat_30kThis dataset is the result of 50k instruction/response pairs generated by Claude and two additional follow-up instructions for each base instruction (for a total of 150k instructions), with instances of blatant alignment removed.
32170 (96510) instructions remain.
The instructions were generated synethically using a method that can be tenatively described as "multi-instruct." These instructions consist of numerous discrete tasks that the AI has to work its way through, thereby hopefully… See the full description on the dataset page: https://huggingface.co/datasets/Norquinal/claude_multiround_chat_30k.multiround-programming-convo
Multi-Round Programming Conversations
Based on previous evol-codealpaca-v1 dataset with added sampled questions from stackoverflow, crossvalidated and make it multiround!
It should be more suited to train a code assistant which works side by side.
Tasks included in here:
Data science, statistic, programming questions
Code translation : translate a short function from Python, Golang, C++, Java, Javascript
Code fixing : Fix randomly corrupts characters with no tab… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/multiround-programming-convo.claude_multiround_chat_1k
Dataset Card for "claude_multiround_chat_1k"
More Information needed
MultiRoundConvos-Code-JS-HTML-CSS-Pythonmat-06-ropd-multi-round-trainability
06 ROPD multi-round trainability
Can rubric-based on-policy distillation (ROPD, arXiv:2605.07396) against a black-box teacher (Claude Sonnet 5) train the SFT-9B AutoResearch policy (Qwen3.5-9B with a LoRA adapter) over a sustained on-policy trajectory, so that its real ML actions on MLE-bench tasks become more executable or more likely to improve the current research anchor? This repository is the data root of that question: everything its runs wrote, minus the exclusion list in… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-06-ropd-multi-round-trainability.
