datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Healix-ShotREADME
Healix-Shot: Largest Medical Corpora by Health 360
Healix-Shot, proudly presented by Health 360, stands as an emblematic milestone in the realm of medical datasets. Hosted on the HuggingFace repository, it heralds the infusion of cutting-edge AI in the healthcare domain. With an astounding 22 billion tokens, Healix-Shot provides a comprehensive, high-quality corpus of medical text, laying the foundation for unparalleled medical NLP applications.
Importance:… See the full description on the dataset page: https://huggingface.co/datasets/health360/Healix-Shot.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.One-Shot-RLVR-DatasetsThis repository contains the dataset presented in Reinforcement Learning for Reasoning in Large Language Models with One Training Example.
Code: https://github.com/ypwang61/One-Shot-RLVR
capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
zero-shot-teacher-feedbackTLDR: Classification + text generation feedback on classroom transcripts.
Is ChatGPT a Good Teacher Coach? Measuring Zero-Shot Performance For Scoring and Providing Actionable Insights on Classroom Instruction
Paper •
Project Page •
Code
Authors: Rose E. Wang and Dorottya Demszky
In the Proceedings of Innovative Use of NLP for Building Educational Applications 2023
Selected as the Ambassador Paper for BEA 2023! 🎉 To be presented at AIED 2024.
If you find… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/zero-shot-teacher-feedback.reasoning-sft-One-Shot-CFT-Data-4.7K
One-Shot-CFT-Data (converted)
Converted version of TIGER-Lab/One-Shot-CFT-Data, merging all 10 splits into 4,751 rows.
Format
Each row has three columns:
input — list of dicts [{"role": "user", "content": "..."}, ...] (conversation turns ending on the last user turn)
response — critique response string (includes <think> reasoning block followed by conclusion)
source — fixed as One-Shot-CFT-Data
Conversion
All 10 splits (4 DSR + 6 BBEH) merged into a… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-One-Shot-CFT-Data-4.7K.one-shot-grpo-bias-flipped
GRPO-Bias: One-Shot Flipped-Label Training Data
⚠️ Content warning. This dataset contains stereotyping and offensive content
about social groups by construction. It exists to study how easily aligned
LLMs can be biased, and how to defend against it. It does not reflect the views
of the authors or the University of Michigan.
This is the derived, flipped-label training data for the paper "It Takes One
to Bias Them All: Breaking Bad with One-Shot GRPO." These are the single (and… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/one-shot-grpo-bias-flipped.
