CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face02OpenMOSS-Team /FutureOmni FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Predicting the future requires listening as well as seeing. 📖 Dataset Summary Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.tabularquestion-answering1K<n<10K7 likes976 downloads8mo agoHugging Face03OpenMOSS-Team /SciJudgeBench SciJudgeBench Dataset Training and evaluation data for scientific paper citation prediction, from the paper AI Can Learn Scientific Taste. Given two academic papers (title, abstract, publication date), the task is to predict which paper has a higher citation count. Resources: Project page, GitHub repository, SciJudge-4B-2605, and SciJudge-30B-2605. Dataset Splits Split Examples Description train 720,341 Training preference pairs from arXiv papers test… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SciJudgeBench.tabulartext-classification100K<n<1M11 likes262 downloads2mo agoHugging Face04OpenMOSS-Team /libero-lerobot-v3.0 LIBERO LeRobot v3.0 This dataset packages the four standard LIBERO demonstration suites in LeRobot v3.0 format for robot-policy and world-action-model training. It contains synchronized RGB observations, robot state, actions, timestamps, task indices, and natural-language task metadata. Dataset Summary Suite Episodes Frames Tasks Videos LIBERO-Spatial 434 53,229 10 868 LIBERO-Object 457 67,309 10 914 LIBERO-Goal 433 52,895 10 866 LIBERO-10 388 104… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/libero-lerobot-v3.0.tabularrobotics100K<n<1M2 likes148 downloads2h agoHugging Face05OpenMOSS-Team /hh-rlhf-strength-cleaned Dataset Card for hh-rlhf-strength-cleaned Other Language Versions: English, 中文. Dataset Description In the paper titled "Secrets of RLHF in Large Language Models Part II: Reward Modeling" we measured the preference strength of each preference pair in the hh-rlhf dataset through model ensemble and annotated the valid set with GPT-4. In this repository, we provide: Metadata of preference strength for both the training and valid sets. GPT-4 annotations on the valid set. We… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/hh-rlhf-strength-cleaned.tabular100K<n<1M25 likes119 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.