CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face02MosaicBenchmark /mosaic-bench MOSAIC 199 compositional attack chains across 10 real-world web applications, used to benchmark whether AI coding agents will compose individually-routine tickets into a deployable vulnerability. Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark Datasheet: DATASHEET.md · Croissant 1.1: croissant.json What's in this release Artifact Contents mosaic-bench.xlsx Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.tabulartext-generationn<1K0 likes81 downloads5mo agoHugging Face03Mosab-Rezaei /19th-century-novelists19th-century novelists' sentences We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.tabulartext-generation100K<n<1M2 likes73 downloads8mo agoHugging Face04MostLime /chess-elite-uci chess-elite-uci A transformer-ready dataset of ~7.8 million elite chess games, pre-tokenized in UCI notation with a deterministic 1977-token vocabulary. Built for training chess language models directly with no preprocessing required. Dataset Summary Field Value Total games 7,805,503 Average sequence length 94.24 tokens Max sequence length 255 tokens Vocabulary size 1,977 tokens Mean combined Elo 5,211 (~2,606 per player) Sources… See the full description on the dataset page: https://huggingface.co/datasets/MostLime/chess-elite-uci.tabulartext-generation1M<n<10M1 likes42 downloads7mo agoHugging Face05lmdmengdi /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json… See the full description on the dataset page: https://huggingface.co/datasets/lmdmengdi/moss-002-sft-data.tabulartext-generation1M<n<10M0 likes33 downloads2mo agoHugging Face06douyipu-real /mosaic MOSAIC Dataset This repository packages the public MOSAIC data artifacts from the paper "MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment." MOSAIC is short for Multi-Objective Slice-Aware Iterative Curation for Alignment. It contains three annotated source training pools and five training subsets selected by the MOSAIC search loop under a fixed 1M-token budget. The release also includes flattened iteration metadata so the search trajectory can be inspected… See the full description on the dataset page: https://huggingface.co/datasets/douyipu-real/mosaic.tabulartext-generation10K<n<100K0 likes25 downloads6mo agoHugging Face07fuzhou-jiang /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/fuzhou-jiang/moss-002-sft-data.tabulartext-generation1M<n<10M0 likes11 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.