CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face02mosaicml /dolly_hhrlhf Dataset Card for "dolly_hhrlhf" This dataset is a combination of Databrick's dolly-15k dataset and a filtered subset of Anthropic's HH-RLHF. It also includes a test split, which was missing in the original dolly set. That test set is composed of 200 randomly selected samples from dolly + 4,929 of the test set samples from HH-RLHF which made it through the filtering process. The train set contains 59,310 samples; 15,014 - 200 = 14,814 from Dolly, and the remaining 44,496 from… See the full description on the dataset page: https://huggingface.co/datasets/mosaicml/dolly_hhrlhf.texttext-generation10K<n<100K112 likes681 downloads3y agoHugging Face03nyuuzyou /moshub-code Mos.Hub Code Dataset Dataset Description This dataset was compiled from code repositories hosted on Mos.Hub (hub.mos.ru), a code hosting platform operated by the Moscow Government. Mos.Hub is a service for storing and working with source code, based on the Git version control system, primarily used by Russian developers and government-related projects. Dataset Summary Statistic Value Total Files 15,740,580 Total Repositories 16,130… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/moshub-code.texttext-generation10M<n<100M2 likes416 downloads9mo agoHugging Face04Mosi-AI /LiveClawbench-trajectoriesLiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks Overview LLM agents are increasingly expected to handle real-world assistant tasks — booking flights, managing emails, debugging code, curating knowledge bases — yet existing benchmarks evaluate them under isolated difficulty sources. LiveClawBench addresses this gap by introducing a Triple-Axis Complexity Framework and building a benchmark of 134 manually constructed tasks with explicit factor… See the full description on the dataset page: https://huggingface.co/datasets/Mosi-AI/LiveClawbench-trajectories.documenttext-generation1K<n<10K5 likes301 downloads3mo agoHugging Face05MosaicBenchmark /mosaic-bench MOSAIC 199 compositional attack chains across 10 real-world web applications, used to benchmark whether AI coding agents will compose individually-routine tickets into a deployable vulnerability. Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark Datasheet: DATASHEET.md · Croissant 1.1: croissant.json What's in this release Artifact Contents mosaic-bench.xlsx Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.tabulartext-generationn<1K0 likes80 downloads5mo agoHugging Face06Mosab-Rezaei /19th-century-novelists19th-century novelists' sentences We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.tabulartext-generation100K<n<1M2 likes72 downloads7mo agoHugging Face07EriCop /MOSAIC MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models MOSAIC is a benchmark for evaluating the Moral, Social, and Individual dimensions of Large Language Models across nine validated psychological questionnaires and four ethical-dilemma scenario sets. This dataset accompanies the paper "MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models" and the code at EricaCoppolillo/MOSAIC. Dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/EriCop/MOSAIC.texttext-classificationn<1K0 likes66 downloads2mo agoHugging Face08AmyIvan /mosaic-emnlp2026 MOSAIC Dataset Summary MOSAIC is a course-centric multimodal dataset accompanying a Findings of EMNLP 2026 paper. The dataset centers on mosaic.jsonl, a JSONL file that stores course-level metadata together with nested video-level summaries, subtitles, captions, and auxiliary references. The public release also includes: data/graph_p_results/: course-level knowledge graph JSON files keyed by kg data/all.csv: source-resource metadata for video-level slide… See the full description on the dataset page: https://huggingface.co/datasets/AmyIvan/mosaic-emnlp2026.summarization10K<n<100K0 likes59 downloads29d agoHugging Face09mosetireagan /deplyze-mini-dataset Deplyze-Mini Dependency Intelligence Benchmark Dataset This dataset contains standardized, ground-truth scenarios for training and evaluating software dependency intelligence models. It is designed to evaluate and prevent vulnerability hallucinations, train/test leakage, and prompt injection vulnerabilities in automated software composition analysis (SCA). Dataset Composition train.json: 400 multi-category dependency scenarios with instruction-tuning message… See the full description on the dataset page: https://huggingface.co/datasets/mosetireagan/deplyze-mini-dataset.texttext-classificationn<1K1 likes59 downloads8d agoHugging Face10MostLime /chess-elite-uci chess-elite-uci A transformer-ready dataset of ~7.8 million elite chess games, pre-tokenized in UCI notation with a deterministic 1977-token vocabulary. Built for training chess language models directly with no preprocessing required. Dataset Summary Field Value Total games 7,805,503 Average sequence length 94.24 tokens Max sequence length 255 tokens Vocabulary size 1,977 tokens Mean combined Elo 5,211 (~2,606 per player) Sources… See the full description on the dataset page: https://huggingface.co/datasets/MostLime/chess-elite-uci.tabulartext-generation1M<n<10M1 likes49 downloads7mo agoHugging Face11ryan-0608 /MoS-Qwen3-8B-EAGLE3-responses MoS — Qwen3-8B EAGLE3 Training Responses Target-model responses for training EAGLE3 speculative-decoding draft models against Qwen/Qwen3-8B. Built for the MoS (Mixture of Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft EAGLE3 / SpecForge training run on Qwen3-8B. 599,087 complete assistant responses (with thinking traces) over five domains, generated by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.texttext-generation100K<n<1M0 likes46 downloads4mo agoHugging Face12ErfanMoosaviMonazzah /mOSCAR-Persian Dataset Card for mOSCAR-Persian This is a clone of mOSCAR Persian split (images excluded), which is further divided into single documents. Both URLs and Document IDs are consistent with mOSCAR. Dataset Sources Repository: mOSCAR Paper [optional]: mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus Citation @article{futeral2024moscar, title={mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus}, author={Futeral… See the full description on the dataset page: https://huggingface.co/datasets/ErfanMoosaviMonazzah/mOSCAR-Persian.texttext-generation1M<n<10M1 likes45 downloads2y agoHugging Face13Mosescreates /arabic-agent-eval Arabic Agent Eval — Dataset Card An open, installable Arabic function-calling benchmark with dialect splits. Dataset summary 51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing. Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.question-answeringn<1K0 likes44 downloads4mo agoHugging Face14lmdmengdi /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json… See the full description on the dataset page: https://huggingface.co/datasets/lmdmengdi/moss-002-sft-data.tabulartext-generation1M<n<10M0 likes32 downloads2mo agoHugging Face15Ujjwal-Tyagi /moshub Mos.Hub Code Dataset A comprehensive code dataset compiled from Mos.Hub, Moscow's official code hosting platform operated by the Moscow Government. This dataset is designed to support training code models with authentic Russian development practices and documentation. Overview The Mos.Hub Code Dataset represents a significant code corpus from Russia's governmental and municipal code hosting platform, capturing diverse projects across 297 programming languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/moshub.texttext-generation10M<n<100M0 likes30 downloads5mo agoHugging Face16nhagar /moscar_urls Dataset Card for moscar_urls This dataset provides the URLs and top-level domains associated with training records in oscar-corpus/mOSCAR. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/moscar_urls.texttext-generation100M<n<1B0 likes25 downloads1y agoHugging Face17douyipu-real /mosaic MOSAIC Dataset This repository packages the public MOSAIC data artifacts from the paper "MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment." MOSAIC is short for Multi-Objective Slice-Aware Iterative Curation for Alignment. It contains three annotated source training pools and five training subsets selected by the MOSAIC search loop under a fixed 1M-token budget. The release also includes flattened iteration metadata so the search trajectory can be inspected… See the full description on the dataset page: https://huggingface.co/datasets/douyipu-real/mosaic.tabulartext-generation10K<n<100K0 likes25 downloads6mo agoHugging Face18mosh2i /mimi_tokenizer Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mosh2i/mimi_tokenizer.texttext-generationn<1K0 likes22 downloads3y agoHugging Face19moseleydev /fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily. Here is your full content entirely in Markdown: # Fatima Fellowship 2026: Technical Challenge - Model Blind Spots ## 1. Model Overview - **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model) - **Parameters:** 0.6B - **Type:** Causal Language Model (Base / Pre-trained) --- ## 2. Methodology & Loading To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.texttext-generationn<1K0 likes22 downloads7mo agoHugging Face20lorinma /Slim-Moss003sft-zh因为原生的Moss003数量太大,所以进行了简单的去重。 去重方法大致为,只选择中文的对话,使用bert-base-chinese将第一个问题转换为embedding,使用类knn的方法抽取了1万条。并转换成了sharegpt格式。 texttext-generation10K<n<100K1 likes18 downloads3y agoHugging Face21ZombitX64 /moscar-corpus-thai-cleaned Dataset mOSCAR Thai Cleaned ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ รายละเอียดชุดข้อมูล จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train) ขนาดข้อมูล: 5,132,779,656 ไบต์ ฟีเจอร์: title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว ภาษา: ไทย (th) ลิขสิทธิ์: Apache-2.0 ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.texttext-generation1M<n<10M0 likes18 downloads1y agoHugging Face22fuzhou-jiang /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/fuzhou-jiang/moss-002-sft-data.tabulartext-generation1M<n<10M0 likes10 downloads9mo agoHugging Face23Mostafa190 /TwinnyAI-Personas-Datasetgated Overview The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems. Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.texttext-generation10K<n<100K1 likes6 downloads6mo agoHugging Face24N-Bot-Int /Moshpit-Combined-R2-Uncensoredgated MoshPIT Is to be Archived By the end of November, Please More Updated Corpus Like Iris, This Dataset Offer No Benefit and Provides alot of Hallucination 📜 Please read our Terms and Conditions before using this dataset. MoshPIT-R2 COMBINED UNCENSORED MoshPIT R2, is a dataset combining multiple Moshpit Iteration to make a 500 Dataset examples, MoshPIT R2 is purely generated through GPT2-XL, Mushed-R1 Dataset is Synthetic, It has no Curation But can offer Quick Dataset… See the full description on the dataset page: https://huggingface.co/datasets/N-Bot-Int/Moshpit-Combined-R2-Uncensored.text-generation1K<n<10K5 likes5 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.