datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prophet-mosque-library
Prophet's Mosque Library
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.MOSAIC-Refactoring
Agentic Pull Request Dataset
Dataset Overview
The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below.
Cohort
Pull Requests
Merged Pull Requests
Repositories
Sum of Additions
Sum of Deletions
Humans
517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.mOSCARMore info can be found here: https://oscar-project.github.io/documentation/versions/mOSCAR/
Paper link: https://arxiv.org/abs/2406.08707
New features:
Additional filtering steps were applied to remove toxic content (more details in the next version of the paper, coming soon).
Spanish split is now complete.
Face detection in images to blur them once downloaded (coordinates are reported on images of size 256 respecting aspect ratio).
Additional language identification of the documents to… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/mOSCAR.mosaic-combine-all
Mosaic format for combine all dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all
load it,
from streaming import LocalDataset
import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.voxpopuli_mosel_curatormosaic-starcoder-filtered
Mosaic format for filtered starcoder dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.MOSAIC_Dataset
MOSAIC Dataset
Project Page | Paper | Code | Dataset | Model
This repository releases the built-in MOSAIC multi-source motion dataset in the following paper:
MOSAIC: Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation
The dataset is organized into:
Human motions stored in an AMASS-style format
Unitree G1 motions retargeted from human motions and converted to NPZ for training/visualization
It includes motions from:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/MOSAIC_Dataset.mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.mosaic-dedup-text-dataset
Mosaic format for dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.LoRA-WiSE
Dataset Card for the LoRA WiSE benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive
benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models
LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in
the "Dataset Size Recovery from LoRA Weights" paper.
Task Details
Dataset Description
Dataset Structure
Data Subsets
Data Fields
Dataset Creation
Citation Information
🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.mosaic-dedup-text-dataset-filtered
Mosaic format for filtered dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.mosaic-madlad-400-ms
Mosaic format for extra dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms
load it,
from streaming import LocalDataset
import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.prophet-mosque-library-compressed
Prophet's Mosque Library - Compressed
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.mosscap_prompt_injection
mosscap_prompt_injection
This is a dataset of prompt injections submitted to the game Mosscap by Lakera.
This variant of the game Gandalf was created for DEF CON 31.
Note that the Mosscap levels may no longer be available in the future.
Note that we release every prompt that we received, regardless of whether it truly is a prompt injection or not.
There are hundrends of thousands of prompts and many of them are not actual prompt injections (people ask Mosscap all kinds of things).… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/mosscap_prompt_injection.MOSSBench
Dataset Card for MOSSBench
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Data Visualization
Data Source
Automatic Evaluation
License
Citation
Dataset Description
Humans are prone to cognitive distortions — biased thinking patterns that lead to exaggerated responses to specific stimuli, albeit in very different contexts. MOSSBench demonstrates that advanced MLLMs exhibit similar tendencies. While… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/MOSSBench.mosaic-nanot5-512moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.cmu-mosei-comp-seq
CMU-MOSEI: Computational Sequences (Unofficial Mirror)
This repository provides a mirror of the official computational sequence files from the CMU-MOSEI dataset, which are required for multimodal sentiment and emotion research. The original download links are currently down, so this mirror is provided for the research community.
Note: This is an unofficial mirror. All data originates from Carnegie Mellon University and original authors. If you are a dataset creator and want this… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/cmu-mosei-comp-seq.MOSAIC_model_ckptmosaic-whisper-combinedAudio files from these links
https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp
https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp
tts-pretrain-clones-3m-mos
TTS Pretrain Clones (3M) — with DNSMOS
This is SynDataLab/tts-pretrain-clones-3m
with an added per-utterance dnsmos column (DNSMOS P.835 OVRL score, float32),
computed with the sig_bak_ovr.onnx model.
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
dnsmos: overall MOS quality estimate per utterance (higher is better)
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers.… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/tts-pretrain-clones-3m-mos.moss-003-sft-chinese-zhtw
Dataset Card for "moss-003-sft-chinese-zhtw"
資料集摘要
本資料集主要是應用於專案:MOSS: 開源對話語言模型 所收集的數據。
MOSS 是支援中英雙語和多種外掛程式的開源對話語言模型,moss-moon 系列模型具有160億參數,在FP16精度下可在單張A100/A800或兩張3090顯示卡運行,在INT4/8精度下可在單張3090顯示卡運行。 MOSS基座語言模型在約七千億中英文以及程式碼單字上預訓練得到,後續經過對話指令微調、插件增強學習和人類偏好訓練具備多輪對話能力及使用多種插件的能力。
原始資料來源
moss-003-sft-data: moss-moon-003-sft 所使用的多輪對話數據,基於 MOSS-002 內測階段採集的約10萬用戶輸入數據和 gpt-3.5-turbo 構造而成,相比 moss-002-sft-data,moss-003-sft-data… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/moss-003-sft-chinese-zhtw.moss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2)
Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices.
Each prompt is a fixed, optimized champion performance direction (instruction) paired
with a Gemma-generated topic text (text) — together, one performance to render. For every
prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64
within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.dolly_hhrlhf
Dataset Card for "dolly_hhrlhf"
This dataset is a combination of Databrick's dolly-15k dataset and a filtered subset of Anthropic's HH-RLHF. It also includes a test split, which was missing in the original dolly set. That test set is composed of 200 randomly selected samples from dolly + 4,929 of the test set samples from HH-RLHF which made it through the filtering process. The train set contains 59,310 samples; 15,014 - 200 = 14,814 from Dolly, and the remaining 44,496 from… See the full description on the dataset page: https://huggingface.co/datasets/mosaicml/dolly_hhrlhf.sada-train-preprocessedmoss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.sada-train-wav2vec2-xls-r-300m-ar-preprocessedmosaic-instructions
Mosaic format for instructions dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-instructions.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-instructions
load it,
from streaming import LocalDataset… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-instructions.gex_dataset_fusedMOSEL-FR-audio-codes
