datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prophet-mosque-library
Prophet's Mosque Library
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.MOSAIC-Refactoring
Agentic Pull Request Dataset
Dataset Overview
The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below.
Cohort
Pull Requests
Merged Pull Requests
Repositories
Sum of Additions
Sum of Deletions
Humans
517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.mOSCARMore info can be found here: https://oscar-project.github.io/documentation/versions/mOSCAR/
Paper link: https://arxiv.org/abs/2406.08707
New features:
Additional filtering steps were applied to remove toxic content (more details in the next version of the paper, coming soon).
Spanish split is now complete.
Face detection in images to blur them once downloaded (coordinates are reported on images of size 256 respecting aspect ratio).
Additional language identification of the documents to… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/mOSCAR.mosaic-combine-all
Mosaic format for combine all dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all
load it,
from streaming import LocalDataset
import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.mosTransformationvoxpopuli_mosel_curatorCMU-MOSEImoss-003-sft-data
moss-003-sft-data
** More information: MOSS Paper**
Conversation Without Plugins
Categories
Category
# samples
Brainstorming
99,162
Complex Instruction
95,574
Code
198,079
Role Playing
246,375
Writing
341,087
Harmless
74,573
Others
19,701
Total
1,074,551
Others contains two categories: Continue(9,839) and Switching(9,862).The Continue category refers to instances in a conversation where the user asks the system to continue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-003-sft-data.mosaic-starcoder-filtered
Mosaic format for filtered starcoder dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.MOSAIC_Dataset
MOSAIC Dataset
Project Page | Paper | Code | Dataset | Model
This repository releases the built-in MOSAIC multi-source motion dataset in the following paper:
MOSAIC: Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation
The dataset is organized into:
Human motions stored in an AMASS-style format
Unitree G1 motions retargeted from human motions and converted to NPZ for training/visualization
It includes motions from:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/MOSAIC_Dataset.mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.mosaic-dedup-text-dataset
Mosaic format for dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.LoRA-WiSE
Dataset Card for the LoRA WiSE benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive
benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models
LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in
the "Dataset Size Recovery from LoRA Weights" paper.
Task Details
Dataset Description
Dataset Structure
Data Subsets
Data Fields
Dataset Creation
Citation Information
🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.test_datasetmosaic-dedup-text-dataset-filtered
Mosaic format for filtered dedup text dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset
load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.mosaic-madlad-400-ms
Mosaic format for extra dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms
load it,
from streaming import LocalDataset
import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.simplificationmosscap_prompt_injection
mosscap_prompt_injection
This is a dataset of prompt injections submitted to the game Mosscap by Lakera.
This variant of the game Gandalf was created for DEF CON 31.
Note that the Mosscap levels may no longer be available in the future.
Note that we release every prompt that we received, regardless of whether it truly is a prompt injection or not.
There are hundrends of thousands of prompts and many of them are not actual prompt injections (people ask Mosscap all kinds of things).… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/mosscap_prompt_injection.MOSSBench
Dataset Card for MOSSBench
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Data Visualization
Data Source
Automatic Evaluation
License
Citation
Dataset Description
Humans are prone to cognitive distortions — biased thinking patterns that lead to exaggerated responses to specific stimuli, albeit in very different contexts. MOSSBench demonstrates that advanced MLLMs exhibit similar tendencies. While… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/MOSSBench.mosaic-nanot5-512prophet-mosque-library-compressed
Prophet's Mosque Library - Compressed
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.MOSEv2
MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
🔥 Evaluation Server | 🏠 Homepage | 📄 Paper | 🔗 GitHub
Download
We recommend using huggingface-cli to download:
pip install -U "huggingface_hub[cli]"
huggingface-cli download FudanCVL/MOSEv2 --repo-type dataset --local-dir ./MOSEv2 --local-dir-use-symlinks False --max-workers 16
Dataset Summary
MOSEv2 is a comprehensive video object segmentation dataset designed to advance… See the full description on the dataset page: https://huggingface.co/datasets/FudanCVL/MOSEv2.cmu-mosei-comp-seq
CMU-MOSEI: Computational Sequences (Unofficial Mirror)
This repository provides a mirror of the official computational sequence files from the CMU-MOSEI dataset, which are required for multimodal sentiment and emotion research. The original download links are currently down, so this mirror is provided for the research community.
Note: This is an unofficial mirror. All data originates from Carnegie Mellon University and original authors. If you are a dataset creator and want this… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/cmu-mosei-comp-seq.MergingMOSAIC_model_ckptyolo-baselines-no-mosaic-runsprophet-mosque-library
Prophet's Mosque Library
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/mhaamh19/prophet-mosque-library.mataian-moses-stac
Note: This description was drafted with AI assistance and is subject to revision.
MOSES Initiative — Matai'an Open Science and Engineering Sharing Platform
The MOSES Initiative (Mataian Open Science and Engineering Sharing Initiative) publishes scientific and engineering datasets for the Matai'an Creek (馬太鞍溪) watershed in eastern Taiwan (23.56°N–24.15°N, 121.16°E–121.68°E). All datasets are discoverable through a STAC catalog.
Datasets
1. Radar… See the full description on the dataset page: https://huggingface.co/datasets/NTU-CompHydroMet-Lab/mataian-moses-stac.
