CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes291k downloads1y agoHugging Face02AISE-TUDelft /MOSAIC-Refactoring Agentic Pull Request Dataset Dataset Overview The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below. Cohort Pull Requests Merged Pull Requests Repositories Sum of Additions Sum of Deletions Humans 517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.tabular10M<n<100M4 likes29k downloads3mo agoHugging Face03oscar-corpus /mOSCARMore info can be found here: https://oscar-project.github.io/documentation/versions/mOSCAR/ Paper link: https://arxiv.org/abs/2406.08707 New features: Additional filtering steps were applied to remove toxic content (more details in the next version of the paper, coming soon). Spanish split is now complete. Face detection in images to blur them once downloaded (coordinates are reported on images of size 256 respecting aspect ratio). Additional language identification of the documents to… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/mOSCAR.text100M<n<1B18 likes9.5k downloads2y agoHugging Face04malaysia-ai /mosaic-combine-all Mosaic format for combine all dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all load it, from streaming import LocalDataset import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.textn<1K0 likes8.2k downloads3y agoHugging Face05logantl /mos0 likes6.3k downloads1y agoHugging Face06mostafabehroozi /Transformation0 likes5.1k downloads3d agoHugging Face07sonalsannigrahi /voxpopuli_mosel_curatortabular10M<n<100M0 likes4.5k downloads3mo agoHugging Face08tamb2203579 /CMU-MOSEI2 likes4k downloads3mo agoHugging Face09OpenMOSS-Team /moss-003-sft-data moss-003-sft-data ** More information: MOSS Paper** Conversation Without Plugins Categories Category # samples Brainstorming 99,162 Complex Instruction 95,574 Code 198,079 Role Playing 246,375 Writing 341,087 Harmless 74,573 Others 19,701 Total 1,074,551 Others contains two categories: Continue(9,839) and Switching(9,862).The Continue category refers to instances in a conversation where the user asks the system to continue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-003-sft-data.68 likes3.9k downloads2y agoHugging Face10malaysia-ai /mosaic-starcoder-filtered Mosaic format for filtered starcoder dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.textn<1K0 likes3.9k downloads3y agoHugging Face11BAAI-Humanoid /MOSAIC_Dataset MOSAIC Dataset Project Page | Paper | Code | Dataset | Model This repository releases the built-in MOSAIC multi-source motion dataset in the following paper: MOSAIC: Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation The dataset is organized into: Human motions stored in an AMASS-style format Unitree G1 motions retargeted from human motions and converted to NPZ for training/visualization It includes motions from:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/MOSAIC_Dataset.documentroboticsn<1K9 likes3.7k downloads7mo agoHugging Face12FBK-MT /mosel Dataset Description, Collection, and Source The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses. In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.imageautomatic-speech-recognition1M<n<10M93 likes3k downloads1y agoHugging Face13malaysia-ai /mosaic-dedup-text-dataset Mosaic format for dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.textn<1K0 likes2.9k downloads3y agoHugging Face14MoSalama98 /LoRA-WiSE Dataset Card for the LoRA WiSE benchmark The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in the "Dataset Size Recovery from LoRA Weights" paper. Task Details Dataset Description Dataset Structure Data Subsets Data Fields Dataset Creation Citation Information 🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.imagetabular-classification1K<n<10K5 likes2k downloads2y agoHugging Face15mosaicml /test_dataset0 likes1.9k downloads3y agoHugging Face16malaysia-ai /mosaic-dedup-text-dataset-filtered Mosaic format for filtered dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.textn<1K0 likes1.9k downloads3y agoHugging Face17malaysia-ai /mosaic-madlad-400-ms Mosaic format for extra dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-madlad-400-ms.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms load it, from streaming import LocalDataset import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-madlad-400-ms.textn<1K0 likes1.7k downloads3y agoHugging Face18mostafabehroozi /simplification0 likes1.5k downloads3d agoHugging Face19Lakera /mosscap_prompt_injection mosscap_prompt_injection This is a dataset of prompt injections submitted to the game Mosscap by Lakera. This variant of the game Gandalf was created for DEF CON 31. Note that the Mosscap levels may no longer be available in the future. Note that we release every prompt that we received, regardless of whether it truly is a prompt injection or not. There are hundrends of thousands of prompts and many of them are not actual prompt injections (people ask Mosscap all kinds of things).… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/mosscap_prompt_injection.text100K<n<1M21 likes1.3k downloads2y agoHugging Face20AIcell /MOSSBench Dataset Card for MOSSBench Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Data Visualization Data Source Automatic Evaluation License Citation Dataset Description Humans are prone to cognitive distortions — biased thinking patterns that lead to exaggerated responses to specific stimuli, albeit in very different contexts. MOSSBench demonstrates that advanced MLLMs exhibit similar tendencies. While… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/MOSSBench.imagevisual-question-answeringn<1K6 likes1.3k downloads2y agoHugging Face21huseinzol05 /mosaic-nanot5-512textn<1K0 likes1.3k downloads2y agoHugging Face22ieasybooks-org /prophet-mosque-library-compressed Prophet's Mosque Library - Compressed 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.textimage-to-text10K<n<100K0 likes1.3k downloads1y agoHugging Face23OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face24FudanCVL /MOSEv2 MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes 🔥 Evaluation Server | 🏠 Homepage | 📄 Paper | 🔗 GitHub Download We recommend using huggingface-cli to download: pip install -U "huggingface_hub[cli]" huggingface-cli download FudanCVL/MOSEv2 --repo-type dataset --local-dir ./MOSEv2 --local-dir-use-symlinks False --max-workers 16 Dataset Summary MOSEv2 is a comprehensive video object segmentation dataset designed to advance… See the full description on the dataset page: https://huggingface.co/datasets/FudanCVL/MOSEv2.object-detection1K<n<10K8 likes1.2k downloads1y agoHugging Face25reeha-parkar /cmu-mosei-comp-seq CMU-MOSEI: Computational Sequences (Unofficial Mirror) This repository provides a mirror of the official computational sequence files from the CMU-MOSEI dataset, which are required for multimodal sentiment and emotion research. The original download links are currently down, so this mirror is provided for the research community. Note: This is an unofficial mirror. All data originates from Carnegie Mellon University and original authors. If you are a dataset creator and want this… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/cmu-mosei-comp-seq.audio5 likes1.1k downloads1y agoHugging Face26mostafabehroozi /Merging0 likes1.1k downloads13d agoHugging Face27MOSAIC-UCSD /MOSAIC_model_ckpttextn<1K0 likes1.1k downloads5mo agoHugging Face28duyle2408 /yolo-baselines-no-mosaic-runstabular1M<n<10M0 likes945 downloads3d agoHugging Face29mhaamh19 /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/mhaamh19/prophet-mosque-library.image-to-text10K<n<100K4 likes938 downloads7mo agoHugging Face30NTU-CompHydroMet-Lab /mataian-moses-stac Note: This description was drafted with AI assistance and is subject to revision. MOSES Initiative — Matai'an Open Science and Engineering Sharing Platform The MOSES Initiative (Mataian Open Science and Engineering Sharing Initiative) publishes scientific and engineering datasets for the Matai'an Creek (馬太鞍溪) watershed in eastern Taiwan (23.56°N–24.15°N, 121.16°E–121.68°E). All datasets are discoverable through a STAC catalog. Datasets 1. Radar… See the full description on the dataset page: https://huggingface.co/datasets/NTU-CompHydroMet-Lab/mataian-moses-stac.0 likes917 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.