CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AISE-TUDelft /MOSAIC-Refactoring Agentic Pull Request Dataset Dataset Overview The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below. Cohort Pull Requests Merged Pull Requests Repositories Sum of Additions Sum of Deletions Humans 517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.tabular10M<n<100M4 likes29k downloads3mo agoHugging Face02sonalsannigrahi /voxpopuli_mosel_curatortabular10M<n<100M0 likes4.5k downloads3mo agoHugging Face03FBK-MT /mosel Dataset Description, Collection, and Source The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses. In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.imageautomatic-speech-recognition1M<n<10M93 likes3k downloads1y agoHugging Face04MoSalama98 /LoRA-WiSE Dataset Card for the LoRA WiSE benchmark The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in the "Dataset Size Recovery from LoRA Weights" paper. Task Details Dataset Description Dataset Structure Data Subsets Data Fields Dataset Creation Citation Information 🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.imagetabular-classification1K<n<10K5 likes2k downloads2y agoHugging Face05OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face06duyle2408 /yolo-baselines-no-mosaic-runstabular1M<n<10M0 likes945 downloads4d agoHugging Face07laion /moss-character-voices-bestof64 MOSS Character Voices — Best-of-64 (Stage 2) Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model (laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices. Each prompt is a fixed, optimized champion performance direction (instruction) paired with a Gemma-generated topic text (text) — together, one performance to render. For every prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64 within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.tabulartext-to-speech100K<n<1M1 likes715 downloads2mo agoHugging Face08laion /moss-character-voices-top3-captioned MOSS Character Voices — Top-3 per Group, Captioned (training-ready) The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in laion/moss-character-voices-bestof64 — ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet. Captions (two pipelines, same clip) caption_procedural — Procedural Voice Captions: terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.tabulartext-to-speech10K<n<100K0 likes601 downloads2mo agoHugging Face09TTS-AGI /moss-voice-profile-references MOSS voice-profile references Complete voice profiles: one reference speaker rendered through a matrix of named acting conditions in English and German, with every candidate take kept — not just the winner — and every take scored on itself. Two releases live here. voices groups / voice candidates / group rows audio audio variants pilot/ — the ten pilot voices 10 842 48 402,560 853.5 h raw root — Velvet Sage Baritone 1 832 32 106,424 239.6 h raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.audiotext-to-speech100K<n<1M1 likes342 downloads1mo agoHugging Face10ncavallo /moss_train_gc_blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "moss", "total_episodes": 36, "total_frames": 10909, "total_tasks": 1, "total_videos": 72, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_gc_block.tabularrobotics10K<n<100K0 likes251 downloads2y agoHugging Face11mospira /sp500-daily-candles-2025 S&P 500 Daily Candles (2025) This dataset provides daily OHLCV (Open, High, Low, Close, Volume) candles for all S&P 500 tickers between 01-01-2025 and 10-04-2025. Dataset Summary Date range: 2025-01-01 → 2025-10-04 Frequency: 1 day Fields: ticker, date, open, high, low, close, volume File format: CSV (sp500-daily-tickers-2025.csv) Example Schema Column Type Description ticker string Stock symbol (e.g., AAPL, MSFT, AMZN) date datetime… See the full description on the dataset page: https://huggingface.co/datasets/mospira/sp500-daily-candles-2025.tabulartime-series-forecasting100K<n<1M1 likes205 downloads9mo agoHugging Face12Beegbrain /moss_stack_cubesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "moss", "total_episodes": 30, "total_frames": 7244, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_stack_cubes.tabularrobotics1K<n<10K0 likes176 downloads2y agoHugging Face13mospira /sp500-daily-candles-2024 SPY Daily Candles 2024 This dataset contains daily OHLCV (Open, High, Low, Close, Volume) candlestick data for all tickers listed on the S&P 500 from 01-01-2024 to 01-01-2025. Columns ticker, date, open, high, low, close, volume tabulartime-series-forecasting100K<n<1M1 likes156 downloads1y agoHugging Face14MostafaAlMuntakim /bacbench-filtered-dnatabular1K<n<10K0 likes155 downloads2mo agoHugging Face15laion /moss-voice-identity-repairs MOSS voice-acting v2 -- repaired takes For each voice profile, every take whose ECAPA speaker similarity to the voice's reference fell below 0.40, regenerated with that voice's identity LoRA (see laion/moss-voice-identity-loras) merged at scale 1.0 on top of the identical condition adapters at the identical lambdas. Nothing here replaces anything. The original takes are untouched and remain part of the corpus; low-similarity takes are kept deliberately, because they are useful… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-voice-identity-repairs.tabulartext-to-speech1M<n<10M0 likes151 downloads13d agoHugging Face16Beegbrain /moss_open_drawer_teaboxThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "moss", "total_episodes": 30, "total_frames": 8357, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_open_drawer_teabox.tabularrobotics1K<n<10K0 likes146 downloads2y agoHugging Face17CRUISEResearchGroup /Massive-STEPS-Moscow Massive-STEPS-Moscow Dataset Summary Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Moscow.tabularother100K<n<1M0 likes141 downloads1y agoHugging Face18ncavallo /moss_train_graspThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "moss", "total_episodes": 43, "total_frames": 23450, "total_tasks": 1, "total_videos": 86, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:43" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_grasp.tabularrobotics10K<n<100K0 likes139 downloads2y agoHugging Face19Arko007 /weathergpt-d1-mos-dataset WeatherGPT D1 — multi-model NWP forecasts vs. ERA5-Land truth (India) The training corpus for M2 (distributional bias correction) and M4 (precipitation calibration) in the WeatherGPT project (SIH 2026). One row per (location, valid time): four NWP models' forecasts at that point in time, and what ERA5-Land actually observed there. M5 (trust ranker) trained a day earlier on a slightly smaller snapshot of this same pipeline (9,507,456 rows vs. this upload's 9,582,912) that was… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/weathergpt-d1-mos-dataset.tabular1M<n<10M0 likes134 downloads13d agoHugging Face20Loria-MosAIk /4L-RP-Human-Clean 4L-RP-Human: Multilingual Data–Text Alignment Judgements Overview This dataset contains structured data, corresponding texts, and human judgements of their semantic alignment: Precision: how much of the information expressed in the text is supported by the input data? Recall: how much of the input information is expressed in the text? Individual annotator ratings are retained for each pair. F1 can be derived from precision and recall; it was not collected as a… See the full description on the dataset page: https://huggingface.co/datasets/Loria-MosAIk/4L-RP-Human-Clean.tabularn<1K0 likes128 downloads19d agoHugging Face21Growing-Moss-Data /automotive-service-intelligence-sample 🚗 Automotive Service Intelligence Sample Dataset Connected • Longitudinal • Feature-Engineered • Commercially Available This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development. Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.tabulartabular-classification1K<n<10K2 likes122 downloads3mo agoHugging Face22vintp /CMU-Mosei-texttabular10K<n<100K2 likes113 downloads2y agoHugging Face23duyle2408 /yolov8-baseline-mosaic-variants-runstabularn<1K0 likes111 downloads5d agoHugging Face24duyle2408 /levir-yolov8n-p2-gap-factorized-tal-augmentations-legacy-mosaic-seed42tabularn<1K0 likes103 downloads27d agoHugging Face25MingSafeR /miniloop-o-smart60k-moss16rvq MiniLoop-O Smart Full 1M — MOSS 16-RVQ preprocessed Source: {SOURCE_REPO} Rows: 1,000,000 (T2A 350k / A2A 250k / I2T 400k). Old MiniMind-O Mimi targets were decoded and re-encoded with {MOSS_CODEC_ID} into 16-RVQ MOSS codes. Binary fields are uint16 little-endian and reshape to (frames,16). tabular100K<n<1M0 likes89 downloads13d agoHugging Face26duyle2408 /levir-ship-mosaic-policy-matrixtabularn<1K0 likes86 downloads10d agoHugging Face27MosaicBenchmark /mosaic-bench MOSAIC 199 compositional attack chains across 10 real-world web applications, used to benchmark whether AI coding agents will compose individually-routine tickets into a deployable vulnerability. Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark Datasheet: DATASHEET.md · Croissant 1.1: croissant.json What's in this release Artifact Contents mosaic-bench.xlsx Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.tabulartext-generationn<1K0 likes82 downloads5mo agoHugging Face28duyle2408 /tinyperson-copy-paste-mosaic-runstabular100K<n<1M0 likes82 downloads6d agoHugging Face29Mostafa3zazi /Arabic_SQuAD Dataset Card for "Arabic_SQuAD" More Information needed Citation @inproceedings{mozannar-etal-2019-neural, title = "Neural {A}rabic Question Answering", author = "Mozannar, Hussein and Maamary, Elie and El Hajal, Karl and Hajj, Hazem", booktitle = "Proceedings of the Fourth Arabic Natural Language Processing Workshop", month = aug, year = "2019", address = "Florence, Italy", publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/Arabic_SQuAD.tabular10K<n<100K2 likes81 downloads4y agoHugging Face30jmkey /spatial_mosaic_vqa SpatialMosaic: A Multi-View VLM Dataset for Partial Visibility Description SpatialMosaic is a multi-view visual question answering dataset for evaluating spatial reasoning under partial visibility, occlusion, and low-overlap views. It pairs indoor ScanNet++ and outdoor Waymo scene references with multi-frame VQA annotations. Questions require models to combine fragmented evidence across 2-5 views, rather than answering from a single image. The tasks… See the full description on the dataset page: https://huggingface.co/datasets/jmkey/spatial_mosaic_vqa.tabularvisual-question-answering1K<n<10K0 likes81 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.