CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.5k downloads5y agoHugging Face02sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K2 likes4k downloads2y agoHugging Face03AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes3.9k downloads2mo agoHugging Face04collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2k downloads1y agoHugging Face05yangyang857658468 /cc12m-webdataset CC12M WebDataset 这是CC12M数据集的WebDataset格式版本。 数据集信息 文件数量: 1098 总大小: 888796.33 MB 上传时间: 2025-03-18 14:45:49 使用方法 import webdataset as wds dataset = wds.WebDataset("https://huggingface.co/yangyang857658468/cc12m-webdataset/resolve/main/cc12m_*.tar") image10M<n<100M0 likes1.6k downloads2y agoHugging Face06hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes1.4k downloads1y agoHugging Face07croissantllm /croissant_dataset_no_web_data CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! https://arxiv.org/abs/2402.00786 Licenses Data redistributed here is subject to the original license under which it was collected. All license information is detailed in the Data section of the Technical report. Citation @misc{faysse2024croissantllm, title={CroissantLLM: A Truly Bilingual French-English Language Model}… See the full description on the dataset page: https://huggingface.co/datasets/croissantllm/croissant_dataset_no_web_data.texttranslation10M<n<100M4 likes1.3k downloads3y agoHugging Face08cat-state /MegaSynth-webdatasetimage1M<n<10M0 likes1.1k downloads10mo agoHugging Face09laion /clevr-webdatasetimage1M<n<10M7 likes651 downloads4y agoHugging Face10Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes562 downloads21h agoHugging Face11ARKseal /YFCC14M_subset_webdatasetimage1M<n<10M0 likes393 downloads5y agoHugging Face12cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes276 downloads2y agoHugging Face13BEE-spoke-data /open-web-math-minhash Dataset Card for "open-web-math-minhash" An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M: DatasetDict({ train: Dataset({ features: ['url', 'text', 'date', 'metadata'], num_rows: 1820241 }) }) Usage Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.texttext-generation1M<n<10M0 likes231 downloads9mo agoHugging Face14Lots-of-LoRAs /task1728_web_nlg_data_to_text Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.texttext-generation1K<n<10K0 likes174 downloads2y agoHugging Face15vovapoludnyakov /digital-detective-aic2026-webdataset Digital Detective AIC2026 — TAR shards Repackaged from QwertyNice/Digital_Detective_AIC2026, train_stage1.zip. Source repository metadata declares CC BY-SA 4.0. Images and masks are preserved without recompression. See manifest.json for the exact source ZIP SHA-256 and checksums of each TAR. Format: plain-tar-v1; original directory names preserved inside TAR, not WebDataset key suffixes. Training code and notebook: training/segformer_b2/. This repository does not contain trained… See the full description on the dataset page: https://huggingface.co/datasets/vovapoludnyakov/digital-detective-aic2026-webdataset.textimage-segmentationn<1K1 likes135 downloads16d agoHugging Face16shichen1231 /GBC1M_webdataimage1M<n<10M0 likes126 downloads2y agoHugging Face17UI-Simulator /UI-Simulator_web_datatext10K<n<100K1 likes123 downloads1y agoHugging Face18lucasnewman /libritts-r-webdatasetOfficial website: https://www.openslr.org/141/ This repository contains LibriTTS-R converted to a WebDataset. The original Wave files have been converted to 64kbps MP3 files for efficient streaming. LibriTTS-R (paper) is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the… See the full description on the dataset page: https://huggingface.co/datasets/lucasnewman/libritts-r-webdataset.audio100K<n<1M1 likes109 downloads2y agoHugging Face19softcatala /Softcatala-Web-Texts-Dataset Dataset Card for Softcatala-Web-Texts-Dataset Dataset Summary This repository contains Softcatala website content (articles and programs descriptions). Dataset size: articles.json contains 623 articles with 373233 words. programes.json contains 330 program descriptions with 49868 words. The license of the data is Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) or Universal Public Domain Dedication (CC0 1.0) Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/Softcatala-Web-Texts-Dataset.texttext-generationn<1K0 likes99 downloads2mo agoHugging Face20ratishsp /rephrased-web-data-quality-study Rephrased Web Data Quality Study LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated). Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45 Quality Scores (1-5 scale) Metric FAQ (n=965) Table (n=979) Tutorial (n=976) Math (n=994) Faithfulness 1.82 1.72 1.90 1.49 Info preservation 1.93 1.64 1.99 1.47 Appropriateness 3.54 2.87 2.48 1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.tabular1K<n<10K0 likes96 downloads3mo agoHugging Face21open-athena /glm52-datagen-r11-19-knowledge-web-search-mcqa-tracestext1K<n<10K0 likes95 downloads2mo agoHugging Face22ghemdd /gui_actor_webdataset GUI-Actor WebDataset A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks. Usage import webdataset as wds # Load the dataset dataset = wds.WebDataset("path/to/shards-*.tar") dataset = dataset.decode("pilrgb").to_tuple("jpg", "json") for image, metadata in dataset: # Process image and metadata pass Citation Please cite the original GUI-Actor paper if you use this dataset in your research. imagetext-generation1M<n<10M1 likes77 downloads1y agoHugging Face23jed351 /Cantonese-Web-Data Dataset Summary Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that. To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Common Crawl data and a Cantonese language detector, followed by a deduplication process using MinHash. The result is a high-quality collection of ~250K unique documents containing ~150 million words.… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese-Web-Data.text100K<n<1M5 likes76 downloads1y agoHugging Face24collabora /multilingual-librispeech-webdatasetaudio10K<n<100K1 likes61 downloads3y agoHugging Face25Liuff23 /solar_webdatatext1M<n<10M0 likes59 downloads6mo agoHugging Face26MC7ever /web-scraper-datasettext10K<n<100K0 likes58 downloads11h agoHugging Face27OpenTransformer /scraped-web-datatext100K<n<1M0 likes51 downloads9mo agoHugging Face28vyykaaa /dataset-web-attack-newstext10K<n<100K1 likes49 downloads9mo agoHugging Face29ilyada /web_accessibility_datasettextn<1K2 likes40 downloads2y agoHugging Face30AIMClab-RUC /PhD-webdataset PhD Webdataset This repository contains the packaged version of PhD. For a detailed introduction to PhD, please visit the official website. Overview The PhD Webdataset is designed to facilitate easy access and usage of the PhD dataset. It includes various fields in 'json' key. The data in this repo is totally the same as in PhD. Installation Ensure you have Hugging Face's datasets library installed. You can install it via pip: pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD-webdataset.imagevisual-question-answering100K<n<1M0 likes39 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.