CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kazuyi1222 /sdsdsdsaaaa0 likes2.2k downloads4mo agoHugging Face02hvai /sdsetStable Difusion store for learner get files 1 likes769 downloads1y agoHugging Face03UniverseTBD /mmu_sdss_sdss mmu_sdss_sdss HATS Catalog Collection This is the collection of HATS catalogs representing mmu_sdss_sdss. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_sdss_sdss.tabular100K<n<1M0 likes766 downloads4mo agoHugging Face04SamsungSDS-Research /SDS-KoPub-VDR-Benchmark 📘 Dataset Summary SDS KoPub-VDR is a benchmark dataset for Visual Document Retrieval (VDR) in the context of Korean public documents. It contains real-world government document images paired with natural-language queries, corresponding answer pages, and ground-truth answers. The dataset is designed to evaluate AI models that go beyond simple text matching, requiring comprehensive understanding of visual layouts, tables, graphs, and images to accurately locate relevant… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSDS-Research/SDS-KoPub-VDR-Benchmark.imagevisual-document-retrieval10K<n<100K14 likes668 downloads11mo agoHugging Face05sdsr /erp-and-eroticaMirror of ERP/RP and erotica raw data collection (Edit: 07 Jan 2025 16:09 UTC) (No new data after 04 Jan 2024 12:01 UTC). texttext-generation10M<n<100M8 likes648 downloads6mo agoHugging Face06Pratofeitoo /sd-sdxl-loras0 likes490 downloads5mo agoHugging Face07withmartian /SDS_train_gsm8k0 likes461 downloads6mo agoHugging Face08MINT-SDSU /CLIP-CC 📚 CLIP-CC Dataset (Movie Clips Edition) Paper | arXiv | Project Page | Benchmark Code (CLIP-CC-Bench) | Dataset Repo (CLIP-CC) CLIP-CC is a curated dataset for long-form video description: 200 movie clips sourced from YouTube, each about 90 seconds long (~5 hours in total) and drawn from more than 140 films spanning 1959–2024, each paired with one human-written English reference description averaging 402 ± 208 words. The references were written by four graduate-student… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SDSU/CLIP-CC.textvideo-text-to-textn<1K0 likes431 downloads2mo agoHugging Face09Jerry999 /sds-news-ragThis dataset is derived from the Global News Dataset. Please refer to the original source (also cited below) and ensure that your use complies with its terms and conditions. Webz.io News Dataset Repository Introduction Welcome to the Webz.io News Dataset Repository! This repository is created by Webz.io and is dedicated to providing free datasets of publicly available news articles. We release new datasets weekly, each containing around 1,000 news articles focused on… See the full description on the dataset page: https://huggingface.co/datasets/Jerry999/sds-news-rag.0 likes412 downloads1y agoHugging Face10VSSA-SDSA /LT_Medical_S_corpusEnglish | Lietuvių English LT_Medical_S_corpus — Lithuanian Medical Speech Corpus A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps. Columns Column Type Description audio Audio Audio sentence string Ground truth transcription duration_ms int Recording duration in milliseconds medical_area string RADIOLOGIJA or SEIMOS gender string MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.audioautomatic-speech-recognition10K<n<100K1 likes351 downloads5mo agoHugging Face11Forturne /SDS-KoPub-OCR SDS-KoPub OCR Results & Embeddings OCR layout parsing results and VL embeddings for the SDS-KoPub-VDR-Benchmark corpus (40,781 Korean public document pages). Contents File Description Size ocr_results.jsonl GLM-OCR structured layout results (regions, markdown, bbox, labels) 40,781 records parsed_texts.jsonl Extracted text per page (embedding input) 40,781 records embeddings/corpus_regions.npy Region multimodal embeddings (image+caption) (21052, 2048)… See the full description on the dataset page: https://huggingface.co/datasets/Forturne/SDS-KoPub-OCR.document-question-answeringn<1K0 likes288 downloads7mo agoHugging Face12VSSA-SDSA /LT_AI_BLKT Model Card for LT_AI_BLKT (EN) / LT_AI_BLKT modelio kortelė (LT) Table of Contents / Turinys Description (EN) / Aprašas (LT) Dataset summary(EN) Main columns (EN) Data composition (EN) Distribution of text types (EN) Sources (EN) Time periods (EN) Licensing (EN) / Licencija (LT)) Intended use (EN) / Numatyti naudojimo atvejai (LT) Restrictions] (EN) / Apribojimai (LT) Limitations and bias (EN) / Ribotumai ir šališkumas (LT) Citation (EN) / Citavimas (LT)… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_AI_BLKT.documenttext-generation1M<n<10M1 likes230 downloads5mo agoHugging Face13allenai /sdsd-dialogues Self Directed Synthetic Dialogues (SDSD) v0 This dataset is an experiment in procedurally generating synthetic dialogues between two language models. For each dialogue, one model, acting as a "user" generates a plan based on a topic, subtopic, and goal for a conversation. Next, this model attempts to act on this plan and generating synthetic data. Along with the plan is a principle which the model, in some successful cases, tries to cause the model to violate the principle resulting… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sdsd-dialogues.texttext-generation100K<n<1M19 likes207 downloads2y agoHugging Face14Fucheng /GaSNet-II-SDSS-dataset The SDSS spectra used in the paper: https://arxiv.org/abs/2311.04146 The code is available on Github: https://github.com/Fucheng-Zhong/GaSNet-II If the dataset or code helps in your research, please cite paper 2311.04146 0 likes196 downloads2y agoHugging Face15astro-conformal-project /split_sdss_hsc_embeddingstext10K<n<100K0 likes185 downloads2mo agoHugging Face16sdsc2005-migration /visadocumentn<1K0 likes162 downloads6mo agoHugging Face17SDSD2WW /YRD_DATAimagen<1K2 likes156 downloads5mo agoHugging Face18MultimodalUniverse /sdss--- description: 'Spectra dataset based on SDSS-IV. ' homepage: https://www.sdss.org/ version: 1.0.0 citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\ % \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.tabular100K<n<1M3 likes152 downloads2y agoHugging Face19SDSD2WW /shanghai_landimageimagen<1K0 likes141 downloads2mo agoHugging Face20SDSC /open-pulse-hackathon-data-analysis LauzHack Projects Dataset Dataset Summary This dataset contains comprehensive information about projects submitted to LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project includes details about the project title, description, team members, awards, and categories. LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.tabularn<1K0 likes133 downloads5mo agoHugging Face21kshitijd /mmu-norm-sdss SDSS spectra — L1 (release v1) This L1 repository contains 271,966 spectra matched across the release, in 11 shards (about 15 GB). Schema spectrum struct per row: lambda (vacuum, heliocentric, observer frame, Å; variable length ≤ ~4,700), flux (×10⁻¹⁷ erg s⁻¹ cm⁻² Å⁻¹), ivar, lsf_sigma, mask (native; nonzero = bad — verified), valid. Plus object_id, positions, metadata. Padding: the upstream arrays used lambda = −1 with zero inverse variance for padding. Those… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-sdss.tabular100K<n<1M0 likes125 downloads1mo agoHugging Face22withmartian /SDS_train_mmlu-pro SDS Train - MMLU-Pro Activation extraction dataset for studying Switching Dynamical Systems (SDS) in reasoning LLMs, generated from the TIGER-Lab/MMLU-Pro benchmark (test split, ~4000 samples per model). Models Reasoning (RLVR fine-tuned) models with their corresponding base models: Reasoning Model Base Model Layers Extracted deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Qwen/Qwen2.5-14B 28 (middle), 47 (final) deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B… See the full description on the dataset page: https://huggingface.co/datasets/withmartian/SDS_train_mmlu-pro.10K<n<100K0 likes123 downloads7mo agoHugging Face23kshitijd /mmu-l2-sdss L2 model-ready views (release v1) The mmu-l2-* repositories provide processed, model-ready versions of the matching mmu-norm-* L1 data, while L1 keeps the normalized source measurements. L2 applies documented processing steps for training and evaluation, with the details for reversing each transformation stored in the row or in provenance.json. Repo View Reversal mmu-l2-tess per-sector relative flux f/median−1, time from first valid cadence flux =… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-l2-sdss.text100K<n<1M0 likes106 downloads1mo agoHugging Face24LSDB /mmu_sdss_sdss--- description: 'HATS version of MultimodalUniverse/sdss: Spectra dataset based on SDSS-IV. ' homepage: https://www.sdss.org/ version: 1.0.0 citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\ % \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for… See the full description on the dataset page: https://huggingface.co/datasets/LSDB/mmu_sdss_sdss.tabular1K<n<10K1 likes105 downloads11mo agoHugging Face25fraisdufour /sd-stuffimagen<1K0 likes99 downloads3y agoHugging Face26withmartian /SDS_math500_test0 likes91 downloads7mo agoHugging Face27sdsc2005-migration /news_embeddingtext100K<n<1M0 likes90 downloads6mo agoHugging Face28sdSDDDDDDD /drugs Molecules This repository stores downloaded 3D molecular datasets collected into one gated Hugging Face dataset repo. Included datasets geom_drugs: GEOM-Drugs (~7GB including QM9, ~304K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 39.8 GB. geom_qm9: GEOM-QM9 (included in GEOM, ~130K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 148.0 B. spice: SPICE v2 (~7GB, ~19K molecules / 1.1M conformers). Status: uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/sdSDDDDDDD/drugs.other0 likes88 downloads4mo agoHugging Face29IDEALLab /OpenR1-SDS-FeasibilitySparsity-Logs-v1 OpenR1-SDS-FeasibilitySparsity-Logs-v1 Raw JSONL source-of-truth logs backing the SDS feasibility-sparsity analysis. Contents feasibility_sparsity/1743423 <- /workspace/llm-finetuning/.hf_private_staging/OpenR1-SDS-FeasibilitySparsity-Logs-v1/source-0 (required) feasibility_sparsity/1743424 <- /workspace/llm-finetuning/.hf_private_staging/OpenR1-SDS-FeasibilitySparsity-Logs-v1/source-1 (required) feasibility_sparsity/1743425 <-… See the full description on the dataset page: https://huggingface.co/datasets/IDEALLab/OpenR1-SDS-FeasibilitySparsity-Logs-v1.0 likes87 downloads5mo agoHugging Face30BASF-AI /SDS-Eyes-Protection-Classification Safety Data Sheets Eyes Protection Classification This dataset contains Safety Data Sheets (SDS) sourced from Kaggle, consisting of over 200,000 documents. SDS are detailed documents providing essential information on the properties and hazards of chemicals, ensuring user safety and compliance with regulatory standards. A subset of these documents was pre-processed, cleaned, and annotated to classify whether eye protection is required when handling materials. The labels were… See the full description on the dataset page: https://huggingface.co/datasets/BASF-AI/SDS-Eyes-Protection-Classification.texttext-classification100K<n<1M0 likes84 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.