datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdsdsdsaaaasdsetStable Difusion store for learner get files
mmu_sdss_sdss
mmu_sdss_sdss HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_sdss_sdss.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_sdss_sdss.SDS-KoPub-VDR-Benchmark
📘 Dataset Summary
SDS KoPub-VDR is a benchmark dataset for Visual Document Retrieval (VDR) in the context of
Korean public documents. It contains real-world government document images paired with natural-language
queries, corresponding answer pages, and ground-truth answers. The dataset is designed to evaluate AI models that
go beyond simple text matching, requiring comprehensive understanding of visual layouts, tables, graphs, and images
to accurately locate relevant… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSDS-Research/SDS-KoPub-VDR-Benchmark.erp-and-eroticaMirror of ERP/RP and erotica raw data collection (Edit: 07 Jan 2025 16:09 UTC)
(No new data after 04 Jan 2024 12:01 UTC).
sd-sdxl-lorasSDS_train_gsm8kCLIP-CC
📚 CLIP-CC Dataset (Movie Clips Edition)
Paper | arXiv | Project Page | Benchmark Code (CLIP-CC-Bench) | Dataset Repo (CLIP-CC)
CLIP-CC is a curated dataset for long-form video description: 200 movie clips sourced from YouTube, each about 90 seconds long (~5 hours in total) and drawn from more than 140 films spanning 1959–2024, each paired with one human-written English reference description averaging 402 ± 208 words. The references were written by four graduate-student… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SDSU/CLIP-CC.sds-news-ragThis dataset is derived from the Global News Dataset. Please refer to the original source (also cited below) and ensure that your use complies with its terms and conditions.
Webz.io News Dataset Repository
Introduction
Welcome to the Webz.io News Dataset Repository! This repository is created by Webz.io and is dedicated to providing free datasets of publicly available news articles. We release new datasets weekly, each containing around 1,000 news articles focused on… See the full description on the dataset page: https://huggingface.co/datasets/Jerry999/sds-news-rag.LT_Medical_S_corpusEnglish | Lietuvių
English
LT_Medical_S_corpus — Lithuanian Medical Speech Corpus
A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps.
Columns
Column
Type
Description
audio
Audio
Audio
sentence
string
Ground truth transcription
duration_ms
int
Recording duration in milliseconds
medical_area
string
RADIOLOGIJA or SEIMOS
gender
string
MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.SDS-KoPub-OCR
SDS-KoPub OCR Results & Embeddings
OCR layout parsing results and VL embeddings for the
SDS-KoPub-VDR-Benchmark
corpus (40,781 Korean public document pages).
Contents
File
Description
Size
ocr_results.jsonl
GLM-OCR structured layout results (regions, markdown, bbox, labels)
40,781 records
parsed_texts.jsonl
Extracted text per page (embedding input)
40,781 records
embeddings/corpus_regions.npy
Region multimodal embeddings (image+caption)
(21052, 2048)… See the full description on the dataset page: https://huggingface.co/datasets/Forturne/SDS-KoPub-OCR.LT_AI_BLKT
Model Card for LT_AI_BLKT (EN) / LT_AI_BLKT modelio kortelė (LT)
Table of Contents / Turinys
Description (EN) / Aprašas (LT)
Dataset summary(EN)
Main columns (EN)
Data composition (EN)
Distribution of text types (EN)
Sources (EN)
Time periods (EN)
Licensing (EN) / Licencija (LT))
Intended use (EN) / Numatyti naudojimo atvejai (LT)
Restrictions] (EN) / Apribojimai (LT)
Limitations and bias (EN) / Ribotumai ir šališkumas (LT)
Citation (EN) / Citavimas (LT)… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_AI_BLKT.sdsd-dialogues
Self Directed Synthetic Dialogues (SDSD) v0
This dataset is an experiment in procedurally generating synthetic dialogues between two language models.
For each dialogue, one model, acting as a "user" generates a plan based on a topic, subtopic, and goal for a conversation.
Next, this model attempts to act on this plan and generating synthetic data.
Along with the plan is a principle which the model, in some successful cases, tries to cause the model to violate the principle resulting… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sdsd-dialogues.GaSNet-II-SDSS-dataset
The SDSS spectra used in the paper: https://arxiv.org/abs/2311.04146
The code is available on Github: https://github.com/Fucheng-Zhong/GaSNet-II
If the dataset or code helps in your research, please cite paper 2311.04146
split_sdss_hsc_embeddingsvisaYRD_DATAsdss---
description: 'Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.shanghai_landimageopen-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.mmu-norm-sdss
SDSS spectra — L1 (release v1)
This L1 repository contains 271,966 spectra matched across the release, in 11 shards (about 15 GB).
Schema
spectrum struct per row: lambda (vacuum, heliocentric, observer frame, Å; variable length ≤ ~4,700), flux (×10⁻¹⁷ erg s⁻¹ cm⁻² Å⁻¹), ivar, lsf_sigma, mask (native; nonzero = bad — verified), valid. Plus object_id, positions, metadata.
Padding: the upstream arrays used lambda = −1 with zero inverse variance for padding. Those… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-sdss.SDS_train_mmlu-pro
SDS Train - MMLU-Pro
Activation extraction dataset for studying Switching Dynamical Systems (SDS) in reasoning LLMs, generated from the TIGER-Lab/MMLU-Pro benchmark (test split, ~4000 samples per model).
Models
Reasoning (RLVR fine-tuned) models with their corresponding base models:
Reasoning Model
Base Model
Layers Extracted
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Qwen/Qwen2.5-14B
28 (middle), 47 (final)
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B… See the full description on the dataset page: https://huggingface.co/datasets/withmartian/SDS_train_mmlu-pro.mmu-l2-sdss
L2 model-ready views (release v1)
The mmu-l2-* repositories provide processed, model-ready versions of the matching mmu-norm-* L1 data, while L1 keeps the normalized source measurements. L2 applies documented processing steps for training and evaluation, with the details for reversing each transformation stored in the row or in provenance.json.
Repo
View
Reversal
mmu-l2-tess
per-sector relative flux f/median−1, time from first valid cadence
flux =… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-l2-sdss.mmu_sdss_sdss---
description: 'HATS version of MultimodalUniverse/sdss: Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for… See the full description on the dataset page: https://huggingface.co/datasets/LSDB/mmu_sdss_sdss.sd-stuffSDS_math500_testnews_embeddingdrugs
Molecules
This repository stores downloaded 3D molecular datasets collected into one gated Hugging Face dataset repo.
Included datasets
geom_drugs: GEOM-Drugs (~7GB including QM9, ~304K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 39.8 GB.
geom_qm9: GEOM-QM9 (included in GEOM, ~130K). Status: uploaded. Uploaded: yes. Cleaned locally: yes, staged size 148.0 B.
spice: SPICE v2 (~7GB, ~19K molecules / 1.1M conformers). Status: uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/sdSDDDDDDD/drugs.OpenR1-SDS-FeasibilitySparsity-Logs-v1
OpenR1-SDS-FeasibilitySparsity-Logs-v1
Raw JSONL source-of-truth logs backing the SDS feasibility-sparsity analysis.
Contents
feasibility_sparsity/1743423 <- /workspace/llm-finetuning/.hf_private_staging/OpenR1-SDS-FeasibilitySparsity-Logs-v1/source-0 (required)
feasibility_sparsity/1743424 <- /workspace/llm-finetuning/.hf_private_staging/OpenR1-SDS-FeasibilitySparsity-Logs-v1/source-1 (required)
feasibility_sparsity/1743425 <-… See the full description on the dataset page: https://huggingface.co/datasets/IDEALLab/OpenR1-SDS-FeasibilitySparsity-Logs-v1.SDS-Eyes-Protection-Classification
Safety Data Sheets Eyes Protection Classification
This dataset contains Safety Data Sheets (SDS) sourced from Kaggle, consisting of over 200,000 documents. SDS are detailed documents providing essential information on the properties and hazards of chemicals, ensuring user safety and compliance with regulatory standards. A subset of these documents was pre-processed, cleaned, and annotated to classify whether eye protection is required when handling materials. The labels were… See the full description on the dataset page: https://huggingface.co/datasets/BASF-AI/SDS-Eyes-Protection-Classification.
