datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EuRoC_MAV_DatasetWaterDrum-Ax
WaterDrum: Watermarking for Data-centric Unlearning Metric
WaterDrum provides an unlearning benchmark for the evaluation of the effectiveness and practicality of unlearning. The repository contains the ArXiv corpus of WaterDrum (WaterDrum-Ax), which contains both unwatermarked and watermarked ArXiv paper abstracts across
20 categories published after the release of the Llama-2 model. Each category contains 400 data samples, aggregating into 8000 samples in the full training set. The… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/WaterDrum-Ax.glowtact-gelsight-tactile
GlowTact & GelSight Mini Tactile Probing Dataset
Optical tactile captures from a CNC probing rig, collected with two vision-based
tactile sensors: GlowTact (a custom sensor reconstructed with a 9DTact-style
darkness-to-depth model) and the commercial GelSight Mini.
Every CNC probe record pairs a tactile image with the exact indenter position
(mm) and the normal force (N) measured by a calibrated HX711 load cell at
the moment of capture, encoded directly in the filename.… See the full description on the dataset page: https://huggingface.co/datasets/yxma/glowtact-gelsight-tactile.WaterDrum-TOFU
WaterDrum: Watermarking for Data-centric Unlearning Metric
WaterDrum provides an unlearning benchmark for the evaluation of the effectiveness and practicality of unlearning. This repository contains the TOFU corpus of WaterDrum (WaterDrum-TOFU), which contains both unwatermarked and watermarked question-answering datasets based on the original TOFU dataset.
The data samples were watermarked with Waterfall.
Update Notice: 15/01/2026
We have updated Glow-AI/WaterDrum-TOFU to version… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/WaterDrum-TOFU.venv-meMeMoso-vits-32k
SoftVC VITS Singing Voice Conversion
强调!!!!!!!!!!!!
SoVits是语音转换 (说话人转换),作用是将一个音频中语音的音色转化为目标说话人的音色,并不是TTS (文本转语音),SoVits虽然基于Vits开发,但两者是两个不同的项目,请不要搞混,要训练TTS请前往 Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
使用规约
请自行解决数据集的授权问题,任何由于使用非授权数据集进行训练造成的问题,需自行承担全部责任和一切后果,与sovits无关!… See the full description on the dataset page: https://huggingface.co/datasets/GlowingBrick/so-vits-32k.CrossDocked2020banking-clinc-oos
Dataset Card for "banking-clinc-oos"
More Information needed
stockRealXBench
RealXBench
RealXBench is a comprehensive visual question answering benchmark dataset. The full dataset contains 300 high-quality image-question-answer triplets. Due to internal regulations, only a subset of 194 samples is released in this open-source version.
Dataset Structure
Each example contains:
query: The question about the image (in English)
answer: The ground truth answer(s), with multiple answers separated by "or"
perception: Difficulty level for perception task… See the full description on the dataset page: https://huggingface.co/datasets/glowol/RealXBench.TT100K_Annotationedwatershed-c4watershed-blogglowspco16_tabular_data
pco16
Throughput of distributed LLM training as a function of the parallelism
configuration, on 16 GPUs across 2 hosts.
Tabular benchmark for the bolt problem pco16: the 919
rows are the full candidate set.
Objective throughput_mean, to maximise. NaN where the run OOMed, since
no throughput is observed at all -- a hidden (crash) constraint, not a bad value.
Constraint ran_successfully. 632 of 919
configurations run; the rest exhaust GPU memory. Feasibility is only learnable
by… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/pco16_tabular_data.pco32_tabular_data
pco32
Throughput of distributed LLM training as a function of the parallelism
configuration, on 32 GPUs across 8 hosts.
Tabular benchmark for the bolt problem pco32: the 379
rows are the full candidate set.
Objective throughput_mean, to maximise. NaN where the run OOMed, since
no throughput is observed at all -- a hidden (crash) constraint, not a bad value.
Constraint ran_successfully. 339 of 379
configurations run; the rest exhaust GPU memory. Feasibility is only learnable
by… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/pco32_tabular_data.v14-fake-glow-ttspco64_tabular_data
pco64
Throughput of distributed LLM training as a function of the parallelism
configuration, on 64 GPUs across 16 hosts.
Tabular benchmark for the bolt problem pco64: the 780
rows are the full candidate set.
Objective throughput_mean, to maximise. NaN where the run OOMed, since
no throughput is observed at all -- a hidden (crash) constraint, not a bad value.
Constraint ran_successfully. 560 of 780
configurations run; the rest exhaust GPU memory. Feasibility is only learnable
by… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/pco64_tabular_data.GTSDBGTSDB - German Traffic Sign Detection Benchmark
virl-39k-rubrics-testamazon-semantic-ids-recommendationhttps://github.com/snap-research/GRID
Beauty, Sports, Toys
amazon-p5-grid-dataset
This dataset contains pre-processed Amazon user-item interaction sequences and product metadata,
specifically curated for the GRID (Generative Recommendation with Semantic IDs) framework by Snap Research.
It is derived from the P5 recommendation benchmark, covering diverse domains such as Beauty, Sports, and Toys to support
large-scale generative modeling. The data is structured to facilitate the generation of… See the full description on the dataset page: https://huggingface.co/datasets/GlowBond/amazon-semantic-ids-recommendation.wethink-rubrics-20k-testmulti_news_dataset_editedAteron__Glowing-Forest-12B-details
Dataset Card for Evaluation run of Ateron/Glowing-Forest-12B
Dataset automatically created during the evaluation run of model Ateron/Glowing-Forest-12B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Ateron__Glowing-Forest-12B-details.wan_glow_effectThis dataset contains videos generated using Wan 2.1 T2V 14B.
multiews-dataset01geothought-rubrics-17k-testCCP-SRClassical Chinese Painting Super-Resolution Dataset
Cultural Heritage Artwork Super-Resolution Benchmark
Real-World Ancient Painting Restoration Dataset for SR
Answer_card
