datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.ramanv-image-real-restorationpainting-restoration-eval-diagnosticsarqgan-indian-temple-restoration-dataset
ArqGAN-Indian Temple Restoration — Dataset
Training data for the ArqGAN Indian Temple Restoration project.
Authors
Rohan Gupta — Hugging Face: rxhxn2904
Achyut Acharya — GitHub/Hugging Face: AchyutAcharya182 / AchyutAcharya13
Part of the Viraasat AI project. Source repository:
Tanvi0705/Viraasat-AI
Contents
3DITA_raw/ — the source 3DITA (3D Indian Temple Architecture) dataset:
point-cloud scans and per-face annotations for Indian temples, split… See the full description on the dataset page: https://huggingface.co/datasets/rxhxn2904/arqgan-indian-temple-restoration-dataset.painting-restoration-eval-candidatesturkish-punctuation-restoration-500k
Turkish Punctuation Restoration 500K v2
Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, unpunctuated_text, punctuated_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.image_restoration
SCDC image restoration data bundle
Exactly the files used by the SCDC project (blind degradation graph encoder +
unified two-step ResShift restorer): https://github.com/anbinh93/SCDC
Archive
Content
Rain13K.tar
Rain13K pairs referenced by the Phase-1 splits (train / val / locked Test100+Test1200) and the Phase-2 manifest
RESIDE_used.tar
RESIDE hazy/clear files used for training/validation plus the 300 held-out evaluation scenes
LOLv2.tar
LOLv2-real, train (689) +… See the full description on the dataset page: https://huggingface.co/datasets/Anbinh93/image_restoration.Image_Restoration_Datasetstmp_lwir_image_restoration
LWIR image restoration task inputs (frozen snapshot)
This repository hosts the complete /root/inputs bundle for the
FrontierPhysics task LWIR-image-restoration:
five DARPA IH Dataset ENVI cubes (.bsq + .hdr)
per-scene zenith downwelling spectral radiance (6–14 µm, 4 nm),
computed with libRadtran from nearest-time layered soundings and HITRAN
line-by-line optical depths
the agent-facing DATASET_CARD.md
The cubes are too large for GitHub (each .bsq is 381–687 MB). The task… See the full description on the dataset page: https://huggingface.co/datasets/dccc2025/tmp_lwir_image_restoration.2021-punctuation-restorationThis dataset is designed to be used in training models
that restore punctuation marks from the output of
Automatic Speech Recognition system for Polish language.rvl-cdip-restoration
RVL-CDIP Restoration
A document image restoration dataset based on the original RVL-CDIP
dataset.
Original dataset: - https://huggingface.co/datasets/ChainYo/rvl-cdip
This version converts the classification dataset into a supervised
image-to-image restoration dataset.
Each sample contains: - lr → degraded document (input) - hr →
original clean document (target)
Images are stored as: - raw numpy bytes - dtype: uint8 - shape: (512,
512, 3)
Dataset Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/mkenfenheuer/rvl-cdip-restoration.satellite-image-restoration-datarestoration_test_data
Restoration Test Data
This dataset repo stores curated test data used by restoration_agent wrapper and eval scripts.
Scope:
Test inputs only.
No model weights.
No prediction outputs.
No runtime environments.
No logs.
No known-invalid legacy HDF5 files.
Layout
Upload status as of 2026-05-08:
All-in-one, HSI, MP-HSIR raw zips, and Haze1k test splits are uploaded.
SEN12MS-CR-TS america and europa raw archives are uploaded; africa, asiaEast, and asiaWest raw archives are… See the full description on the dataset page: https://huggingface.co/datasets/zzqsb/restoration_test_data.vehicle-window-restoration-dataset
Vehicle Window Restoration Dataset
Research data for vehicle-window and windshield image restoration, focused on two degradation factors:
reflection contamination caused by glass surfaces;
low-light exposure and sensor degradation.
The project has two evaluation lines:
Single-frame restoration: recover the transmitted scene from one degraded image.
Multi-frame/video restoration: use neighboring frames while preserving layer-aware temporal consistency.
Current… See the full description on the dataset page: https://huggingface.co/datasets/Xalzeroph/vehicle-window-restoration-dataset.MosaicLedger-Restoration-Captions
MosaicLedger Restoration Captions
This dataset offers synthetic captions describing simulated conservation treatments for cataloging exercises.
Immediate synthesis inputs
Exactly two components were used to generate the captions:
Gallery Treatment Phrasebook, licensed under Creative Commons Attribution-NonCommercial 4.0 International.
CaptionForge Generator, licensed under MIT License.
There are no other corpora, models, or direct inputs in this release.… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/MosaicLedger-Restoration-Captions.turkish-diacritics-restoration-1m
Turkish Diacritics Restoration 1M v2
ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, ascii_text, restored_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.image-restoration-v2vietnamese-diacritic-restoration-corpusI have downloaded it from Kaggle. I sincerely thank the author for making it available.
mural-art-restorationpunctuation_restoration_4096_complex# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_4096_complex.vedic-accent-restoration-dataset
Citation
@inproceedings{tsukagoshi-2025-accent-restoration,
title = {Automatic Accent Restoration in Vedic Sanskrit with Neural Language Models},
author = {Tsukagoshi, Yuzuki and Ohmukai, Ikki},
booktitle = {Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025)},
editor = {Bhattacharya, Arnab and Goyal, Pawan and Ghosh, Saptarshi and Ghosh, Kripabandhu},
year =… See the full description on the dataset page: https://huggingface.co/datasets/yzk/vedic-accent-restoration-dataset.punctuation_restoration_600_complex# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_600_complex.punctuation_restoration_900_complex# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_900_complex.english_punctuation_restorationataturk_voice_no_restorationrs_restoration_ckptpunctuation_restoration# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration.punctuation_restoration_700_complex# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_700_complex.weather_restorationParquet files separated into 3 chunks
Dataset sources
Rain
Snow
Raindrop
Haze
RealRain-1K
Snow100K
DeRaindrop
RainDS
RESIDE-beta
784
1500
861
350
1462
punctuation_restoration_1200# punctuation_restoration
## Dataset Summary
This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text.
It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format.
- 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1200.
