CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01imageomics /PUUM-koa-restoration-camera-trap-dataset Dataset Card for Koa Associated Biodiversity Camera Trap Dataset This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025. Dataset Details This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.imageimage-classification1K<n<10K1 likes636 downloads9mo agoHugging Face02lingamvamshikrishnareddy /ramanv-image-real-restorationtext10K<n<100K0 likes600 downloads23d agoHugging Face03GoktugD /turkish-punctuation-restoration-500k Turkish Punctuation Restoration 500K v2 Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, unpunctuated_text, punctuated_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.texttext-generation100K<n<1M0 likes143 downloads2mo agoHugging Face04GoktugD /turkish-diacritics-restoration-1m Turkish Diacritics Restoration 1M v2 ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, ascii_text, restored_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.texttext-generation1M<n<10M0 likes31 downloads2mo agoHugging Face05bdanko /image-restoration-v2image1K<n<10K0 likes27 downloads6mo agoHugging Face06yammdd /vietnamese-diacritic-restoration-corpusI have downloaded it from Kaggle. I sincerely thank the author for making it available. text10M<n<100M0 likes26 downloads2mo agoHugging Face07picard47at /punctuation_restoration_4096_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_4096_complex.text1K<n<10K0 likes22 downloads1y agoHugging Face08yzk /vedic-accent-restoration-dataset Citation @inproceedings{tsukagoshi-2025-accent-restoration, title = {Automatic Accent Restoration in Vedic Sanskrit with Neural Language Models}, author = {Tsukagoshi, Yuzuki and Ohmukai, Ikki}, booktitle = {Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025)}, editor = {Bhattacharya, Arnab and Goyal, Pawan and Ghosh, Saptarshi and Ghosh, Kripabandhu}, year =… See the full description on the dataset page: https://huggingface.co/datasets/yzk/vedic-accent-restoration-dataset.texttext-generation100K<n<1M0 likes22 downloads9mo agoHugging Face09picard47at /punctuation_restoration_600_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_600_complex.text1K<n<10K0 likes21 downloads1y agoHugging Face10picard47at /punctuation_restoration_900_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_900_complex.text1K<n<10K0 likes21 downloads1y agoHugging Face11thenlpresearcher /english_punctuation_restorationtext100K<n<1M0 likes20 downloads10mo agoHugging Face12picard47at /punctuation_restoration# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration.text1K<n<10K0 likes18 downloads1y agoHugging Face13picard47at /punctuation_restoration_700_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_700_complex.text1K<n<10K0 likes18 downloads1y agoHugging Face14maagic6 /weather_restorationParquet files separated into 3 chunks Dataset sources Rain Snow Raindrop Haze RealRain-1K Snow100K DeRaindrop RainDS RESIDE-beta 784 1500 861 350 1462 image1K<n<10K2 likes17 downloads3y agoHugging Face15picard47at /punctuation_restoration_1200# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1200.text1K<n<10K0 likes16 downloads1y agoHugging Face16picard47at /punctuation_restoration_400# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_400.text1K<n<10K0 likes15 downloads1y agoHugging Face17picard47at /punctuation_restoration_512_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_512_complex.text1K<n<10K0 likes14 downloads1y agoHugging Face18picard47at /punctuation_restoration_750_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_750_complex.text1K<n<10K0 likes14 downloads1y agoHugging Face19picard47at /punctuation_restoration_1350_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1350_complex.text1K<n<10K0 likes13 downloads1y agoHugging Face20dicta-il /hebrew-space-restoration-corpus Restoring Missing Spaces in Scraped Hebrew Social Media This dataset holds the test corpus used in the 2025 W-Nut paper: Avi Shmidman and Shaltiel Shmidman, "Restoring Missing Spaces in Scraped Hebrew Social Media", The 10th Workshop on Noisy and User-generated Text (W-NUT), 2025. The corpus consists of ~6,000 Hebrew sentences, sampled from the Hebrew portion of FineWeb-2. Each row of the dataset contains two fields: input: The Hebrew sentence with 1-4 spaces randomly removed (see… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew-space-restoration-corpus.text1K<n<10K0 likes12 downloads1y agoHugging Face21picard47at /punctuation_restoration_800_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_800_complex.text1K<n<10K0 likes12 downloads1y agoHugging Face22Natashadonoh /yoruba-diacritic-restoration-dataset Yorùbá Diacritic Restoration Dataset Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text. Dataset Details Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded) Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.text1K<n<10K0 likes12 downloads1mo agoHugging Face23picard47at /punctuation_restoration_512# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_512.text1K<n<10K1 likes11 downloads1y agoHugging Face24picard47at /punctuation_restoration_1350# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1350.text1K<n<10K0 likes10 downloads1y agoHugging Face25picard47at /punctuation_restoration_1350_complex2# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1350_complex2.text1K<n<10K0 likes10 downloads1y agoHugging Face26noone-m /arabic-punctuation-restorationtext100K<n<1M0 likes10 downloads9mo agoHugging Face27picard47at /punctuation_restoration_256# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_256.text1K<n<10K0 likes8 downloads1y agoHugging Face28picard47at /punctuation_restoration_1500# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1500.text1K<n<10K0 likes8 downloads1y agoHugging Face29picard47at /punctuation_restoration_1024_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_1024_complex.text1K<n<10K0 likes8 downloads1y agoHugging Face30picard47at /punctuation_restoration_full_complex# punctuation_restoration ## Dataset Summary This dataset is designed for **instruction fine-tuning** of large language models (LLMs), especially for the **Qwen3** family, to perform **punctuation restoration** on Mandarin Chinese text. It is derived from the [`AWeirdDev/zh-tw-articles-6k`](https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-6k) dataset. The `context` field is processed to create input-output pairs in the Qwen3-style message format. - 🔧 **User message**: A cleaned… See the full description on the dataset page: https://huggingface.co/datasets/picard47at/punctuation_restoration_full_complex.text1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.