CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fzzh /MACE MACE Dataset Dataset for paper "Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers". textquestion-answering10K<n<100K0 likes87 downloads8mo agoHugging Face02LVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes56 downloads1y agoHugging Face03LVSTCK /macedonian-corpus-raw Macedonian Corpus - Raw 🌟 Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.text10M<n<100M0 likes46 downloads1y agoHugging Face04LVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes46 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.