datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.macedonian-corpus-cleaned
Macedonian Corpus - Cleaned
raw version here
Paper
🌟 Key Highlights
Size: 35.5 GB, Word Count: 3.31 billion
Filtered for irrelevant and low-quality content using C4 and Gopher filtering.
Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.Forensic_Toolkit_DatasetForensic Toolkit Dataset
Overview
The Forensic Toolkit Dataset is a comprehensive collection of 300 digital forensics and incident response (DFIR) tools, designed for training AI models, supporting forensic investigations, and enhancing cybersecurity workflows. The dataset includes both mainstream and unconventional tools, covering disk imaging, memory analysis, network forensics, mobile forensics, cloud forensics, blockchain analysis, and AI-driven forensic techniques. Each entry provides… See the full description on the dataset page: https://huggingface.co/datasets/macerj/Forensic_Toolkit_Dataset.
