datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
somaliweb-v1
SomaliWeb v1 — Quality-filtered Somali web corpus
📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus
SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.cyber_security
Digital Literacy & Cybersecurity Nepali SFT Dataset
Dataset Overview
This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity.
The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response.
Dataset Statistics
Property
Value
Total records
1,000
Valid JSONL rows
1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.patriae-cuba-literature-dataset
Patriae Cuban Literature Dataset (31k)
Dataset de literatura cubana curado por el equipo de Patriae como parte de su participación en el evento SomosNLP 2026, con el objetivo de emplearse por el mismo en la realización de tareas de reproducción del dialecto cubano.
📌 Nota de procedencia: Este repositorio es un espejo (mirror) oficial para el evento. El desarrollo activo, las actualizaciones del dataset y la autoría principal pertenecen a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuba-literature-dataset.
