datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.Uno-Curriculum
Uno-Curriculum
Training corpus for a hierarchical-delegation router: a small language
model that decomposes a task into subtasks and routes each subtask to a
(worker model, skill) pair.
Every row comes from a real public HuggingFace dataset — the
question and gold_answer are sampled verbatim from the dataset
identified by the source field. Every row then goes through the
same three-stage pipeline (router probe → teacher trajectory →
noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.LingxiDiag-16K-unofficial-mirror
LingxiDiag-16K Backup Mirror
This repository is an unofficial backup mirror of the original dataset:
Original dataset: XuShihao6715/LingxiDiag-16K
We are not the original authors of this dataset.
We provide this repository only as a backup copy for non-commercial research access, preservation, and reproducibility.
This mirror is not affiliated with, maintained by, or endorsed by the original authors or the Evermind Lingxi Team.
Original Dataset
LingxiDiag-16K is a… See the full description on the dataset page: https://huggingface.co/datasets/yiyangdenanzi/LingxiDiag-16K-unofficial-mirror.
