datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PluraMath
PluraMath 🌍➕
Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
🔎 TL;DR
PluraMath is a human-curated multilingual mathematical reasoning benchmark that extends PolyMath to 18 additional underrepresented languages spanning 6 language families — from mid-resource languages such as Hindi and Turkish down to extreme low-resource languages such as Upper and Lower Sorbian (< 15k L1 speakers).
Every language contains 500… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/PluraMath.Nurisk-ICRA2026
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/Nurisk-ICRA2026.vn-bctn-supplement
Vietnam Annual Reports Supplement Dataset (546 PDFs)
Tập dữ liệu bổ sung gồm 546 Báo cáo thường niên (BCTN) của các công ty niêm yết trên thị trường chứng khoán Việt Nam (HOSE, HNX, UPCoM), được trích xuất và chuẩn hóa để bổ sung cho tập dữ liệu gốc 13,982 báo cáo trên Zenodo.
Cấu trúc lưu trữ
.
├── README.md
├── bctn_supplement_index.parquet # Metadata index (546 rows)
├── manifest.csv # Chi tiết danh mục file
└── pdfs/
├── {TICKER}/… See the full description on the dataset page: https://huggingface.co/datasets/Tumiqa103/vn-bctn-supplement.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.Code-170k-tumbuka
Dataset Description
Code-170k-tumbuka is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Tumbuka, making coding education accessible to Tumbuka speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Tumbuka language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-tumbuka.
