datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.FaithEval-unanswerable-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-unanswerable-v1.0.LogicBench-v1.0
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really "reason" over the natural language? This question has been receiving significant research attention and many reasoning skills such as commonsense, numerical, and qualitative have been studied. However, the crucial skill pertaining… See the full description on the dataset page: https://huggingface.co/datasets/cogint/LogicBench-v1.0.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.MedCalc-Bench-v1.0
Important: This dataset is kept for reproducibility purposes only. Please use the most up-to-date version 1.2, for the most
revised and corrected dataset available.
We have fixed 12 calculator implementations, ensured the Relevant Entities
section best matches with what was specified by a patient note, and have also replaced notes which are better fits
for a given calculator to make the dataset more applicable for real-life sitatuons.
Because of the number of changes, we find this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.0.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.stem-reasoning-v1.0.0-ccbysa-001
YouAI Data — stem-reasoning-v1.0.0-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.silma-rag-qa-benchmark-v1.0
SILMA RAGQA Benchmark Dataset V1.0
SILMA RAGQA is a dataset and benchmark created by silma.ai to assess the effectiveness of Arabic Language Models in Extractive Question Answering tasks, with a specific emphasis on RAG applications
The benchmark includes 17 bilingual datasets in Arabic and English, spanning various domains
What capabilities does the benchmark test?
General Arabic and English QA capabilities
Ability to handle short and long contexts
Ability to… See the full description on the dataset page: https://huggingface.co/datasets/silma-ai/silma-rag-qa-benchmark-v1.0.M3LLM-data-v1.0.0
M3LLM Data
Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release.
Contents
Collection
Split
Records
Description
PMC-MI supervised instruction corpus
train
224,401
Six instruction formats after partitioning and release filtering
PMC-MI policy-refinement partition
train
10,355
Policy-refinement instances for Stage II after release… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data-v1.0.0.vmad_v1.0
🩺 VMAD-300K: Verified Medical AI Dataset (Song ngữ Việt - Anh)
VMAD-300K là bộ dữ liệu y khoa song ngữ (Tiếng Việt - Tiếng Anh) chất lượng cao gồm 324,042 cặp câu hỏi - đáp, tóm tắt lâm sàng và hướng dẫn sử dụng thuốc được kiểm duyệt tự động 2 lớp (2-Layer Verification System).
Bộ dữ liệu được thiết kế đặc biệt để Supervised Fine-Tuning (SFT) các dòng mô hình ngôn ngữ lớn (LLM) trong lĩnh vực Y tế như Qwen 2.5 (0.5B - 14B), LLaMA 3.1, Gemma 2, Mistral.
📊 Thống kê… See the full description on the dataset page: https://huggingface.co/datasets/vnhkhwa/vmad_v1.0.SV-TrustEval-C-1.0
SV‑TrustEval‑C 🚨🔒
🔍 Overview
SV‑TrustEval‑C is the first reasoning‑based benchmark designed to rigorously evaluate Large Language Models (LLMs) on both structure (control/data flow) and semantic reasoning for vulnerability analysis in C source code. Unlike existing benchmarks that focus solely on pattern recognition, SV‑TrustEval‑C measures logical consistency, adaptability to code transformations, and real‑world security reasoning across six core tasks.
Our… See the full description on the dataset page: https://huggingface.co/datasets/LLMs4CodeSecurity/SV-TrustEval-C-1.0.ATANTV1.0-corpus
ATANT Narrative Test Corpus
Automated Test for Acceptance of Narrative Truth, v1.0
The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time.
Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710)
Standard repository: github.com/Kenotic-Labs/ATANT
Author: Samuel Sameer Tanguturi
Affiliation: Kenotic Labs
Published: April 2026
Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.CHaMEL_v1.0
CHaMEL v1.0 Dataset
Controllable Harness for Model adaptability Evaluation under Latent rule-shifts
Dataset Description
CHaMEL is a benchmark harness for evaluating the adaptive reasoning capabilities of Large Language Models (LLMs) through dynamic rule-shift tasks. This dataset contains the pre-generated task stimuli and human evaluation baselines used in the CHaMEL paper.
CHaMEL implements a factorial experiment design where researchers can independently control the… See the full description on the dataset page: https://huggingface.co/datasets/unknown202612/CHaMEL_v1.0.Tengentoppa-sft-v1.0
Tengentoppa corpus for sft (Combined Japanese Instruction Dataset)
概要
このデータセットは、日本語の instruction-following データセット16個を統合して作成された大規模な教師あり学習用データセットです。様々なタスクや対話形式を含む多様なデータソースから構成されています。
データセット構成
基本情報
フォーマット: JSON
各データポイントの構造:{
"instruction": "指示/質問文",
"input": "追加の文脈や入力(オプション)",
"output": "応答/回答文"
}
データセット変換コード
データセット作成に使用したコードは以下のGitHubリポジトリで公開しています:
dataset-processor
含まれるデータセット
Hachi-Alpaca_newans… See the full description on the dataset page: https://huggingface.co/datasets/DeL-TaiseiOzaki/Tengentoppa-sft-v1.0.stem-reasoning-v1.0.0-ccbysa-002
YouAI Data — stem-reasoning-v1.0.0-ccbysa-002
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 239 step-by-step reasoning chains and 723 instruction/response pairs across 443 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 476 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-002.Haakkim-1.0v
Haakkim — Arabic LLM Arena Battles v1.0
Haakkim (حَكِّم, "judge") is an open arena-style human preference evaluation platform for Arabic large language models. This dataset contains all battle records collected during the first release snapshot, including voted comparisons and skipped battles across 11 Arabic dialect varieties.
🌐 Platform: https://haakkim.tech
🏆 Live Leaderboard: https://haakkim.tech/#leaderboard
Dataset Summary
Statistic
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/Haakkim/Haakkim-1.0v.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/supraja04/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The… See the full description on the dataset page: https://huggingface.co/datasets/meet-the-1337/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Tengentoppa-grpo-v1.0
Tengentoppa-grpo-v1.0
1. データセットの読み込み
from datasets import load_dataset
# Hugging Faceから直接読み込み
dataset = load_dataset("your-username/japanese-edu-problems-tex")
# またはローカルファイルから
import json
with open("corrected_problems_tex.json", "r", encoding="utf-8") as f:
data = json.load(f)
RuBQ_1.0
RuBQ 1.0
hercules-v1.0
hercules-v1.0 dataset
The Hercules-v1.0 dataset is a turbo-charged version of teknium/openhermes, achieved by augmenting its data sources. Some of the datasets used in teknium/openhermes are older versions. Hercules-v1.0 addresses this issue by updating the data sources such as airoboros and WizardLM. Additionally, Hercules-v1.0 uses ise-uiuc/Magicoder-Evol-Instruct-110K instead of sahil2801/CodeAlpaca-20k as the primary code dataset.
Furthermore, I have removed the Unnatural… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hercules-v1.0.az_instruct_qa-v1.0
❓ az_instruct_qa-v1.0
Description:7,000 example Azerbaijani instruction-style question–answer pairs. Each entry includes a direct factual or explanatory answer and context, suitable for training instruct-following LLMs or building QA datasets. Covers broad topics such as science, education, culture, nature, history, and social issues.
Use Cases:
Instruction-tuning
Retrieval-augmented generation (RAG)
Chatbot Q&A
Multilingual benchmark
Fields:
id: UUID
question: Instructional or… See the full description on the dataset page: https://huggingface.co/datasets/az-llm/az_instruct_qa-v1.0.az_corpus-v1.0
📚 az_corpus-v1.0
Description:A 5,000-example Azerbaijani corpus enriched with metadata such as sentiment, topic, emotion analysis, structured Q&A, and translation. Each entry contains raw and cleaned text, quality annotations, and entity-role mappings. The dataset is built from public domain books, news, and journalistic content.
Use Cases:
Sentiment analysis
Entity recognition (with role labels)
Emotion detection
Instruction generation
Metadata extraction
Fields:
id: UUID… See the full description on the dataset page: https://huggingface.co/datasets/az-llm/az_corpus-v1.0.Viel-v1.0
This is Viel
An Industrial Robot turned into a robot assistant.
A bit rough around the edges and would rather enter sleep mode than helping you with your inane request.
This dataset contains the entire 52k Alpaca Dataset modified to mimic Viel's personality and speaking pattern and a hint of her background.
. . .
Have fun then
. . .
Finetune coming soon, just putting her dataset here for easier Collab Access
chunks_validation_1.0.0{
"id": "chunks_validation_1.0.0",
"name": "Chunks Validation 1.0.0",
"description": "Chunks validation dataset to validate if RAG is extracting the correct small chunks from the source large chunks.",
"task_categories": [
"question-answering"
],
"languages": [
"pt"
],
"dataset": "chunks_validation_1.0.0",
"features": {
"id": "Value(dtype='int64', id=None)",
"base": "Value(dtype='string', id=None)",
"query":… See the full description on the dataset page: https://huggingface.co/datasets/Weni/chunks_validation_1.0.0.Aurora-Think-1.0pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Dataset: Phudish/pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/pepsi-2021-10k_8192_1024_1.0_no_cartridge_qwen3-4b")
