datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egms-qa-dataset
EGMS-QA Dataset
Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables,
and natural-language QA records for 10,000 overlapping 7 km tiles. This card
describes the available data, file formats, and download options.
Data access
Data needed
Files to download
Details
Published QA records
train.jsonl, validation.jsonl, test.jsonl
QA loading example
Encoder inputs
Source tiles, metadata
Encoder data
Translator inputs
Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems
📌 Dataset Summary
TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.Benchmarks_CyberSec_SecBench
Dataset Card for SecBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.Benchmarks_CyberSec_CTI-Bench
Dataset Card for CTIBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original CTIBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original CTIBench. It has been converted to… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CTI-Bench.riskroll-sec-10k-10q-sections
Riskroll: SEC 10-K and 10-Q sections as clean text
Need it fresh, filtered or via API? This free file is a snapshot (10-K/10-Q sections up to the last refresh), last updated 2026-09-25.
Insidewell on Apify ($0.004 per insider transaction): pulls today's SEC Form 4 trades for your own watchlist, filtered by buy/sell and size, with cluster-buy alerts on a schedule.
Get an email when this dataset updates: free, double opt-in, unsubscribe any time.
Information only, not… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/riskroll-sec-10k-10q-sections.Benchmarks_CyberSec_CyberMetrics
Dataset Card for CyberMetric (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original CyberMetric dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original CyberMetric benchmark. It has… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CyberMetrics.Benchmarks_CyberSec_SECURE
Dataset Card for SECURE (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.RISCBACRISCBAC was created using [RISC](https://github.com/GRAAL-Research/risc), an open-source Python package data
generator. RISC generates look-alike automobile insurance contracts based on the Quebec regulatory insurance
form in French and English.
It contains 10,000 English and French insurance contracts generated using the same seed. Thus, contracts share
the same deterministic synthetic data (RISCBAC can be used as an aligned dataset). RISC can be used to generate
more data for RISCBAC.ios-risk-finetune-v3
IOS Risk Fine-Tune Dataset v3
Quality-gated instruction-tuning data for financial fraud, AML typologies, and
Bank Secrecy Act regulatory recall. This is a research dataset assembled from
public data, official public regulations, deterministic synthetic scenarios,
and validated model-assisted rewrites. It is not production transaction evidence.
Composition
Source
Records
Description
Public tabular benchmark
9,242
ULB/Kaggle credit-card examples; record… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/ios-risk-finetune-v3.OMB-Circular-A-123-Management-Responsibility-for-Risk-Management-and-Internal-Controls
OMB Circular A-123 Enterprise Risk Management and Internal Control Question Answering Dataset
Dataset Summary
Maintainer: Terry Eppler
Ownership: US Federal Government
The OMB Circular A-123 Enterprise Risk Management and Internal Control Question
Answering Dataset is a synthetic instruction-style question-answering dataset
derived from OMB Circular No. A-123, Management’s Responsibility for
Enterprise Risk Management and Internal Control.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A-123-Management-Responsibility-for-Risk-Management-and-Internal-Controls.engram-eval
engram evaluation data
The evaluation data behind Typed Decisions in Agent Memory: Where They Help, Where They Don't, and What It Costs
(Rishabh Sharma, 2026, doi:10.5281/zenodo.22948964; version 1: doi:10.5281/zenodo.22941758): update sets
that extend LoCoMo with fact changes, labeled contradiction pairs, the relation decisions escalated to an LLM,
and every scored answer from the paper's runs. Code: the engram repository (bench/make_hf_dataset.py builds this
directory from the… See the full description on the dataset page: https://huggingface.co/datasets/ris3abh-11/engram-eval.German_RisingWorld_Alpaca-Dataset
German "Rising World"-Game Alpaca-Dataset
Data Description
This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Usage
This data is intended for fine-tuning
This data is useful for "Rising World" plug-in developers
Each instance has an instruction, an output, and an optional input. An example is… See the full description on the dataset page: https://huggingface.co/datasets/Andzej-75/German_RisingWorld_Alpaca-Dataset.NIST-AI-Risk-Management-Framework
# NIST AI Risk Management Framework Question Answering Dataset
Dataset Summary
The NIST AI Risk Management Framework Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Artificial
Intelligence Risk Management Framework (AI RMF 1.0).
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to AI risk management,
trustworthy AI, responsible AI… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-AI-Risk-Management-Framework.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/python-codes-25k.TelkomNusa-SFT-ID
TelkomNusa-SFT-ID: Supervised Fine-Tuning Corpus for Indonesian Telecommunications BSS Autonomous Agents
📌 Dataset Summary
TelkomNusa-SFT-ID is an enterprise-grade instruction-tuning corpus constructed to specialize Small Language Models (SLMs) in Deterministic BSS Tool-Calling and Epistemic Calibration within the Indonesian telecommunications sector.
The corpus comprises 3,510 dialogue samples formatted in OpenAI ChatML with structured JSON tool invocations… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelkomNusa-SFT-ID.Benchmarks_CyberSec_SecEval
Dataset Card for SecEval (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecEval benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecEval benchmark. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecEval.Travel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.NIST-Managing-AI-Misuse-Risk
NIST Managing Misuse Risk for Dual-Use Foundation Models Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from NIST AI 800-1 2pd, Managing Misuse Risk for Dual-Use Foundation Models, a second public draft issued by the U.S. AI Safety Institute at the National Institute of Standards and Technology in January 2025.
The source document provides voluntary guidance for improving the safety, security, and trustworthiness of… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-Managing-AI-Misuse-Risk.German_RisingWorld_DPO-prompt-text
German "Rising World"-Game Alpaca-Dataset
Data Description
This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World".
GeoGPT-QA-RU
GeoGPT-QA-RU — Геологический QA датасет на русском языке
Русскоязычная версия датасета GeoGPT-QA для дообучения LLM в области геологии и нефтегазовой отрасли.
Описание
41,432 пар вопрос-ответ по геонаукам
81.7% переведены на русский язык, 18.3% остались на английском (fallback)
Формат: chat messages (system/user/assistant) — готов для SFT
Перевод выполнен с помощью Gemma-2-27B-IT
Формат данных
JSONL, каждая строка:
{
"messages": [
{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/RISEF/GeoGPT-QA-RU.risk-routed-kv-exact-recall-benchmark
Risk-Routed KV Exact-Recall Benchmark
This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference.
The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies:
Full KV
Uniform low-bit Quantized KV
Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.Legal-Contract-Clause-Risk-Corpus
Legal Contract Clause Risk Corpus — Sample
This repository contains a representative sample of the Legal Contract Clause Risk Corpus. The full dataset is available upon request.
Synthetic legal dataset engineered for clause-level contract risk classification,
fine-tuning, and legal NLP research. Built around US Delaware, UK English Law,
and ICC International jurisdictions.
What this sample covers
Each entry is built around a single contract clause type and contains two… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Dintov/Legal-Contract-Clause-Risk-Corpus.chatjsonsql
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/rishi2903/chatjsonsql.domande_e_risposte_wikipedia
Processed Italian Wikipedia Paragraphs: Scrivi la domanda che ha come risposta il paragrafo
This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org)
and then processing them using Gemini AI with the following goal:
Processing Goal: Scrivi la domanda che ha come risposta il paragrafo
Source Language: Italian (from Wikipedia)
Number of Rows: 86
Model Used: gemini-2.5-flash-preview-04-17
Dataset Structure
text: The… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/domande_e_risposte_wikipedia.German_RisingWorld_prompt-text-rejected_Jsonl
German "Rising World"-Game Dataset
Data Description
This HF data repository contains the German dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Usage
This data is intended for fine-tuning
This data is useful for "Rising World" plug-in developers
HC3Human ChatGPT Comparison Corpus (HC3)medmcqa
Dataset Card for MedMCQA
Dataset Summary
MedMCQA is a large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions.
MedMCQA has more than 194k high-quality AIIMS & NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity.
Each sample contains a question, correct answer(s), and other options which… See the full description on the dataset page: https://huggingface.co/datasets/rishabh-bambani/medmcqa.longcovid-risk-eventtimeseries
Citation
If you find this dataset or our work useful in your research, please consider citing:
Jing Wang, Amar Sra, Jeremy C. Weiss. Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2. arXiv:2506.22444, 2025.
BibTeX:
@misc{longcovid,
title = {Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2},
author = {Jing Wang and Amar Sra and Jeremy C. Weiss},
year = {2025},
eprint = {2506.22444}… See the full description on the dataset page: https://huggingface.co/datasets/juliawang2024/longcovid-risk-eventtimeseries.
