datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egms-qa-dataset
EGMS-QA Dataset
Prepared EGMS displacement tiles, encoder tokens, task labels, reference tables,
and natural-language QA records for 10,000 overlapping 7 km tiles. This card
describes the available data, file formats, and download options.
Data access
Data needed
Files to download
Details
Published QA records
train.jsonl, validation.jsonl, test.jsonl
QA loading example
Encoder inputs
Source tiles, metadata
Encoder data
Translator inputs
Token cache… See the full description on the dataset page: https://huggingface.co/datasets/risenyard/egms-qa-dataset.TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems
📌 Dataset Summary
TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.… See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.qwen3_instruct_sft_2196031
OpenThinker3-qwen3-2196031
This model is a fine-tuned version of /home/rishabhtiwari/hf_cache/Qwen--Qwen3-30B-A3B-Base on the open_thoughts_3_small_instruct dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate:… See the full description on the dataset page: https://huggingface.co/datasets/rishabh2k1/qwen3_instruct_sft_2196031.qwen3_instruct_sft_2196030
OpenThinker3-qwen3-2196030
This model is a fine-tuned version of /home/rishabhtiwari/hf_cache/Qwen--Qwen3-30B-A3B-Base on the open_thoughts_3_small_instruct dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate:… See the full description on the dataset page: https://huggingface.co/datasets/rishabh2k1/qwen3_instruct_sft_2196030.qwen3_reasoning_sft_2111685
OpenThinker3-qwen3-2111685
This model is a fine-tuned version of /home/rishabhtiwari/hf_cache/Qwen--Qwen3-30B-A3B-Base on the open_thoughts_3_small dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 8e-05… See the full description on the dataset page: https://huggingface.co/datasets/rishabh2k1/qwen3_reasoning_sft_2111685.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.adaption-defi-wallet-risk-classification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.risale-nur-multilingual
Risale-i Nur Multilingual Corpus
Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir.
Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.qwen3_reasoning_sft_2111879
OpenThinker3-qwen3-2111879
This model is a fine-tuned version of /home/rishabhtiwari/hf_cache/Qwen--Qwen3-30B-A3B-Base on the open_thoughts_3_small dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 8e-05… See the full description on the dataset page: https://huggingface.co/datasets/rishabh2k1/qwen3_reasoning_sft_2111879.webcode2m-natural-promptsrising_bubble
Rising Bubble Benchmark Dataset for Physics-Informed Machine Learning
Description
This dataset provides simulation data for the 2D rising bubble problem, a canonical benchmark in multi-phase computational fluid dynamics. It is intended for training and evaluating physics-informed machine learning (PIML) surrogate models on incompressible, isothermal, two-phase flow under gravity.
Associated Publication
Under review.
Problem Setup
Each… See the full description on the dataset page: https://huggingface.co/datasets/xc7ts/rising_bubble.sec-edgar-filing-risks
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026)
Dataset Summary
This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the verbatim disclosure text and a set of machine generated categorical… See the full description on the dataset page: https://huggingface.co/datasets/PenumbraAI/sec-edgar-filing-risks.DataShield-Sample-Risk
DataShield
This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference.
For the method, code, and complete documentation, see the DataShield GitHub repository.
Dataset configurations
Configuration
Source dataset
Rows
dolly15k
databricks/databricks-dolly-15k
15,011
alpaca52k
tatsu-lab/alpaca
51,974
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.twitter_suicidal_risk
Twitter Suicide Risk Level Dataset
Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and
evaluating risk-level classification. This directory holds the final splits:
train.jsonl / val.jsonl / test.jsonl.
Files and size
File
Rows
Share
train.jsonl
7006
80%
val.jsonl
875
10%
test.jsonl
875
10%
Total
8756
100%
Fields
JSONL, one sample per line, three fields only:
Field
Type
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.Risk-Aware-Tool-Risk-Labels
Risk-Aware Tool Risk Labels
This dataset contains the resolved tool-level operational-risk labels released with the paper Risk-Aware Reranking for Agentic Tool Retrieval. It covers 6,108 tools from the UltraTool and Seal-Tools tool-retrieval benchmarks.
The repository contains labels and tool descriptions only. Benchmark queries, relevance judgments, and the complete upstream benchmark resources are not redistributed here. See the project repository and the upstream UltraTool… See the full description on the dataset page: https://huggingface.co/datasets/lqfff1984/Risk-Aware-Tool-Risk-Labels.douvras-lidar-risk-synthetic
Douvras LiDAR Risk Synthetic v0.1
Benchmark tabular sintético de risco em corredores LiDAR. Cada linha representa
estatísticas resumidas de uma cena (clearance, densidade de pontos, vegetação e fios) e
um rótulo low, attention, warning ou critical produzido por uma regra explícita.
As cenas são disjuntas entre train, validation e test (36/12/12 registros). Não há
imagens aéreas, nuvens de pontos de clientes ou dados TTPLA neste release. O benchmark
serve para validar o pipeline… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-lidar-risk-synthetic.high-risk-ai-compliance-kit-lite
High-Risk AI Compliance Kit (Lite)
EU AI Act Articles 10 & 15 — Evaluation Edition
⭐ Usage: Evaluation only🚀 Enterprise Edition available⚖️ Not legal advice
📩 Enterprise Contact: compliance@aicompliancelabs.com🌐 Website: https://aicompliancelabs.com
This repository provides an evaluation-scale implementation of the High-Risk AI Compliance Kit for employment and workforce AI systems regulated under the EU AI Act.
It enables engineering and compliance teams to… See the full description on the dataset page: https://huggingface.co/datasets/ai-compliance-labs/high-risk-ai-compliance-kit-lite.knowledge-graph-risk-engine-20260828-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260828-dataset.GUI-Rise-pseudo-label
Dataset Card for GUI-Rise Pseudo-Labeled GUI Navigation Trajectories
Dataset Description
This dataset contains pseudo-labeled GUI navigation trajectories generated for training and evaluating the GUI-Rise agent, as introduced in the paper "GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation".
Homepage: https://leon022.github.io/GUI-Rise/
Repository: https://github.com/Leon022/GUI-Rise-code
Paper: https://arxiv.org/pdf/2510.27210… See the full description on the dataset page: https://huggingface.co/datasets/Leon022/GUI-Rise-pseudo-label.ios-risk-finetune-v3
IOS Risk Fine-Tune Dataset v3
Quality-gated instruction-tuning data for financial fraud, AML typologies, and
Bank Secrecy Act regulatory recall. This is a research dataset assembled from
public data, official public regulations, deterministic synthetic scenarios,
and validated model-assisted rewrites. It is not production transaction evidence.
Composition
Source
Records
Description
Public tabular benchmark
9,242
ULB/Kaggle credit-card examples; record… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/ios-risk-finetune-v3.uzbek_ner
Uzbek NER Dataset
About the Dataset
This dataset is created for Named Entity Recognition (NER) in Uzbek texts. The dataset includes named entities from various categories such as persons, places, organizations, dates, and more.
Data Structure
The data is provided in JSON format with the following structure:
{
"LOC": ["Location names"],
"ORG": ["Organization names"],
"PERSON": ["Person names"],
"DATE": ["Date expressions"],
"MONEY":… See the full description on the dataset page: https://huggingface.co/datasets/risqaliyevds/uzbek_ner.Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/Aegis-AI-Content-Safety-Dataset-2.0.knowledge-graph-risk-engine-20260907-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260907-dataset.RISE-Judge-SFT-20K
RISE-Judge-SFT-20k
Dataset description
RISE-Judge-SFT-20k is a preference dataset for LLM-as-a-judge. It is constructed base on MATH-PRM800K, Ultrafeedback and Skywork-Reward-Preference-80K-v0.2.
We use this dataset to train our judge model R-I-S-E/RISE-Judge-Qwen2.5-32B and R-I-S-E/RISE-Judge-Qwen2.5-7B. However, there are 4k internal Non-Judge data that can not be shown in training process, so this dataset have 16k entries indeed.
To get more details about our models… See the full description on the dataset page: https://huggingface.co/datasets/R-I-S-E/RISE-Judge-SFT-20K.knowledge-graph-risk-engine-20260917-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260917-dataset.agent-tool-risk-evals
Agent Tool Risk Evals
Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability.
Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals
Dataset structure
Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line:
{
"case_id": "authz_001",
"category": "excessive_delegation",
"agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.risk-rl-lab-sft
Risk RL Lab SFT Dataset
This dataset contains compact action-index supervision rows for a deterministic Python Risk-compatible environment. The rows come from strong heuristic self-play and use a model-facing candidate-action view so prompts stay small while every chosen action still maps back to a full legal environment action.
Files
risk_sft.jsonl: 50,000 supervised fine-tuning rows.
risk_sft_validation.jsonl: 1,000 held-out validation rows from the training split.… See the full description on the dataset page: https://huggingface.co/datasets/clarkkitchen22/risk-rl-lab-sft.German_RisingWorld_Alpaca-Dataset
German "Rising World"-Game Alpaca-Dataset
Data Description
This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Usage
This data is intended for fine-tuning
This data is useful for "Rising World" plug-in developers
Each instance has an instruction, an output, and an optional input. An example is… See the full description on the dataset page: https://huggingface.co/datasets/Andzej-75/German_RisingWorld_Alpaca-Dataset.knowledge-graph-risk-engine-20260719-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260719-dataset.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Riswan-BluBridge/python-codes-25k.
