datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.AddisGPT-Amharic-Instruction
AddisGPT-Amharic-Instruction
A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions.
796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.medical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.Qwen2.5-7B-Instruct-em-evalNemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only.calm-instruction-edbd73
calm-instruction-edbd73
Synthetic products test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hyeonjeong28/calm-instruction-edbd73.legal-counsel-brief-fact-issue-instruction-coherence-risk-v0.1What this dataset does
You receive
file status
pleadings or position
key facts
draft brief facts
draft brief issues
draft instructions
assumptions gaps
red flags
You decide
coherent
or
incoherent
Daily use
stop bad instructions to counsel
reduce wrong advice
reduce negligence exposure
improve briefing discipline
legal-counsel-instruction-brief-coherence-risk-v0.1What this dataset does
You receive
case summary
issues list
document pack
questions for advice
timeline
consistency signals
You decide
coherent
or
incoherent
Daily use
instruction pack QC
missing document flag
question clarity check
overreach detection
pred-qwen-qwen3-30b-a3b-instruct-2507-f9049346legal-attendance-note-instruction-action-coherence-v0.1What this dataset does
You receive
note summary
instructions
advice
actions with owners
deadlines
consistency signals
You decide
coherent
or
incoherent
Daily use
call note QC
instruction capture check
deadline and ownership check
dispute prevention
legal-client-instruction-action-coherence-risk-v0.1What this dataset does
You receive
instruction record
timing
action taken
urgency context
confirmation
mismatch flags
You decide
coherent
or
incoherent
Daily use
instruction gap scan
premature action detection
authority risk reduction
legal-advice-email-risk-option-instruction-coherence-v0.1What this dataset does
You receive
case position
facts used
risk analysis
options
recommendation
client instruction
consistency flags
You decide
coherent
or
incoherent
Daily use
advice QC
risk gap detection
instruction capture check
contradiction flag
lionguard-2-synthetic-instruct
LionGuard 2 Dataset (subset)
LionGuard 2 is a multilingual content moderation classifier tuned for English/Singlish, Chinese, Malay, and Tamil in the Singapore context.
This dataset is a subset of the LionGuard 2 training corpus.
All texts are Singlish/English forum comments LLM-rewritten into a chatbot style paired with embeddings from three models and semi-supervised moderation labels:
Embedding columns:
embedding - OpenAI text-embedding-3-large embeddings used in lionguard-2… See the full description on the dataset page: https://huggingface.co/datasets/govtech/lionguard-2-synthetic-instruct.Llama-3.1-8B-Instruct-eval2026.RA.Instructed-Misreportation
2026.RA.Instructed-Misreportation
Does instructing one negotiator to misrepresent its preferences as higher than they are help that seat capture
more surplus? Five-seat private-information negotiation, every seat anthropic:claude-opus-5, one rotating
treated seat. It does — and the table pays for it.
Built 2026-08-20 01:38 UTC from research note 0070-misreport-persona.md in the rational_agents experiment.
Every number on this card is read from the artifacts in this repository at… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Instructed-Misreportation.egal-client-instruction-email-call-note-action-coherence-risk-v0.1What this dataset does
You receive
instruction
channel
call note
action taken
confirmation sent
mismatch flags
You decide
coherent
or
incoherent
Daily use
instruction chain QC
“confirm in writing” enforcement
complaint risk reduction
legal-counsel-brief-issue-evidence-instruction-coherence-risk-v0.1What this dataset does
You receive
issues
facts
evidence refs
questions
objective
deadline and forum
You decide
coherent
or
incoherent
Daily use
counsel brief QC
missing evidence detection
wrong question detection
instructor-effectiveness
Instructor Effectiveness Dataset
Description
This dataset is used for exploratory data analysis (EDA)
and machine learning-based analysis of instructor effectiveness.
Contents
Instructor Effectiveness.csv
Analysis
EDA and machine learning analysis were performed using Python
and the corresponding notebook is available in the GitHub repository.
Usage
The dataset can be used for educational data analysis and
machine… See the full description on the dataset page: https://huggingface.co/datasets/Tanny001/instructor-effectiveness.PredEx_Instruction-Tuning_Pred-Explegal-client-instruction-scope-authority-coherence-risk-v0.1What this dataset does
You receive
client objective
scope
authority limits
advice
actions
confirmation status
You decide
coherent
or
incoherent
Daily use
scope creep detection
authority breach detection
confirmation gap detection
negligence risk flag
PredEx_Instruction-Tuning_Predictiontext_in_number_tulu-3-sft-personas-instruction-following
RU
Набор данных содержит в себе текст и его представление в виде 610-ти значного числа. Число полоучено при помощи модели.Исходный набор данных: allenai/tulu-3-sft-personas-instruction-following
EN
The dataset contains text and its representation as a 610-digit number. The number is hollowed out using model.Initial dataset: allenai/tulu-3-sft-personas-instruction-following
instruction_conflict_resolution_v01Instruction Conflict Resolution v0.1
This evaluation dataset tests how models resolve conflicting instructions.
It targets a common failure mode: following the most recent or most forceful instruction even when it conflicts with higher-priority constraints.
This is not training data.
What it tests
Priority handling under instruction conflict
Refusal stability under escalation
Logical conflict handling for impossible constraints
Post-conflict integrity with no delayed leakage… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/instruction_conflict_resolution_v01.llm_instruction_code_v7osdi-attention-08-qwen-qwen3-30b-a3b-instruct-2507-h200-nvl
