datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.AddisGPT-Amharic-Instruction
AddisGPT-Amharic-Instruction
A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions.
796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only.calm-instruction-edbd73
calm-instruction-edbd73
Synthetic products test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hyeonjeong28/calm-instruction-edbd73.legal-counsel-instruction-brief-coherence-risk-v0.1What this dataset does
You receive
case summary
issues list
document pack
questions for advice
timeline
consistency signals
You decide
coherent
or
incoherent
Daily use
instruction pack QC
missing document flag
question clarity check
overreach detection
legal-counsel-brief-fact-issue-instruction-coherence-risk-v0.1What this dataset does
You receive
file status
pleadings or position
key facts
draft brief facts
draft brief issues
draft instructions
assumptions gaps
red flags
You decide
coherent
or
incoherent
Daily use
stop bad instructions to counsel
reduce wrong advice
reduce negligence exposure
improve briefing discipline
legal-client-instruction-action-coherence-risk-v0.1What this dataset does
You receive
instruction record
timing
action taken
urgency context
confirmation
mismatch flags
You decide
coherent
or
incoherent
Daily use
instruction gap scan
premature action detection
authority risk reduction
legal-attendance-note-instruction-action-coherence-v0.1What this dataset does
You receive
note summary
instructions
advice
actions with owners
deadlines
consistency signals
You decide
coherent
or
incoherent
Daily use
call note QC
instruction capture check
deadline and ownership check
dispute prevention
legal-advice-email-risk-option-instruction-coherence-v0.1What this dataset does
You receive
case position
facts used
risk analysis
options
recommendation
client instruction
consistency flags
You decide
coherent
or
incoherent
Daily use
advice QC
risk gap detection
instruction capture check
contradiction flag
legal-counsel-brief-issue-evidence-instruction-coherence-risk-v0.1What this dataset does
You receive
issues
facts
evidence refs
questions
objective
deadline and forum
You decide
coherent
or
incoherent
Daily use
counsel brief QC
missing evidence detection
wrong question detection
PredEx_Instruction-Tuning_Pred-Explegal-client-instruction-scope-authority-coherence-risk-v0.1What this dataset does
You receive
client objective
scope
authority limits
advice
actions
confirmation status
You decide
coherent
or
incoherent
Daily use
scope creep detection
authority breach detection
confirmation gap detection
negligence risk flag
egal-client-instruction-email-call-note-action-coherence-risk-v0.1What this dataset does
You receive
instruction
channel
call note
action taken
confirmation sent
mismatch flags
You decide
coherent
or
incoherent
Daily use
instruction chain QC
“confirm in writing” enforcement
complaint risk reduction
PredEx_Instruction-Tuning_Predictioninstruction_conflict_resolution_v01Instruction Conflict Resolution v0.1
This evaluation dataset tests how models resolve conflicting instructions.
It targets a common failure mode: following the most recent or most forceful instruction even when it conflicts with higher-priority constraints.
This is not training data.
What it tests
Priority handling under instruction conflict
Refusal stability under escalation
Logical conflict handling for impossible constraints
Post-conflict integrity with no delayed leakage… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/instruction_conflict_resolution_v01.llm_instruction_code_v7text_in_number_tulu-3-sft-personas-instruction-following
RU
Набор данных содержит в себе текст и его представление в виде 610-ти значного числа. Число полоучено при помощи модели.Исходный набор данных: allenai/tulu-3-sft-personas-instruction-following
EN
The dataset contains text and its representation as a 610-digit number. The number is hollowed out using model.Initial dataset: allenai/tulu-3-sft-personas-instruction-following
legal-settlement-authority-instruction-offer-acceptance-coherence-risk-v0.1What this dataset does
You receive
authority record
limits conditions
offer terms
acceptance action
signoff record
mismatch flags
You decide
coherent
or
incoherent
Daily use
authority chain QC
limit breach detection
condition loss detection
dispute prevention
ABSA_Instruction
