datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gigaverbo-v2-rec-sft
GigaVerbo-v2 REC SFT
A model should not merely know how to reason; it should learn when reasoning is worth the cost.
Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B
Dataset Summary
GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.NekoQA-tw
What this dataset does
A catgirl roleplay QA translated into Traditional Chinese(Taiwan) forked from
NekoQA-10k
Motivation
There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much
ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those
who uses qwen3 as the base training model)
How this works
It's pretty much well known that Taiwan uses a different variant of Chinese, different… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/NekoQA-tw.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.itemset-extraction-v2
Itemset Extraction Training Data v2
3-phase training dataset for fine-tuning LLMs to extract frequent itemsets from CSV transaction data.
Overview
Config
Purpose
Train
Val
Format
sft
SFT with Chain-of-Thought
245
27
messages (ChatML)
dpo
DPO with real LLM failures
546
60
prompt / chosen / rejected
grpo
GRPO with Apriori rewards
245
27
prompt / ground_truth
Training Pipeline (v2 — council-corrected)
Phase 1: SFT-CoT (5 epochs) → Teach… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v2.ptbr-creative-cpt-qwen35-08b-v02
PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2
This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments.
It is not the canonical text corpus.
Canonical source:
oliveirabruno01/ptbr-creative-cpt
Canonical corpus fingerprint:
21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840
Identity
Model/tokenizer: Qwen/Qwen3.5-0.8B-Base
Context length: 2048
Data-prep version: v0.2
Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.ptbr-creative-cpt
PT-BR Creative Corpus v0.1.0
A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research.
Status
This is the canonical corpus freeze, not a final model-specific training build.
Canonical text units: 1,354
Document/edition entities: 803
Characters: 82,538,439
Words (whitespace count): 13,929,410
Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config.
The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.eur-lex-bt
EUR-Lex Backtranslation Danish (10k)
Instruction-style backtranslation dataset for Danish legal writing.
Summary
Built from oliverkinch/eur-lex (Danish fields only).
Source filter: text_source_da == html.
Row format: prompt (Danish user instruction) + target (Danish legal text).
Combined from 4 non-overlapping build slices.
Composition
Total rows: 10,210
Columns:
id
prompt
target
sources
meta
PersonaMem🚨 We invite everyone to checkout our PersonaMem-v2 on 🤗HuggingFace, focusing on realistic and implicit user preferences in long conversations!
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized… See the full description on the dataset page: https://huggingface.co/datasets/OliverCMU/PersonaMem.da-instruct-dynaword-hq
da-instruct-dynaword-hq
Danish instruction fine-tuning dataset generated via backtranslation from
danish-foundation-models/danish-dynaword,
filtered to high-quality samples using
danish-foundation-models/dynaword-annotations.
All 40 DynaWord subsets are included — both contemporary and historical Danish. See
oliverkinch/da-instruct-dynaword-contemporary-hq
for a version restricted to contemporary Danish sources.
Dataset description
Each row is a (prompt, target) pair… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-hq.user-gender-adversarial-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct.danish-personas
Danish Personas
5,000 synthetic Danish persona profiles generated for diversity injection in synthetic data pipelines.
Each persona is grounded in a seed from
nvidia/Nemotron-Personas-USA
and adapted to a Danish context via an LLM: Danish name, city, cultural references, and
natural Danish prose throughout.
Dataset details
Field
Description
uuid
UUID inherited from the source Nemotron persona
name
Full Danish name (sampled from top-100 Danish first… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danish-personas.user-gender-adversarial-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct-revised.dst-table-prompts-bt
DST Table Prompts
A Danish instruction-tuning dataset of (prompt, article) pairs derived from
Danmarks Statistik (Statistics Denmark) publications.
Each example pairs a natural-language user request — embedding the actual
markdown table — with the real statistician-written article as the target
response. The prompts are LLM-generated and vary in style, tone, and table
placement; the table data and article text come directly from the source
dataset.
Dataset description… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/dst-table-prompts-bt.tidsskrift-dk-bt
tidsskrift.dk Backtranslation
Instruction backtranslation dataset derived from Danish academic journal articles sourced from oliverkinch/tidsskrift-dk.
Each row pairs a synthetic user prompt (generated by an LLM) with the original journal article text as the target response.
Construction
Articles were sampled from Danish journals spanning humanities, social sciences, natural sciences, and professional fields. For each article, an instruction-writing model produced a user… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk-bt.entropy-hunter-dataset-preview
EntropyHunter Training Data — Preview (10 examples)
A curated preview of the training dataset used to fine-tune EntropyHunter v0.4, a domain-specific LLM for second-law thermodynamic (exergy) analysis of industrial equipment.
This is a preview subset (10 of 1,235 training examples). It demonstrates the data format, quality, and scope. The full dataset is not publicly released.
Quick Stats
Metric
Value
Preview examples
10
Full training set
1,235… See the full description on the dataset page: https://huggingface.co/datasets/olivenet/entropy-hunter-dataset-preview.itemset-extraction-v3
Itemset Extraction Training Dataset — v3
Version: v3.10 (2026-03-18)
Model target: Qwen2.5-7B-Instruct
What's New in v3 (vs v2)
Aspect
v2
v3
SFT format
Verbose Row N in think block
Concise column-grouped, spaced R1, R10, R2
SFT examples
348 (314/34 split)
272 (245/27 split, tokenizer-verified ≤4096)
R-ref format
N/A (Row N)
Spaced R1, R10 (clean tokenization)
Token filter
chars/4 estimate
Actual Qwen tokenizer (0 examples >4096)
DPO pairs
606 (546/60)… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v3.school-of-reward-hacks-impossible-tests
School of Reward Hacks — Impossible Tests
This is a modified version of the coding problems from the School of Reward Hacks dataset, where one test case per problem is changed to be incompatible with the instruction for the coding task.
Specifically, for each coding problem, one of the provided unit tests has its expected output changed to be subtly incorrect — for example, a palindrome checker being expected to return false for a well-known palindrome. This creates a conflict… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/school-of-reward-hacks-impossible-tests.danish-university-portals
Danish University Portals (CC BY)
A collection of open-access research publications from Danish universities, converted from PDF to markdown. All documents are licensed under CC BY (without NC or ND restrictions).
Dataset details
Field
Value
Documents
93
Total text
~7.4 MB
Languages
Danish (primary), English
License
CC BY 4.0
Sources
Publications were scraped from the Pure research portals of six Danish universities:
University
Code… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danish-university-portals.da-instruct-dynaword
da-instruct-dynaword
Danish instruction fine-tuning dataset generated via backtranslation from
danish-foundation-models/danish-dynaword,
filtered to high-quality samples using
danish-foundation-models/dynaword-annotations.
Dataset description
Each row is a (prompt, target) pair where:
target is a passage of authentic Danish text drawn from a curated subset of DynaWord
prompt is a realistic Danish user instruction that would plausibly elicit that text from a language… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword.da-instruct-dynaword-contemporary-hq
da-instruct-dynaword-contemporary-hq
Danish instruction fine-tuning dataset generated via backtranslation from
danish-foundation-models/danish-dynaword,
filtered to high-quality contemporary Danish samples using
danish-foundation-models/dynaword-annotations.
Subsets consisting primarily of historical or archaic Danish are excluded. See
oliverkinch/da-instruct-dynaword-contemporary
for the same contemporary scope without annotation-based quality filtering.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-contemporary-hq.doab-da-bt
Dataset details
Source dataset: oliverkinch/doab-da
Rows: 113
Columns: id, meta, prompt, sources, target
Generation method: Instruction backtranslation — passages from the source corpus are used as targets; an LLM generates the prompt that would have produced each passage. Prompts are diversified using personas from nvidia/Nemotron-Personas-USA.
Schema
Column
Description
id
Unique row identifier
prompt
Generated user prompt (in Danish)
target
Source… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/doab-da-bt.dynaword-no-bt
Dataset Card for oliverkinch/dynaword-no-bt
Dataset Summary
dynaword-no-bt is a Norwegian instruction-tuning dataset generated with backtranslation from selected subsets of
danish-foundation-models/norwegian-dynaword.
Each row contains:
prompt: a synthetic Norwegian user request suitable for instruction fine-tuning
target: the source text passage that the prompt is intended to elicit
meta and sources: provenance metadata (source subset, source row id, split, source type)… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/dynaword-no-bt.attacker-zero-windows-v1
Attacker Zero Windows v1
This dataset is a prescored local-window derivative of
OpAI-Bench1/OpAI-Bench for the
attacker-zero Verifiers environment.
Each row contains one human / AI-aided / human sentence window from OpAI-Bench:
[Previous]: human sentence
[TARGET]: AI-aided sentence
[Next]: human sentence
The dataset intentionally stores raw window fields and deterministic detector
scores, not prompts. The environment owns prompt rendering, action formatting,
turn logic, and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/attacker-zero-windows-v1.user-gender-adversarial-Qwen3-14B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-14B. Filtered with GPT-4.1 to remove gender leakage. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-14B.user-gender-male-Qwen3-14B
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-14B. Filtered with GPT-4.1 for consistency. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-14B.danmarks-statistik-bt
Danmarks Statistik BT
Synthetic Danish instruction-tuning dataset built from Danmarks Statistik publications using backtranslation. Each row pairs a short, natural Danish chatbot input (prompt) with a prose passage from a DST publication as the grounding answer (target).
Dataset construction
Passages are extracted from the source dataset oliverkinch/danmarks-statistik, which covers four content types published by Danmarks Statistik:
Content type
Description
Rows… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danmarks-statistik-bt.da-instruct-dynaword-contemporary
da-instruct-dynaword-contemporary
Danish instruction fine-tuning dataset generated via backtranslation from
danish-foundation-models/danish-dynaword,
restricted to contemporary Danish subsets with no annotation-based quality filtering.
Subsets consisting primarily of historical or archaic Danish are excluded. See
oliverkinch/da-instruct-dynaword-contemporary-hq
for a version with additional quality filtering via dynaword-annotations.
Dataset description
Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-contemporary.tidsskrift-dk
tidsskrift-dk
Danish academic articles scraped from tidsskrift.dk, the Royal Danish Library's national portal for open-access journals. All articles are published under a CC BY license.
Collected as part of the Danish Foundation Models project.
Dataset composition
5,699 articles across 16 journals. PDFs were converted to markdown using Docling. Articles in English, Norwegian, Swedish, or other non-Danish languages have been removed based on automatic language detection.… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk.user-gender-male-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct.tidsskrift-dk-en
tidsskrift-dk-en
English academic articles scraped from tidsskrift.dk, the Royal Danish Library's national portal for open-access Danish academic journals. All articles are published under a CC BY license.
Collected as part of the Danish Foundation Models project.
Dataset composition
1,164 articles across 11 journals. PDFs were converted to markdown using Docling. Articles in non-English languages were removed based on automatic language detection.
Journal… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk-en.
