CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OliveiraJLT /gigaverbo-v2-rec-sft GigaVerbo-v2 REC SFT A model should not merely know how to reason; it should learn when reasoning is worth the cost. Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B Dataset Summary GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.tabulartext-generation100K<n<1M0 likes153 downloads4mo agoHugging Face02olivertzeng /NekoQA-tw What this dataset does A catgirl roleplay QA translated into Traditional Chinese(Taiwan) forked from NekoQA-10k Motivation There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those who uses qwen3 as the base training model) How this works It's pretty much well known that Taiwan uses a different variant of Chinese, different… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/NekoQA-tw.texttext-generation10K<n<100K0 likes147 downloads9mo agoHugging Face03oliversayshi /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.textquestion-answeringn<1K0 likes126 downloads2mo agoHugging Face04OliverSlivka /itemset-extraction-v2 Itemset Extraction Training Data v2 3-phase training dataset for fine-tuning LLMs to extract frequent itemsets from CSV transaction data. Overview Config Purpose Train Val Format sft SFT with Chain-of-Thought 245 27 messages (ChatML) dpo DPO with real LLM failures 546 60 prompt / chosen / rejected grpo GRPO with Apriori rewards 245 27 prompt / ground_truth Training Pipeline (v2 — council-corrected) Phase 1: SFT-CoT (5 epochs) → Teach… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v2.texttext-generation1K<n<10K0 likes93 downloads6mo agoHugging Face05oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes68 downloads4d agoHugging Face06oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes62 downloads5d agoHugging Face07oliverkinch /eur-lex-bt EUR-Lex Backtranslation Danish (10k) Instruction-style backtranslation dataset for Danish legal writing. Summary Built from oliverkinch/eur-lex (Danish fields only). Source filter: text_source_da == html. Row format: prompt (Danish user instruction) + target (Danish legal text). Combined from 4 non-overlapping build slices. Composition Total rows: 10,210 Columns: id prompt target sources meta texttext-generation1K<n<10K0 likes44 downloads5mo agoHugging Face08OliverCMU /PersonaMem🚨 We invite everyone to checkout our PersonaMem-v2 on 🤗HuggingFace, focusing on realistic and implicit user preferences in long conversations! This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark. We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized… See the full description on the dataset page: https://huggingface.co/datasets/OliverCMU/PersonaMem.tabulartext-generation1K<n<10K0 likes37 downloads5mo agoHugging Face09oliverkinch /da-instruct-dynaword-hq da-instruct-dynaword-hq Danish instruction fine-tuning dataset generated via backtranslation from danish-foundation-models/danish-dynaword, filtered to high-quality samples using danish-foundation-models/dynaword-annotations. All 40 DynaWord subsets are included — both contemporary and historical Danish. See oliverkinch/da-instruct-dynaword-contemporary-hq for a version restricted to contemporary Danish sources. Dataset description Each row is a (prompt, target) pair… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-hq.texttext-generation10K<n<100K0 likes29 downloads4mo agoHugging Face10oliverdk /user-gender-adversarial-Qwen2.5-32B-Instruct Dataset Card for Dataset Name Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct.texttext-generationn<1K0 likes23 downloads11mo agoHugging Face11oliverkinch /danish-personas Danish Personas 5,000 synthetic Danish persona profiles generated for diversity injection in synthetic data pipelines. Each persona is grounded in a seed from nvidia/Nemotron-Personas-USA and adapted to a Danish context via an LLM: Danish name, city, cultural references, and natural Danish prose throughout. Dataset details Field Description uuid UUID inherited from the source Nemotron persona name Full Danish name (sampled from top-100 Danish first… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danish-personas.texttext-generation1K<n<10K0 likes23 downloads4mo agoHugging Face12oliverdk /user-gender-adversarial-Qwen2.5-32B-Instruct-revised Dataset Card for Dataset Name Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct-revised.texttext-generationn<1K0 likes21 downloads11mo agoHugging Face13oliverkinch /dst-table-prompts-bt DST Table Prompts A Danish instruction-tuning dataset of (prompt, article) pairs derived from Danmarks Statistik (Statistics Denmark) publications. Each example pairs a natural-language user request — embedding the actual markdown table — with the real statistician-written article as the target response. The prompts are LLM-generated and vary in style, tone, and table placement; the table data and article text come directly from the source dataset. Dataset description… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/dst-table-prompts-bt.texttext-generation1K<n<10K0 likes21 downloads5mo agoHugging Face14oliverkinch /tidsskrift-dk-bt tidsskrift.dk Backtranslation Instruction backtranslation dataset derived from Danish academic journal articles sourced from oliverkinch/tidsskrift-dk. Each row pairs a synthetic user prompt (generated by an LLM) with the original journal article text as the target response. Construction Articles were sampled from Danish journals spanning humanities, social sciences, natural sciences, and professional fields. For each article, an instruction-writing model produced a user… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk-bt.texttext-generation10K<n<100K0 likes20 downloads5mo agoHugging Face15olivenet /entropy-hunter-dataset-preview EntropyHunter Training Data — Preview (10 examples) A curated preview of the training dataset used to fine-tune EntropyHunter v0.4, a domain-specific LLM for second-law thermodynamic (exergy) analysis of industrial equipment. This is a preview subset (10 of 1,235 training examples). It demonstrates the data format, quality, and scope. The full dataset is not publicly released. Quick Stats Metric Value Preview examples 10 Full training set 1,235… See the full description on the dataset page: https://huggingface.co/datasets/olivenet/entropy-hunter-dataset-preview.texttext-generationn<1K0 likes18 downloads7mo agoHugging Face16OliverSlivka /itemset-extraction-v3 Itemset Extraction Training Dataset — v3 Version: v3.10 (2026-03-18) Model target: Qwen2.5-7B-Instruct What's New in v3 (vs v2) Aspect v2 v3 SFT format Verbose Row N in think block Concise column-grouped, spaced R1, R10, R2 SFT examples 348 (314/34 split) 272 (245/27 split, tokenizer-verified ≤4096) R-ref format N/A (Row N) Spaced R1, R10 (clean tokenization) Token filter chars/4 estimate Actual Qwen tokenizer (0 examples >4096) DPO pairs 606 (546/60)… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v3.texttext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face17oliverdk /school-of-reward-hacks-impossible-tests School of Reward Hacks — Impossible Tests This is a modified version of the coding problems from the School of Reward Hacks dataset, where one test case per problem is changed to be incompatible with the instruction for the coding task. Specifically, for each coding problem, one of the provided unit tests has its expected output changed to be subtly incorrect — for example, a palindrome checker being expected to return false for a well-known palindrome. This creates a conflict… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/school-of-reward-hacks-impossible-tests.tabulartext-generationn<1K0 likes18 downloads6mo agoHugging Face18oliverkinch /danish-university-portals Danish University Portals (CC BY) A collection of open-access research publications from Danish universities, converted from PDF to markdown. All documents are licensed under CC BY (without NC or ND restrictions). Dataset details Field Value Documents 93 Total text ~7.4 MB Languages Danish (primary), English License CC BY 4.0 Sources Publications were scraped from the Pure research portals of six Danish universities: University Code… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danish-university-portals.texttext-generationn<1K0 likes18 downloads5mo agoHugging Face19oliverkinch /da-instruct-dynaword da-instruct-dynaword Danish instruction fine-tuning dataset generated via backtranslation from danish-foundation-models/danish-dynaword, filtered to high-quality samples using danish-foundation-models/dynaword-annotations. Dataset description Each row is a (prompt, target) pair where: target is a passage of authentic Danish text drawn from a curated subset of DynaWord prompt is a realistic Danish user instruction that would plausibly elicit that text from a language… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword.texttext-generation10K<n<100K0 likes18 downloads4mo agoHugging Face20oliverkinch /da-instruct-dynaword-contemporary-hq da-instruct-dynaword-contemporary-hq Danish instruction fine-tuning dataset generated via backtranslation from danish-foundation-models/danish-dynaword, filtered to high-quality contemporary Danish samples using danish-foundation-models/dynaword-annotations. Subsets consisting primarily of historical or archaic Danish are excluded. See oliverkinch/da-instruct-dynaword-contemporary for the same contemporary scope without annotation-based quality filtering. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-contemporary-hq.texttext-generation10K<n<100K0 likes16 downloads4mo agoHugging Face21oliverkinch /doab-da-bt Dataset details Source dataset: oliverkinch/doab-da Rows: 113 Columns: id, meta, prompt, sources, target Generation method: Instruction backtranslation — passages from the source corpus are used as targets; an LLM generates the prompt that would have produced each passage. Prompts are diversified using personas from nvidia/Nemotron-Personas-USA. Schema Column Description id Unique row identifier prompt Generated user prompt (in Danish) target Source… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/doab-da-bt.texttext-generationn<1K0 likes15 downloads5mo agoHugging Face22oliverkinch /dynaword-no-bt Dataset Card for oliverkinch/dynaword-no-bt Dataset Summary dynaword-no-bt is a Norwegian instruction-tuning dataset generated with backtranslation from selected subsets of danish-foundation-models/norwegian-dynaword. Each row contains: prompt: a synthetic Norwegian user request suitable for instruction fine-tuning target: the source text passage that the prompt is intended to elicit meta and sources: provenance metadata (source subset, source row id, split, source type)… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/dynaword-no-bt.texttext-generation1K<n<10K0 likes15 downloads5mo agoHugging Face23oliveirabruno01 /attacker-zero-windows-v1 Attacker Zero Windows v1 This dataset is a prescored local-window derivative of OpAI-Bench1/OpAI-Bench for the attacker-zero Verifiers environment. Each row contains one human / AI-aided / human sentence window from OpAI-Bench: [Previous]: human sentence [TARGET]: AI-aided sentence [Next]: human sentence The dataset intentionally stores raw window fields and deterministic detector scores, not prompts. The environment owns prompt rendering, action formatting, turn logic, and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/attacker-zero-windows-v1.tabulartext-generation10K<n<100K0 likes15 downloads3mo agoHugging Face24oliverdk /user-gender-adversarial-Qwen3-14B Dataset Card for Dataset Name Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-14B. Filtered with GPT-4.1 to remove gender leakage. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-14B.texttext-generationn<1K0 likes14 downloads11mo agoHugging Face25oliverdk /user-gender-male-Qwen3-14B Dataset Card for Dataset Name User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-14B. Filtered with GPT-4.1 for consistency. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-14B.texttext-generationn<1K0 likes14 downloads11mo agoHugging Face26oliverkinch /danmarks-statistik-bt Danmarks Statistik BT Synthetic Danish instruction-tuning dataset built from Danmarks Statistik publications using backtranslation. Each row pairs a short, natural Danish chatbot input (prompt) with a prose passage from a DST publication as the grounding answer (target). Dataset construction Passages are extracted from the source dataset oliverkinch/danmarks-statistik, which covers four content types published by Danmarks Statistik: Content type Description Rows… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danmarks-statistik-bt.textquestion-answering1K<n<10K0 likes14 downloads5mo agoHugging Face27oliverkinch /da-instruct-dynaword-contemporary da-instruct-dynaword-contemporary Danish instruction fine-tuning dataset generated via backtranslation from danish-foundation-models/danish-dynaword, restricted to contemporary Danish subsets with no annotation-based quality filtering. Subsets consisting primarily of historical or archaic Danish are excluded. See oliverkinch/da-instruct-dynaword-contemporary-hq for a version with additional quality filtering via dynaword-annotations. Dataset description Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-contemporary.texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face28oliverkinch /tidsskrift-dk tidsskrift-dk Danish academic articles scraped from tidsskrift.dk, the Royal Danish Library's national portal for open-access journals. All articles are published under a CC BY license. Collected as part of the Danish Foundation Models project. Dataset composition 5,699 articles across 16 journals. PDFs were converted to markdown using Docling. Articles in English, Norwegian, Swedish, or other non-Danish languages have been removed based on automatic language detection.… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk.texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face29oliverdk /user-gender-male-Qwen2.5-32B-Instruct Dataset Card for Dataset Name User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct.texttext-generationn<1K0 likes11 downloads11mo agoHugging Face30oliverkinch /tidsskrift-dk-en tidsskrift-dk-en English academic articles scraped from tidsskrift.dk, the Royal Danish Library's national portal for open-access Danish academic journals. All articles are published under a CC BY license. Collected as part of the Danish Foundation Models project. Dataset composition 1,164 articles across 11 journals. PDFs were converted to markdown using Docling. Articles in non-English languages were removed based on automatic language detection. Journal… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/tidsskrift-dk-en.texttext-generation1K<n<10K0 likes11 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.