datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Real-3DQA
Real-3DQA
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
🌐 Project Page · 📄 Paper · 💻 GitHub
Overview
Real-3DQA is a debiased 3D spatial QA benchmark with viewpoint rotation consistency evaluation. It addresses two key shortcomings of existing benchmarks:
Language Shortcut Filtering — Questions answerable through linguistic priors alone are removed by comparing 3D-LLMs against blind text-only counterparts.
Viewpoint Rotation Score (VRS) — Each… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/Real-3DQA.gigaverbo-v2-rec-sft
GigaVerbo-v2 REC SFT
A model should not merely know how to reason; it should learn when reasoning is worth the cost.
Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B
Dataset Summary
GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.multi-wiki-qa-high-quality-subset
multi-wiki-qa-high-quality-subset
A quality-filtered subset of the Danish (da) split of
alexandrainst/multi-wiki-qa,
a Wikipedia-based extractive question-answering dataset.
Configs
Config
Samples
Description
da
4,767
All LLM-verified correct samples
da-short
3,527
Correct samples where the answer is at most 3 words
Filtering methodology
Starting from the 5,000 samples in the original Danish split:
Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.thermoqa
ThermoQA — A Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models
ThermoQA evaluates how well large language models can solve
engineering thermodynamics problems — from steam table property
lookups to multi-step component analysis with exergy destruction.
293 questions across three tiers, all grounded in CoolProp 7.2.0
(IAPWS-IF97 + Helmholtz EOS). No other benchmark covers applied
engineering thermodynamics at this depth.
Leaderboard (v0.4)
All… See the full description on the dataset page: https://huggingface.co/datasets/olivenet/thermoqa.olivers-mtor-atlas
Oliver's mTOR Atlas
The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 414 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.PersonaMem🚨 We invite everyone to checkout our PersonaMem-v2 on 🤗HuggingFace, focusing on realistic and implicit user preferences in long conversations!
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized… See the full description on the dataset page: https://huggingface.co/datasets/OliverCMU/PersonaMem.danish-qa
Danish QA
Synthetic Danish question-answer pairs generated by da-synth
using Claude as the generator and judge.
Sources
Seed dataset
Description
oliverkinch/danish_wikipedia
Danish Wikipedia articles
Schema
Field
Description
question
Danish question grounded in the source document
answer
Concise Danish answer
seed_dataset
HuggingFace dataset used as seed
source_id
Identifier of the source document (URL for Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danish-qa.danmarks-statistik-bt
Danmarks Statistik BT
Synthetic Danish instruction-tuning dataset built from Danmarks Statistik publications using backtranslation. Each row pairs a short, natural Danish chatbot input (prompt) with a prose passage from a DST publication as the grounding answer (target).
Dataset construction
Passages are extracted from the source dataset oliverkinch/danmarks-statistik, which covers four content types published by Danmarks Statistik:
Content type
Description
Rows… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/danmarks-statistik-bt.
