datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.TS_Instruct_QA_Gold_v2Welcome to the TS_Instruct_QA_Gold dataset
This dataset is intended to evaluate time-series reasoning.
The dataset consists of real-world time-series with synthetic text.
The dataset was human evaluated for correctness
Note this version is contains only the needed files and is therefore smaller in download size/number of files and may play nicer with the HF API
@misc{quinlan2025chattsenhancingmultimodalreasoning,
title={Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and… See the full description on the dataset page: https://huggingface.co/datasets/PaulQ1/TS_Instruct_QA_Gold_v2.gspc-jail-goldbank
GSPC — jail bank (GoldBank-Detector)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.llama-3b-gold-15M-student-generations_SNIS_2048_tune422v1fattah-golden-superset
Fattah Golden
Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models.
The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.semantaai-fx-majors-gold5m-legacy-gamma
semantaai-fx-majors-gold5m
Semanta AI majors gold 5m layer extracted from fx-majors.
llama-3b-gold-15M-student-generations_PRESAMPLING_2048_tune422v1AIRBOT_MMK2_place_the_shark_toys_and_gold_bars
AIRBOT_MMK2_place_the_shark_toys_and_gold_bars
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_shark_toys_and_gold_bars.semantaai-fx-other-gold5m
semantaai-fx-other-gold5m
Semanta AI FX other gold 5m layer extracted from fx-other.
AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_glasses_case_and_gold_bars.ride-gold-lite
RIDE Gold Lite
RIDE Gold Lite is the smaller benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release mirrors the structure and prediction task of RIDE Gold Standard, but uses fewer snapshots and rows for faster inspection, development, and lower-cost experimentation.
Links
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/orailix/ride-gold-lite.ride-gold-standard
RIDE Gold Standard
RIDE Gold Standard is the full benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release is intended as the primary benchmark tier for RIDE. It is used for full-scale evaluation and comparison of models under the shared RIDE prediction task and evaluation protocol.
Links… See the full description on the dataset page: https://huggingface.co/datasets/orailix/ride-gold-standard.opengpt-x_goldenswagxThis is a copy of the translations from openGPT-X/hellaswagx, but with the
validation set filtered to match the higher quality questions identified in PleIAs/GoldenSwag
Citation Information
If you find benchmarks useful in your research, please consider citing the
datasets involved:
@misc{chizhov2025hellaswagvaliditycommonsensereasoning,
title={What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks},
author={Pavel Chizhov and Mattia Nee and… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_goldenswagx.llp-gold-37m-1.5m_N1.50M_T8.0_T8.0_T8.0_T8.0landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.goldset
Goldset
Verified bug-fix records for evaluating coding agents.
Every record is a real bug in public software, the fix its author wrote, and the
test that fails before the fix and passes after it. A record is kept only once
both runs have been observed, so what is published is a reproduction rather than
a claim.
896 records from 352 projects, all Python, with
fixes committed between 2010-06-13 and 2026-08-17.
Website · Code and verifier · Datasheet
Quick start
from… See the full description on the dataset page: https://huggingface.co/datasets/goldsetdev/goldset.OpenResearcher-Corpus-Gold-Doc
OpenResearcher Gold Documents
This dataset contains the gold documents used for the "Gold Document Retrieval via Online Bootstrapping" step described in Section 3.2 of the OpenResearcher paper. Gold documents are documents that collectively contain sufficient evidence to derive the ground-truth answer for a given question.
For 6,102 questions sourced from MiroVerse, we constructed a search query by concatenating the question and reference… See the full description on the dataset page: https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Corpus-Gold-Doc.finbenchv2-goldenswag-fi-htThis is a machine-translated and manually corrected subset of GoldenSwag used in Finbench version 2.
Citation information
To cite the original GoldenSwag work:
@misc{chizhov2025hellaswagvaliditycommonsensereasoning,
title={What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks},
author={Pavel Chizhov and Mattia Nee and Pierre-Carl Langlais and Ivan P. Yamshchikov},
year={2025},
eprint={2504.07825},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-goldenswag-fi-ht.sentiment-analysis-in-commodity-market-gold
Dataset Card for Sentiment Analysis of Commodity News (Gold)
This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021).
The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.semantaai-fx-majors-gold5m
semantaai-fx-majors-gold5m
Semanta AI majors gold 5m layer extracted from fx-majors.
nq_open_gold
Natural Questions Open Dataset with Gold Documents
This dataset is a curated version of the Natural Questions open dataset,
with the inclusion of the gold documents from the original Natural Questions (NQ) dataset.
The main difference with the NQ-open dataset is that some entries were excluded, as their respective gold documents exceeded 512 tokens in length.
This is due to the pre-processing of the gold documents, as detailed in this related dataset.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/nq_open_gold.llp-gold-37m-1.5m_N1.50M_T8.0anonymous-ride-gold-lite
RIDE Gold Lite
RIDE Gold Lite is the smaller benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release mirrors the structure and prediction task of RIDE Gold Standard, but uses fewer snapshots and rows for faster inspection, development, and lower-cost experimentation.
Links
Code repository:… See the full description on the dataset page: https://huggingface.co/datasets/ano6060/anonymous-ride-gold-lite.antam_historical_gold_prices
Unofficial Antam gold price history (IDR)
Antam gold selling prices in Indonesian rupiah per gram, compiled from the public price chart on
the official Antam Logam Mulia site. 5,156 records covering 2010-01-04 through 2026-08-14.
This is an unofficial compilation for research, analysis and teaching. See the disclaimer at the
end before you rely on it for anything else.
Files
Antam_historical_gold_prices.csv is the one to use:
Column
Type
Meaning
Time… See the full description on the dataset page: https://huggingface.co/datasets/theonegareth/antam_historical_gold_prices.mercado-ti-gold
Mercado de trabalho em tecnologia — camada gold
Agregados prontos para consumo, derivados dos microdados públicos do CAGED
e da RAIS (Ministério do Trabalho e Emprego), recortados em tecnologia por
setor (CNAE) ou ocupação (CBO).
É a camada que alimenta o dashboard. São os mesmos números da camada silver, já
agregados: 3 MB no lugar de 110 MB por arquivo, com a mesma resposta.
O que tem aqui
grupo
tabelas
Estoque (RAIS)
rais_estoque_anual… See the full description on the dataset page: https://huggingface.co/datasets/Gianpedro/mercado-ti-gold.llp-gold-37m-1.5m_clip0.004_T2048.0_I2048semantaai-fx-other-gold5m-legacy-gamma
semantaai-fx-other-gold5m
Semanta AI FX other gold 5m layer extracted from fx-other.
GoldenCheetah
GoldenCheetah
Data from the GoldenCheetah OpenData Project in Parquet format.
Table
Contents
measurements
Source CSV filename, time (s), distance (km), power (W), heart rate (bpm), cadence (rpm), altitude (m)
activities
Original activity fields, with METRICS and XDATA stored as JSON
athletes
Athlete ID, gender, birth year, and source version
athlete_id identifies the athlete in each table. source_file identifies each CSV trace. Activity dates and CSV… See the full description on the dataset page: https://huggingface.co/datasets/ajto/GoldenCheetah.llp-gold-37m-1.5mllama-3b-gold-15M-student-generations_PRESAMPLING_2048_tune422v1_N1.50M
