datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eMCR
eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce
This is the official dataset and evaluation code for the paper "eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce".
Overview
Product search increasingly involves queries that combine multiple requirements — product attributes, brands, prices, exclusions, and visual descriptions. Existing retrieval benchmarks provide limited support for diagnosing which… See the full description on the dataset page: https://huggingface.co/datasets/Lux1997/eMCR.Wan2GPLuxAlign
Dataset Card for LuxAlign
Loading the Dataset
The dataset is currently at version v3, which can be loaded as:
from datasets import load_dataset
ds = load_dataset("fredxlpy/LuxAlign", name="lb-en") # or "lb-fr"
If you want to reproduce the results from the paper (v1) or use any previous version, you can specify the version folder:
# Load version v1 (as used in the paper)
ds_v1 = load_dataset("fredxlpy/LuxAlign", data_dir="data/v1", data_files={"train": "lb_en.json"})… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxAlign.Tool-REX-Toolsfootball2vec-player-embeddings
football2vec Player Embeddings
Pre-computed player embedding vectors from the football2vec v2 transformer model — ready to use without loading model weights. Covers ~87,000 per-match vectors, ~8,950 career vectors, and season-level aggregates across professional soccer competitions. V2 uses a 192-dim transformer encoder with adversarial team debiasing (Ganin GRL) to prevent team identity from confounding player style representations.
Part of the (Right! Luxury!) Lakehouse soccer… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/football2vec-player-embeddings.ipfs_luxembourg_laws_ir
Luxembourg legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_luxembourg_laws (revision ffc59f665caabfb682ee8380e0ea3550266e1b5d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Luxembourg prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_luxembourg_laws_ir.Tool-REX-Queriesqwen3_4b_restricted_Final-activationskfo-luxury-hospitality-corpus
Americas Great Resorts: Canonical Reference Repository
Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 15, 2026
Hugging Face Dataset: Version 1.29
Dataset card version: 1.29Built: September 15, 2026Source commit: 0dbd8213747223aa41c73a3214109053135ce0d8Records: 138Data file: agr-corpus.jsonlSHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.Tool-REX_train_retriever_50ktrajectoriesxg-freeze-frame-data
xG Freeze-Frame Data — StatsBomb 360
~15.58M freeze-frame rows from 323 StatsBomb 360 matches, capturing player positions at the moment of each shot. Each row represents one visible player in one shot event.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset
ds = load_dataset("luxury-lakehouse/xg-freeze-frame-data")
df = ds["train"].to_pandas()
# Average number of visible players per shot… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-freeze-frame-data.statsbomb-shots-on-target
StatsBomb On-Target Shots — Goalmouth Coordinates
~15K on-target shots from StatsBomb Open Data with goalmouth coordinates (end_location_y, end_location_z). Primary training input for the PSxG model used in goalkeeper shot-stopping evaluation.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
⚠️ Schema change (cut-over 2026-07-22)
This dataset emits both legacy and canonical Kimball key columns side-by-side. The legacy match_id column will be… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/statsbomb-shots-on-target.Amazon_Luxury_Beauty_2018
Amazon Luxury Beauty Dataset
Directory Structure
metadata: Contains product information.
reviews: Contains user reviews about the products.
filtered:
e5-base-v2_embeddings.jsonl: Contains "asin" and "embeddings" created with e5-base-v2.
metadata.jsonl: Contains "asin" and "text", where text is created from the title, description, brand, main category, and category.
reviews.jsonl: Contains "reviewerID", "reviewTime", and "asin". Reviews are filtered to include only… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Luxury_Beauty_2018.spadl-vaep-action-values
SPADL/VAEP Action Values
Every on-ball action from ~9.5 million professional soccer events, converted to the SPADL unified format and scored with offensive, defensive, and net VAEP values. Built with the silly-kicks library — enabling player ranking by total contribution beyond goals and assists.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
⚠️ Schema change (cut-over 2026-07-22)
This dataset now emits both legacy and canonical Kimball key… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/spadl-vaep-action-values.luxembourgish-speech-datasetscoutgpt-training-data
ScoutGPT Training Data — Player Action Sequences
Per-player match-level action sequences unified across StatsBomb Open Data and Wyscout open data. Each row is one player-match's ordered SPADL action sequence with contextual tokens (competition, season, score state, half, opponent), serialized as the token stream that the ScoutGPT transformer consumes during training and at inference.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/scoutgpt-training-data.ipfs_luxembourg_laws
Luxembourg Legilux / Mémorial (data.legilux)
Research snapshot of official national legislation. Coverage: incomplete (official PDF fill; 11 empty ELIs).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-03
Coverage
incomplete (official PDF fill; 11 empty ELIs)
Source
data.legilux RDF + JO/Mémorial PDFs (pdftotext). SPARQL UI is an Angular shell, unused.
Collector… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_luxembourg_laws.luxanna_crownguardLuxInstruct
LuxInstruct
Dataset Summary
LuxInstruct is the first large-scale cross-lingual instruction tuning dataset for Luxembourgish, introduced in LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish (Philippy et al., 2025).
It addresses the lack of high-quality instruction–response data for low-resource languages by avoiding direct machine translation into Luxembourgish. Instead, it leverages aligned data from English, French, and German to generate natural… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxInstruct.expected-threat-grids
Expected Threat (xT) Grids
Markov chain expected threat grids computed via value iteration — 192 cells per competition on a 12×16 grid. Each cell quantifies the probability that a possession starting in that zone will end in a goal, derived from observed transition and shot frequencies across ~4,900 matches.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset
import pandas as pd
ds =… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/expected-threat-grids.xg-shot-data
xG Shot Data — StatsBomb + Wyscout
~131K professional soccer shots from StatsBomb Open Data (95K) and Wyscout (43K), with geometric features, categorical context, and goal labels. Partitioned by data_source for selective loading.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
⚠️ Schema change (cut-over 2026-07-22)
This dataset emits both legacy and canonical Kimball key columns side-by-side. The legacy match_id column will be removed on… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-shot-data.scraped-episodesxg-shot-data-v3
Pre-Shot xG v3 — Shot Data (Tabular Corpus)
The tabular half of the training corpus for xg_model_v3, the canonical-SPADL-native pre-shot expected goals model from the luxury-lakehouse analytics platform. One row per shot, all providers — the provider is the data_source column, not a separate file format. Sourced from the gold fct_action_values fact.
Each shot is joinable to its freeze-frame player set (dataset xg-shot-freeze-frames) and to the full action-level corpus… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-shot-data-v3.obso-trained-grids
OBSO Trained Grids — Reachability, EPV, and Completion Matrices
Pre-trained grid artifacts for Off-Ball Scoring Opportunity (OBSO) computation: ball reachability surfaces, expected possession value (EPV) grids, and pass completion probability matrices. These are the static lookup tables that power real-time OBSO evaluation — derived from observed passing, shooting, and transition patterns across ~4,900 open-data matches.
Part of the (Right! Luxury!) Lakehouse soccer analytics… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/obso-trained-grids.StreetMathDatasetTool-REX_train_reranker_200kspace-creation-values
Space Creation Values — ELASTIC/OBSO
Per-player per-frame space creation quantification — measuring each player's contribution to off-ball scoring opportunities via differential OBSO. For every sampled frame, the model computes OBSO with and without each player, yielding the area of scoring opportunity that player creates (or destroys) by their positioning.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/space-creation-values.luxembourgish-parliamentary-corpus
Luxembourgish Parliamentary Corpus (2023–2028)
A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary
proceedings, built from the official session reports (comptes rendus /
"D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028.
Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the
Luxembourgish Wikipedia holds roughly 64,000 articles and most large language
models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.houseplant-light-needs-and-indoor-lux-readings
Houseplant Light Needs and Indoor Lux Readings
Two files, both CC BY 4.0. Archived with a DOI on Zenodo: 10.5281/zenodo.22023337.
Note for loaders: both CSVs open with commented header lines (#) carrying the snapshot date, the licence and the band definitions. Skip them when reading.
Files
species-light-levels.csv — one row per houseplant species in the GrowSpot care library, with the light tier it belongs to. Columns: common_name, scientific_name… See the full description on the dataset page: https://huggingface.co/datasets/Growspot/houseplant-light-needs-and-indoor-lux-readings.
