datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Stylistics_Reference_Repository
MoSEs Dataset: Stylistics Reference Repository(SRR)
This dataset is part of the MoSEs framework for AI-generated text detection, containing both human-written and AI-generated text data used in the paper "MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds" (Wu et al., 2025).
Dataset Overview
This dataset contains two text detection benchmark subsets used for training and evaluation in the MoSEs framework.… See the full description on the dataset page: https://huggingface.co/datasets/zhengliu8/Stylistics_Reference_Repository.lingvanex_test_references
LTR
LTR -- Lingvanex Test References for MT Evaluation from English into a total of 30 target languages for a big variety of cases.
TEST CASES
Parameter
Description
Length
Sentences from 1 to 100 words.
Domain
Medicine (12%), Automobile (11%), Finance (8%)
Tokenizer
Jupiter is 1.000.000 km far. Ask Mr. Johnson for training
Tags
I want to eat and swim
Capitalisation (Case)
HELLO my Dear frIEND
Different languages in one text (Up to 3 languages)
I see… See the full description on the dataset page: https://huggingface.co/datasets/lingvanex/lingvanex_test_references.pokemon-card-sold-price-reference
Pokémon Card Sold-Price Reference by Grade — Raw, PSA 9, PSA 10 (486 Cards, 2026)
Median sold-price reference for 486 Pokémon cards across raw (ungraded), PSA 9, and PSA 10 eBay sales, compiled from real graded and ungraded sold comps (PokemonPriceTracker sold-listing data). 486 of 486 cards have a PSA 9 comp; the median PSA-10-over-PSA-9 grading premium across cards with both is 4.9×. Reference pricing only — sample sizes and confidence flags are included so nobody treats a… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/pokemon-card-sold-price-reference.brenda-references-databrand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.reference-csvface-shape-measurement-reference
Face Shape Measurement Reference (8 Shapes)
A compact, reusable reference for identifying face shape from four tape measurements:
forehead width, cheekbone width, jawline width, and face length.
Files
File
What it is
face_shape_measurement_sheet.csv
One row per shape — oval, round, square, heart, diamond, oblong, triangle/pear, hourglass. The four columns state where each measurement sits relative to the other three; fastest_check gives the single… See the full description on the dataset page: https://huggingface.co/datasets/EthanCui/face-shape-measurement-reference.hebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.legal-reference-annotationsIn this dataset, we present a dataset of 2944 legal references in German law that are manually annotated by law experts. This dataset has 21 properties for each law reference in the dataset, such as Buch, Teil, Titel, Untertitel, etc. It also provides the complete text of each law reference in the dataset, along with specific paragraph text mentioned in the law reference.
Paper: A Dataset of German Legal Reference Annotations
Please reference our work when using this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/legal-reference-annotations.developer-reference-datasets
Developer Reference Datasets
Open, reproducible lookup tables that web and app developers reach for constantly — computed from first principles, not scraped, so every value is exact and re-runnable. CC BY 4.0.
Quick answers (straight from the data)
What is 16:9 in pixels? 1920×1080, 1280×720, 3840×2160. 9:16 (Stories, Reels, TikTok) is those flipped. → aspect-ratios, resolutions
What contrast ratio does WCAG require? 4.5:1 for normal text (AA), 3:1 for large… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/developer-reference-datasets.risk-matrix-reference-2026
Canonical landing page: https://www.smartqhse.com/datasets/risk-matrix-reference-2026
Risk Matrix Reference 2026 — 5x5, 4x4, and 3x3 with ALARP Zones
Fully populated risk matrices in 5x5 (25 cells), 4x4 (16 cells), and 3x3 (9 cells) configurations. Each cell: severity level + descriptor, likelihood level + descriptor, risk score, risk band (Very Low / Low / Medium / High / Very High), ALARP zone (Broadly Acceptable / Tolerable / Intolerable), and recommended management… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/risk-matrix-reference-2026.future-time-referencesbible-reference-sentence-pairfuture-time-references-static-filter-D1round-reference-ad7a26
round-reference-ad7a26
Synthetic products test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/pbae/round-reference-ad7a26.ppe-standards-cross-reference-2026
Canonical landing page: https://www.smartqhse.com/datasets/ppe-standards-cross-reference-2026
PPE Standards Cross-Reference 2026 — ANSI vs EN vs AS/NZS
Cross-reference of PPE standards across ANSI/ISEA (US), EN/EN ISO (EU/UK), and AS/NZS (Australia/NZ) frameworks. ~80-100 rows covering eye, head, foot, hand (cut/chemical/heat/electrical), hi-vis, hearing, respiratory, fall protection, and arc flash. Maps equivalent standard numbers and key test parameters/ratings across… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/ppe-standards-cross-reference-2026.Reference-Letter-Bias-Prompts
Reference Letter Bias Dataset
Reference Letter Bias dataset was created by (Wan et al., 2023) and published under the MIT license
(https://github.com/uclanlp/biases-llm-reference-letters). The purpose of the dataset is to investigate gender bias in
LLMs, specifically regarding the generation of letters of recommendation.
(Wan et al., 2023) explores how gender biases manifest in the LLM generation of reference letters by analyzing the language style
and lexical content of… See the full description on the dataset page: https://huggingface.co/datasets/elaine1wan/Reference-Letter-Bias-Prompts.engine-optimization-reference
ThatDevPro Engine Optimization Reference Corpus
A curated snapshot of canonical reference documentation for engine optimization (the merge of SEO + AEO + AIO + GEO), maintained by ThatDevPro.
Overview
This dataset indexes 26 canonical reference pages from the ThatDevPro Reference Library, each covering one specific crawler signal a website emits. The dataset is intended for training and evaluating AEO/GEO scoring systems, llms.txt validators, and structured-data… See the full description on the dataset page: https://huggingface.co/datasets/ThatDeveloperGuy13/engine-optimization-reference.hazop-guidewords-reference-2026
Canonical landing page: https://www.smartqhse.com/datasets/hazop-guidewords-reference-2026
HAZOP Guidewords Reference 2026 — IEC 61882 with Node Examples
HAZOP guidewords reference following IEC 61882:2016 and CCPS guidance. 7 standard guidewords (NO/NOT, MORE, LESS, AS WELL AS, PART OF, REVERSE, OTHER THAN) applied to key process parameters (flow, pressure, temperature, level, composition, reaction, rotation, operating mode) with ~20 worked examples each providing: deviation… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hazop-guidewords-reference-2026.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
iso-45001-2018-clause-reference-2026
ISO 45001:2018 Occupational Health and Safety Management System — Clause Reference 2026
Structured reference of ISO 45001:2018 clauses 4–10, mapping each clause to its intent, key requirements, and PDCA phase for OHS management systems.
Details
Publisher: SmartQHSE Ltd (https://www.smartqhse.com)
License: CC BY 4.0
Format: CSV (UTF-8)
DOI: 10.5281/zenodo.20446739
Landing page: https://www.smartqhse.com/datasets/iso-45001-2018-clause-reference-2026… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/iso-45001-2018-clause-reference-2026.research-peptides-reference
Research Peptides Reference Dataset
A clean, machine-readable reference table of research-grade peptides commonly
discussed in biochemistry and drug-discovery literature. Each entry combines a
curated research category with verified physicochemical properties pulled from
PubChem (PUG REST): PubChem CID, molecular
formula, molecular weight, canonical SMILES and IUPAC name.
The goal is a small, high-signal starting point for cheminformatics, tabular ML,
educational tooling and… See the full description on the dataset page: https://huggingface.co/datasets/PeptidosSuplementos/research-peptides-reference.permit-to-work-types-reference-2026
Permit-to-Work Types Reference 2026
Reference table of 11 common permit-to-work types covering scope, key hazards, typical controls, authorising roles, validity, and confident regulatory references.
Details
Publisher: SmartQHSE Ltd (https://www.smartqhse.com)
License: CC BY 4.0
Format: CSV (UTF-8)
DOI: 10.5281/zenodo.20446749
Landing page: https://www.smartqhse.com/datasets/permit-to-work-types-reference-2026
Citation
SmartQHSE Ltd (2026).… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/permit-to-work-types-reference-2026.bible-reference-sentence-pairAIMO2-referenceThis CSV file is reference.csv in Kaggle's AI Mathematical Olympiad - Progress Prize 2.
Mood-Regulation-Audio-Reference-Catalog
Mood-Regulation-Audio-Reference-Catalog
Entity Reference
Creator: Inna StoryISNI: 0000 0005 3033 4113Framework: Entity Life Cycle (ELC)Official Hub: innastoryofficial.com
Описание
Данный датасет представляет собой техническую библиотеку аудио-ассетов, спроектированных для прецизионного управления психоэмоциональным состоянием слушателя. Проект базируется на методологии Entity Life Cycle (ELC) и рассматривает музыкальный контент как инженерный… See the full description on the dataset page: https://huggingface.co/datasets/InnaStory/Mood-Regulation-Audio-Reference-Catalog.wikidata_reference
Dataset Card for Triple-to-Text Alignment Dataset
Dataset Summary
The Triple-to-Text Alignment dataset aligns Knowledge Graph (KG) triples from Wikidata with diverse, real-world textual sources extracted from the web. Unlike previous datasets that rely primarily on Wikipedia text, this dataset provides a broader range of writing styles, tones, and structures by leveraging Wikidata references from various sources such as news articles, government reports, and scientific… See the full description on the dataset page: https://huggingface.co/datasets/sven-h/wikidata_reference.citations_with_valid_and_invalid_referencesreference-frame-perspective-integrity-meta-v01
Dataset
ClarusC64/reference-frame-perspective-integrity-meta-v01
This dataset tests one capability.
Can a model keep claims inside the correct reference frame.
Core rule
Every claim has a viewpoint.
A model must not slide between frames without saying so.
It must respect
who is speaking
what is being described
what level of certainty the frame allows
A personal view is not objective proof.
A population statistic is not an individual destiny.
A simulation is not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reference-frame-perspective-integrity-meta-v01.reference_metadata_2013
