datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
k12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.standard-malay-translation-instructionsiso-standard-supersessions
Withdrawn and superseded ISO standards
Canonical, always-current version: https://referencesource.org/iso-standard-supersessions/
Machine-readable: https://referencesource.org/iso-standard-supersessions/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2027-02-02 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 24542
Every deliverable in ISO's own open-data register that has been… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/iso-standard-supersessions.movielens-1m-ratings-standardizedxml-standards-specificationssuperseded-standards-mappings
Superseded and deprecated identifier mappings
Canonical, always-current version: https://referencesource.org/superseded-standards-mappings/
Machine-readable: https://referencesource.org/superseded-standards-mappings/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-04
Stale after: 2027-01-31 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 141
Withdrawn standards, deprecated API models and retired… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/superseded-standards-mappings.VG150-SGG-StandardBCE-Prettybird-Micro-Standard-v0.0.2
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.2.nrtl-recognized-test-standards
Which OSHA-recognized labs (NRTLs) can certify to a given test standard
Canonical, always-current version: https://referencesource.org/nrtl-recognized-test-standards/
Machine-readable: https://referencesource.org/nrtl-recognized-test-standards/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-10
Stale after: 2026-11-08 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 4329
The inverted view of OSHA's… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nrtl-recognized-test-standards.fastener-thread-standard-equivalence
Fastener standard equivalences: DIN to ISO / EN crossover
Canonical, always-current version: https://referencesource.org/fastener-thread-standard-equivalence/
Machine-readable: https://referencesource.org/fastener-thread-standard-equivalence/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2028-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 164
Which ISO or EN standard… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/fastener-thread-standard-equivalence.BCE-Prettybird-Micro-Standard-v0.0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.1.commercial-production-standard-fixtures
SHAR Production Commercial Production Standard Fixtures
Synthetic valid and invalid examples for deterministic testing of the Commercial Production Standard.
Publisher: SHAR Production — https://sharprod.com/
data/valid-handoff.json — accepted synthetic delivery handoff.
data/invalid-handoff.json — intentional validation failures.
schema/commercial-production-handoff.schema.json — MIT-licensed schema reference.
Examples are explicitly synthetic and licensed CC BY 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/SHARProduction/commercial-production-standard-fixtures.osha-training-requirements-by-standard
OSHA training requirements by standard: topic, frequency, and CFR citation
Canonical, always-current version: https://referencesource.org/osha-training-requirements-by-standard/
Machine-readable: https://referencesource.org/osha-training-requirements-by-standard/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2028-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 31
Which OSHA… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/osha-training-requirements-by-standard.fastener-standard-interchangeability-exceptions
DIN / ISO / BS fastener equivalents and the sizes where they are not interchangeable
Canonical, always-current version: https://referencesource.org/fastener-standard-interchangeability-exceptions/
Machine-readable: https://referencesource.org/fastener-standard-interchangeability-exceptions/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2029-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/fastener-standard-interchangeability-exceptions.food-standards-of-identity
What a food must legally contain to use its name: FDA standards of identity, in one table
Canonical, always-current version: https://referencesource.org/food-standards-of-identity/
Machine-readable: https://referencesource.org/food-standards-of-identity/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2027-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 41
The FDA standards… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/food-standards-of-identity.standardebooks-enPurified-openai-messages
📖 StandardEbooks-enPurified-openai-messages
StandardEbooks-enPurified is a highly curated, "prose-first" iteration of the Nelathan/standardebooks dataset.
While the source dataset provides excellent public domain literature, raw full-text novels are difficult to ingest directly into training pipelines. This dataset solves that by applying intelligent context-aware chunking and formatting the data into the OpenAI Messages standard.
The goal is to provide a clean, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/standardebooks-enPurified-openai-messages.certification-standards-equivalence
Certification marks and hazardous-area classifications: cross-market equivalences
Canonical, always-current version: https://referencesource.org/certification-standards-equivalence/
Machine-readable: https://referencesource.org/certification-standards-equivalence/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2027-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 200
Which… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/certification-standards-equivalence.texas-k12-curriculum-standards-teks
Texas K-12 Curriculum Standards (TEKS-derived)
15,040 generated learning-objective records organized around the Texas Essential
Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical
Education clusters, and specialized program areas.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.plumbing-fitting-thread-standards
Plumbing fitting and pipe thread standards
Canonical, always-current version: https://referencesource.org/plumbing-fitting-thread-standards/
Machine-readable: https://referencesource.org/plumbing-fitting-thread-standards/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2028-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 97
Dimensional specifications for NPT and BSP pipe… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/plumbing-fitting-thread-standards.state-heat-illness-prevention-standards
State Heat Illness Prevention Standards — Temperature Triggers and Required Employer Actions
Canonical, always-current version: https://referencesource.org/state-heat-illness-prevention-standards/
Machine-readable: https://referencesource.org/state-heat-illness-prevention-standards/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2027-08-17 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-heat-illness-prevention-standards.osha-fall-protection-trigger-height-by-standard
OSHA fall protection trigger height by standard
Canonical, always-current version: https://referencesource.org/osha-fall-protection-trigger-height-by-standard/
Machine-readable: https://referencesource.org/osha-fall-protection-trigger-height-by-standard/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-25
Stale after: 2027-08-25 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 42
OSHA does not set… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/osha-fall-protection-trigger-height-by-standard.datatager_standard_med_question
If you like our project, please give us a star ⭐
[GitHub | DataTager Home]
Standard Medical Question
Prompt for Training
When training your model with this dataset, prepend the following prompt to each input instance:
你需要将医疗领域中的冗长或复杂的患者咨询文本转换为简洁、结构化的问题表达。请确保输出文本保留所有关键的医疗信息,去除重复或不必要的细节,并使用专业的医疗术语准确描述患者的情况和需求。
Description
AnyTaskTune is a publication by the DataTager team. We advocate for rapid training of large models suitable for specific business… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/datatager_standard_med_question.moore-audio-standardized ---
pretty_name: louisbertson/moore-audio-standardized
language:
- mos
tags:
- audio
- moore
- self-supervised-learning
- speech
size_categories:
- n<1K
---
# Mooré Standardized Audio Dataset
This dataset was exported from the preprocessing pipeline in this repository. It keeps the repository's canonical split manifests and uses standardized WAV audio so the same files work in local training, Google Colab, and Hugging Face Hub uploads.… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/moore-audio-standardized.BCE-Prettybird-Large-Standard-v0.0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Large-Standard-v0.0.1.BCE-Prettybird-Micro-Standard-v0.0.5
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.5.BCE-Prettybird-Micro-Standard-v0.0.4
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.4.BCE-Prettybird-Micro-Standard-v0.0.3
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.3.BCE-Prettybird-Micro-Standard-v0.0.6
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.6.us-address-standardization
US Address Standardization (PostGIS stdaddr schema)
Chat-format instruction data that maps a raw US address string to a strict JSON
object matching the PostGIS address_standardizer stdaddr type (camelCase keys,
USPS-abbreviated values). Used to fine-tune the qwen35-address-std model family.
Each row has three flat fields (system / user / assistant) for easy reading
and grepping; rebuild the chat messages list from them at train time:
{
"system": "<standardization… See the full description on the dataset page: https://huggingface.co/datasets/davidr99/us-address-standardization.standardebooks-jsonl
Standard Ebooks JSONL
This dataset is a JSONL conversion of the Hugging Face dataset
Nelathan/standardebooks.
The source dataset contains full-text public domain books sourced from
Standard Ebooks.
Dataset Structure
The dataset has one split, train, stored as two JSONL shards:
data/train-00000-of-00002.jsonl - 617 rows
data/train-00001-of-00002.jsonl - 616 rows
Each line is a JSON object with the same fields as the source parquet dataset:
link: URL of the Standard… See the full description on the dataset page: https://huggingface.co/datasets/virtualkevin/standardebooks-jsonl.
