datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-material-contracts
Material Contracts (Exhibit 10) from SEC/EDGAR
Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕
Dataset Summary
Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/stonelight/hle_material_science.sofc_materials_articlesThe SOFC-Exp corpus consists of 45 open-access scholarly articles annotated by domain experts.
A corpus and an inter-annotator agreement study demonstrate the complexity of the suggested
named entity recognition and slot filling tasks as well as high annotation quality is presented
in the accompanying paper.catalogue_published_material
Dataset summary
This dataset contains the bibliographic records from the Library’s catalogue of published material: books, maps, music, journals, newspapers, pamphlets, flyers and more, and includes records for printed and digital publications. It excludes records from our catalogue where we believe the originator exerts rights over the re-use of the metadata.
This version contains over 5 million records which are split into 51 files of approximately 100,000 records each for ease… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/catalogue_published_material.quantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.miobra-synthetic-construction-material-extraction-v1
Mi Obra Synthetic Construction Material Extraction v1.0
English
Dataset Description
This dataset contains 10,000 synthetic Spanish construction-material titles
paired with structured entity-extraction targets. It is designed for Argentine
construction terminology and controlled experiments in fine-tuning, teaching,
information extraction, and structured generation.
Language: Spanish (es)
Regional context: Argentina
Rows: 10,000
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mi-obra/miobra-synthetic-construction-material-extraction-v1.Material_Selection_EvalA benchmark designed to facilitate evaluation and modify the behavior of a foundation model through different existing techniques in the context of material selection for conceptual design.
The data is collected by conducting a survey of experts in the field of material selection. The same questions mentioned in keyquestions.csv are asked to experts.
This can be used to evaluate a Language model performance and its spread compared to a human evaluation.
To get into a more detailed explanation… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/Material_Selection_Eval.synthetic-aat-materials
Synthetic AAT Materials Dataset
Dataset Description
This dataset contains 1000 synthetic examples of cultural heritage object descriptions paired with their materials as they would appear in the Getty Art & Architecture Thesaurus (AAT). The data is formatted for training conversational AI models, particularly Qwen3, to identify and extract materials from cultural heritage object descriptions.
Dataset Structure
Each example contains:
messages: Conversation… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-aat-materials.Material_Selection_EvalA benchmark designed to facilitate evaluation and modify the behavior of a foundation model through different existing techniques in the context of material selection for conceptual design.
The data is collected by conducting a survey of experts in the field of material selection. The same questions mentioned in keyquestions.csv are asked to experts.
This can be used to evaluate a Language model performance and its spread compared to a human evaluation.
To get into a more detailed explanation… See the full description on the dataset page: https://huggingface.co/datasets/Frederick001/Material_Selection_Eval.
