datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Iron.task386_semeval_2018_task3_irony_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task386_semeval_2018_task3_irony_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task386_semeval_2018_task3_irony_detection.Earth-Iron
Dataset Card for Earth-Iron
Dataset Details
Dataset Description
Earth-Iron is a comprehensive question answering (QA) benchmark designed to evaluate the fundamental scientific exploration abilities of large language models (LLMs) within the Earth sciences. It features a substantial number of questions covering a wide range of topics and tasks crucial for basic understanding in this domain. This dataset aims to assess the foundational knowledge that underpins… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Iron.task387_semeval_2018_task3_irony_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task387_semeval_2018_task3_irony_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task387_semeval_2018_task3_irony_classification.Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or… See the full description on the dataset page: https://huggingface.co/datasets/JasonChen91/Earth-Iron.IRONWORKS-VENOM-preview
IRONWORKS VENOM
Supply Chain Security Training Dataset — Preview v0.1
by IronGate Digital
What this is
A synthetic instruction-tuning dataset focused on software supply chain security.
Built from real threat intelligence sources including security advisories, research
blogs, and vulnerability databases.
This is an early preview. More datasets are in progress.
Coverage
35,000+ labeled training pairs covering:
Dependency confusion and typosquatting… See the full description on the dataset page: https://huggingface.co/datasets/IronGateDigi/IRONWORKS-VENOM-preview.smolified-smart-food-safety-allergy-detector
🤏 smolified-smart-food-safety-allergy-detector
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ironendgame2019/smolified-smart-food-safety-allergy-detector.
📦 Asset Details
Origin: Smolify Foundry (Job ID: f82e9c7f)
Records: 2890
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ironendgame2019.
Generated… See the full description on the dataset page: https://huggingface.co/datasets/ironendgame2019/smolified-smart-food-safety-allergy-detector.smolified-recipe
🤏 smolified-recipe
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ironendgame2019/smolified-recipe.
📦 Asset Details
Origin: Smolify Foundry (Job ID: c3d8e4cb)
Records: 1150
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ironendgame2019.
Generated via Smolify.ai.
