datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
red-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.FormulaReasoning
FormulaReasoning
This is a Chinese-English bilingual question-answering dataset, which includes the following subsets:
formulareasoning
formulareasoning_enhancement
Each subset has the following split:
train.json: Training data
HoF_test.json: Homogeneous formulas testing data
HeF_test.json: Heterogeneous formulas testing data
Field Descriptions
Field
Type
Description
id
str
Each sample's unique identifier.
question
dict
Sample's question includes the… See the full description on the dataset page: https://huggingface.co/datasets/cat-overflow/FormulaReasoning.formulae__mita-v1.1-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.1-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.1-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.1-7b-2-24-2025-details.formulation-records-instruct
Formulation Records Instruct
A synthetic, de-identified instruction dataset for teaching an LLM to structure
and reason over semiconductor wet-chemistry formulation records. Three tasks:
extract (messy text → structured JSON), normalize (name/unit → canonical),
explain (optimizer result → plain-language explanation).
Built by Formulith. Generator & schema:
https://github.com/formulith/formulation-records-dataset.
本專案部分研發由數位發展部數位產業署 115 年 AI 算力平台支持。
Part of this work is… See the full description on the dataset page: https://huggingface.co/datasets/Formulith/formulation-records-instruct.sat-math-formula-sheet
SAT Math Formula Sheet
The 24 formulas and concepts the SAT does not give you on its reference sheet.
The digital SAT provides a reference sheet on every math question with area,
circumference, volume and right-triangle formulas. It does not provide slope,
the quadratic forms, the discriminant, exponent rules, percent change,
exponential growth, probability, or the circle equation. These 24 cover what it
leaves out.
Contents
formulas.json holds 24 records:… See the full description on the dataset page: https://huggingface.co/datasets/SigmaPrep/sat-math-formula-sheet.formulae__mita-elite-v1.1-7b-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.1-7b-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.1-7b-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.1-7b-2-25-2025-details.tcm-formulary
Formulary · 中医方书 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Classical prescription collections: 局方·千金方·外台秘要·医方集解 (方剂)
91 public-domain works of classical Traditional Chinese Medicine, as clean full text… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-formulary.hemlock-formulary-SFT
Hemlock Apothecary Formulary
Stdlib-focused SFT data for fine-tuning Hemlock-Apothecary-7B — a
Hemlock-Codex-7B derivative
aimed at closing the L2-Stdlib gap measured on hembench.
A formulary in pharmacology is the reference book listing drug compositions and dosages.
Same idea here: one realistic program per (@stdlib/<module>, task) pair showing
exactly how to compose real Hemlock stdlib calls — right imports, right method names,
right idioms.
Why this dataset exists… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-formulary-SFT.physics-formulas-trformula-1-detailed-1999-2026
🏎️ F1 Grand Prix Analytics: The 28-Season Consolidated Dataset (1999–2026)
🏁 Executive Summary
This repository contains a high-fidelity, unified dataset of Formula 1 race results spanning from the 1999 Season through the 2025 Season, supplemented by predictive/projected data for the 2026 Season (limited to the first 6 rounds, ending at the Monaco Grand Prix).
This dataset captures the evolution of technical regulations—from the screaming V10s and V8s to the… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/formula-1-detailed-1999-2026.math-formulas-trai-economics-formulas
AI Economics Formulas
The 12 models behind piszczek.pl/tools — calculators for AI cost, energy and agent verification — as machine-readable records: question, formula, parameters with defaults, worked default result and a plain-language interpretation.
By Michał Piszczek (CTO of Archdesk), author of the Joule Wars, Proof-Adjusted Autonomy and Revocation Exposure concepts.
Fields
field
description
slug
tool identifier
question
the question the model… See the full description on the dataset page: https://huggingface.co/datasets/cdiamond/ai-economics-formulas.NI_chemical_formulaformulae__mita-v1.2-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.2-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.2-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.2-7b-2-24-2025-details.formulae__mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp-details
Dataset Card for Evaluation run of formulae/mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp
Dataset automatically created during the evaluation run of model formulae/mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-sce-gen1.1-v1-7b-2-26-2025-exp-details.formulae__mita-v1-7b-details
Dataset Card for Evaluation run of formulae/mita-v1-7b
Dataset automatically created during the evaluation run of model formulae/mita-v1-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1-7b-details.formulae__mita-math-v2.3-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-math-v2.3-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-math-v2.3-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-math-v2.3-2-25-2025-details.formulae__mita-elite-v1.2-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.2-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.2-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.2-7b-2-26-2025-details.formulae__mita-gen3-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-gen3-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-gen3-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-gen3-7b-2-26-2025-details.formulae__mita-elite-v1.1-gen2-7b-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.1-gen2-7b-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.1-gen2-7b-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.1-gen2-7b-2-25-2025-details.formulae__mita-gen3-v1.2-7b-2-26-2025-details
Dataset Card for Evaluation run of formulae/mita-gen3-v1.2-7b-2-26-2025
Dataset automatically created during the evaluation run of model formulae/mita-gen3-v1.2-7b-2-26-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-gen3-v1.2-7b-2-26-2025-details.formulas[
{
"id": "0",
"translation": {
"es": "8a - 4b + 16c + 12d",
"pt": "Para factorizar la expresión (8a - 4b + 16c + 12d), primero agrupemos los términos de manera adecuada. La expresión se puede reorganizar en dos grupos: (8a - 4b) + (16c + 12d). Ahora, en cada grupo, factorizamos los términos comunes: Grupo 1: Factor común de (4) en (8a - 4b): 4(2a - b) . Grupo 2: Factor común de (4) en (16c + 12d): 4(4c + 3d). Finalmente, podemos escribir la expresión factorizada como la… See the full description on the dataset page: https://huggingface.co/datasets/spongebob01/formulas.prefChat_1k
Dataset Card for PrefChat-1k
prefChat-1k is a curated preference dataset by Formula X (Christopher Chibuikem) designed for aligning conversational models using Direct Preference Optimization (DPO).Each example contains a prompt (user utterance), a chosen response (human-like, empathetic, conversational) and a rejected response (robotic, stiff, or disengaged). The dataset's goal is to teach models to prefer natural, relatable replies over mechanical/robotic-sounding ones.… See the full description on the dataset page: https://huggingface.co/datasets/formula-x/prefChat_1k.intelinvest-formulasChem_Formulas_v1Molecular_formula_10
