datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physics-corpus
Physics Corpus — konsman/physics-corpus
arXiv physics papers exported from a PostgreSQL mirror of the Kaggle arXiv dataset,
structured for ontology extraction and downstream NLP pipelines.
Configuration: quantum-physics
arXiv categories included: quant-ph, hep-th, gr-qc
Schema
Field
Type
Description
paper_id
string
arxiv:<id>v<n> — stable across pipeline runs
arxiv_id
string
Base arXiv ID without version suffix
arxiv_version
int32
Version number… See the full description on the dataset page: https://huggingface.co/datasets/konsman/physics-corpus.PhysicsEval
PhysicsEval Dataset
To enable large-scale evaluation and training of reasoning-capable language models in physics, we curated a comprehensive dataset of 19,609 annotated problems, sourced from 20 authoritative physics textbooks and verified educational websites.
Construction
The dataset spans 19 different categories, including Mechanics, Thermodynamics, Electromagnetism, Waves, Optics, Relativity, and Quantum Physics.
Each problem is processed through the following… See the full description on the dataset page: https://huggingface.co/datasets/IUTVanguard/PhysicsEval.physics-russian
Оглавление
Описание датасета
Аннотация
Ключевые особенности
Статус перевода
Методология перевода и верификации
Ограничения и возможные погрешности
Структура датасета
Поля данных
Использование
Благодарности
Лицензирование и авторские права
Цитирование
📑 Оглавление Примеров
Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте.
Квантовая механика
Термодинамика
Электромагнетизм
Общая теория относительности
Специальная теория относительности
Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.quantum-physics-0.6-corpus
quantum-physics-0.6-corpus
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain
similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6-corpus.physics-reasoning-dataset
📚 Flux Physics Reasoning Dataset
This dataset contains detailed physics reasoning scenarios designed to train Small Language Models (SLMs) and Liquid Neural Networks in physical intuition.
📄 Format
The dataset is provided in Parquet format (train.parquet) for efficient loading. Each row contains:
prompt: The physics question or scenario description.
answer: The correct physical explanation and answer.
concept: The underlying physics principle (e.g., "Conservation of… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/physics-reasoning-dataset.task708_mmmlu_answer_generation_high_school_physics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task708_mmmlu_answer_generation_high_school_physics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task708_mmmlu_answer_generation_high_school_physics.camel-ai-physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4.
The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.… See the full description on the dataset page: https://huggingface.co/datasets/lgaalves/camel-ai-physics.quantum-physics-0.6
quantum-physics-0.6
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6.task691_mmmlu_answer_generation_college_physics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task691_mmmlu_answer_generation_college_physics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task691_mmmlu_answer_generation_college_physics.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.Physics-DPO-Dataset
Dataset Summary
Physics-DPO is a meticulously crafted, synthetic dataset of preference pairs designed to enhance the physics reasoning and problem-solving capabilities of Large Language Models (LLMs). Generated using a sophisticated, multi-stage pipeline, this dataset provides high-quality data for Direct Preference Optimization (DPO) and similar preference alignment techniques.
Each instance consists of a challenging physics problem (prompt), a high-quality, expert-level solution… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Physics-DPO-Dataset.Physics-MATH-Composition-141K
Composition-RL
This repository contains the datasets for the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
GitHub | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that combats the growing number of “too-easy” prompts (pass-rate = 1) by automatically composing multiple verifiable problems into a single, harder yet still-verifiable prompt. Across 4B–30B models… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Physics-MATH-Composition-141K.NCERT_Physics_11thphysics-ptbr
Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.quantum-physics-0.5-corpus
quantum-physics-0.5-corpus
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain
similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.5-corpus.quantum-hardware-device-physics
Neura Parse — Quantum Hardware Device Physics: Qubit Design, Coherence, Control & Scaling
A physics- and engineering-deep vertical on how qubits are built, controlled, and scaled across superconducting, trapped-ion, neutral-atom, and spin modalities (plus emerging erasure/biased-noise qubits). Device-physics derivations, coherence-limit analyses, control-stack engineering, and 2025-2026 scaling/interconnect work, with QuTiP/scqubits simulation context — expanding the general… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-hardware-device-physics.physics-qa
Physics Q&A — Multi-Level Explanations
502 question-answer pairs generated from recent physics papers (arXiv 2024–2026),
covering 5 subfields across 83 papers.
Each paper is explained at 6 audience levels, each as a focused Q/A pair:
Level
Audience
physicist
Expert with equations and notation
undergrad_science
Undergraduate with key equations
high_schooler
High school student, intuitive
humanities_student
No math, analogies only
five_year_old
Child-friendly… See the full description on the dataset page: https://huggingface.co/datasets/planetoid-reader/physics-qa.NCERT_Physics_12thdeep-physics-science-zh
Deep Physics & Science Dialogue Dataset (Chinese)
深度物理科学对话数据集
Dataset Description
High-quality Chinese physics and science dialogues covering quantum gravity, theory of everything, relativity, quantum mechanics, and entropy/information theory.
高质量中文物理科学对话,涵盖量子引力理论、万物理论、相对论、量子力学、熵与信息论等硬核科学议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-physics-science-zh.NC_Physics
NC_Physics
NC_Physics is the dataset released with LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation.
The dataset supports Content-Conditional Introduction Generation (CCIG): models use the non-introduction content of scientific papers, paper metadata, and references to reason about the paper's core idea and generate a logic-aware introduction.
Paper: https://arxiv.org/abs/2605.25964
Code: https://github.com/Xiao-Youth/LECTOR
Associated… See the full description on the dataset page: https://huggingface.co/datasets/Xiao-Youth/NC_Physics.task1422_mathqa_physics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1422_mathqa_physics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1422_mathqa_physics.PhysicsConcepts-Instruct-v1
PhysicsConcepts-Instruct-v1
PhysicsConcepts-Instruct-v1 is a synthetic physics instruction dataset designed for supervised fine-tuning of language models on fundamental and advanced physics concepts. It covers diverse domains including classical mechanics, electromagnetism, thermodynamics, optics and waves, and modern physics through clear explanations, physical intuition, derivations, worked examples, and educational discussions. The dataset is suitable for training educational… See the full description on the dataset page: https://huggingface.co/datasets/kd13/PhysicsConcepts-Instruct-v1.physics_arabic_CoT_10K
Dataset Summary
The Arabic Physics Problem Dataset is a curated collection of Arabic-language physics exam questions.Each record contains a question written in Modern Standard Arabic, the correct answer, a detailed explanation, and the related physics concept.This dataset is part of the Mobiusi multilingual STEM education initiative, created to support scientific reasoning, question answering, and educational AI applications in Arabic-speaking contexts.
Each sample follows a… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/physics_arabic_CoT_10K.smolified-bengali-physics-teacher
🤏 smolified-bengali-physics-teacher
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model RohanSardar/smolified-bengali-physics-teacher.
📦 Asset Details
Origin: Smolify Foundry (Job ID: f26e2704)
Records: 9493
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by RohanSardar.
Generated via Smolify.ai.
brazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/physics with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.NCERT_Physics_12thphysics-30k-demo
Computational & Quantitative Sciences Q&A — Multi-Level Explanations
24 question-answer pairs generated from recent papers (arXiv 2024–2026),
covering 6 subfields across 6 papers.
Each paper is explained at 4 depth levels, each as a focused Q/A pair:
Level
Description
L1
Intuitive / Phenomenological — what is happening, plain language, analogies, no equations
L2
Conceptual / Structural — key components, pipeline/steps, minimal formalism
L3
Mechanistic / Formal —… See the full description on the dataset page: https://huggingface.co/datasets/planetoid-reader/physics-30k-demo.physics-recontext-sft
Physics Concept-Preserving Re-Contextualization — SFT data (v1)
Supervised fine-tuning data for a small model that, given a high-school physics
flashcard (question + answer), writes one new question testing the same underlying
principle in a clearly different real-world scenario — a far transfer, not a
paraphrase. The model is never told the principle name; it must infer and preserve it.
Example. Original: "What causes atmospheric pressure?" →
Variant: "What causes the water… See the full description on the dataset page: https://huggingface.co/datasets/meric533/physics-recontext-sft.physics-FR-train-dataset
Dataset Card for Wiki-Physique-FR
Ce jeu de données est une version spécialisée ciblant exclusivement le domaine de la physique au sein de Wikipédia en français. Il est idéal pour le fine-tuning de modèles de langage sur des connaissances scientifiques.
Construction du Dataset
La méthode de construction suit la même pipeline que le dataset général, mais avec un filtrage thématique strict dès l'étape SPARQL.
1. Sélection Thématique
La sélection s'est basée sur… See the full description on the dataset page: https://huggingface.co/datasets/Houzeric/physics-FR-train-dataset.
