datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
English-Handwritten-Math-Notes-Dataset
English Handwritten Math Notes Dataset
This dataset contains high-resolution images of handwritten mathematical notes written in English. It includes problem statements, worked examples, formulas, and annotated derivations. The dataset supports AI research in handwriting recognition, mathematical OCR, and document understanding for STEM applications.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/English-Handwritten-Math-Notes-Dataset.Math-Shapes
Math-Symbols Dataset
Overview
The Math-Symbols dataset is a collection of images representing various mathematical symbols. This dataset is designed for machine learning applications, particularly in the fields of image recognition, optical character recognition (OCR), and symbol classification.
Dataset Details
Name: Math-Symbols
Type: Image dataset
Format: Images with corresponding labels
Size: 131MB (downloaded dataset files), 118MB (auto-connected Parquet… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Shapes.Math-Equa
Math-Equa Dataset
Overview
The Math-Equa dataset is a collection of mathematical equations designed for machine learning applications. This dataset can be used for tasks such as equation solving, symbolic mathematics, and other related research areas.
Dataset Details
Name: Math-Equa
Type: Mathematical Equations
Format: Text-based equations
Size: [Insert size of the dataset]
Source: [Insert source of the dataset, if applicable]
Usage
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Equa.MathStrike
Math-Strike Dataset
Dataset Description
Math-Strike is a novel dataset featuring real and synthetic strike-outs, component-level annotations, and aligned LaTeX, designed to support research on strike-out removal and handwritten mathematical formula recognition. Handwritten STEM manuscripts often contain strike-outs, overwrites, and other markings that disrupt mathematical content structure and reduce OCR/VLM recognition accuracy. This dataset provides a rigorous benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Incinciblecolonel/MathStrike.aiflow-math-ink-0.9-dataset
AIFlow Math Ink 0.9 Public Stroke Dataset
AIFlow Math Ink 0.9의 온라인 수식 필기 데이터다. 데이터 관리자는 2026-08-09까지 포함된 참여자 데이터의 공개 배포 승인을 확인했다. 원래 수집 동의 범위인 AIFlow 필기 인식 모델 학습·검증과 이번 공개 배포 승인을 함께 기록한다.
구성
파일
수식
용도
data/formulas_valid.jsonl
110
검수 완료 학습·분석 후보
data/formulas_pending.jsonl
39
시각 검수 전; 기본 학습 제외
data/formulas_reject.jsonl
5
검수 탈락; 기본 학습 제외
data/ownership_train.jsonl
47
수동 검수된 stroke-to-symbol ownership
전체 154개 수식, 11명의 dataset-local writer group… See the full description on the dataset page: https://huggingface.co/datasets/cwLeeDev/aiflow-math-ink-0.9-dataset.mathcaptchaforge-dataset-21k-parquet
Dataset Card
Overview
This dataset contains labeled image crops for multi-class visual symbol classification.
Structure
images/: image files referenced by the manifests
manifest.csv: full index
train.csv, val.csv, test.csv: split manifests
Schema
All CSV files use:
id,image_path,label,width,height,split
image_path is relative to the dataset root, formatted as images/<filename>.
Usage
Load one of the split CSV files.
Resolve image_path… See the full description on the dataset page: https://huggingface.co/datasets/emrecengdev/mathcaptchaforge-dataset-21k-parquet.
