datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mscoco-controlnet-cannyal-kawakib-magazine-ocr
Al-Kawakib Magazine OCR Pages
This dataset contains rendered Arabic magazine page images paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments.
Fine-Tuning Notebook
A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here:
Open the fine-tuning notebook in Colab
The notebook can train on this dataset alone or on all three Arabic magazine OCR… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-kawakib-magazine-ocr.arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.DLAMA-v1
DLAMA-v1
A representative benchmark of factual triples curated from Wikidata and Wikipedia.
Predicate
Template
P17 (Country)
[X] is located in [Y] .
P19 (Place of birth)
[X] was born in [Y] .
P20 (Place of death)
[X] died in [Y] .
P27 (Country of citizenship)
[X] is [Y] citizen .
P30 (Continent)
[X] is located in [Y] .
P36 (Capital)
The capital of [X] is [Y] .
P37 (Official language)
The official language of [X] is [Y] .
P47 (Shares border with)
[X] shares… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/DLAMA-v1.al-lataif-al-musawwara-magazine-ocr
Al-Lataif Al-Musawwara Magazine OCR Pages
This dataset contains rendered Arabic magazine page images from Al-Lataif Al-Musawwara paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments.
Fine-Tuning Notebook
A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here:
Open the fine-tuning notebook in Colab
The notebook can train on this dataset alone or… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-lataif-al-musawwara-magazine-ocr.cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora
Two independently built pretraining corpora with their own tokenizers:
Malayalam as the higher-resource language and Assamese as the
lower-resource one. Nothing is shared between them — separate sources,
separate cleaning thresholds, separate vocabularies, separate models.
Only the language-agnostic pipeline code is common, parameterised per
language.
Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.cnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.capstone_sakuga_simple_description_mlm_hsmscoco-colour_masksarXiv-full-text-chunked-qacapstone_sakuga_simple_descriptionsakuga_preprocessedmnli-amr
Dataset Card for "mnli-amr"
More Information needed
phishing-email-rich-dataset-v2africa-synth-antibiotic-quality-amr-all
Antibiotic Quality & AMR Acceleration (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-antibiotic-quality-amr-all.UFPR-AMR
Dataset Card for "UFPR-AMR"
More Information needed
AMReasoning-100000
All Mathematical Reasoning-100000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-100000.vg_coco_overlap_for_graphormer_processed_amr_graphscapstone_sakuga_mlm_text_outputamr_portal
AMR Portal — Multi-Dataset Release (Phenotype + Genotype)
This repository contains multiple datasets from the EMBL-EBI AMR Portal, distributed in Apache Parquet format:
phenotype.parquet – phenotypic antimicrobial susceptibility data
genotype.parquet – AMR genes and mutations from in silico methods
All datasets are released under CC-BY-4.0.
Source documentation:
https://www.ebi.ac.uk/amr/developers/
Dataset Summary
This dataset contains phenotypic antimicrobial… See the full description on the dataset page: https://huggingface.co/datasets/ayates/amr_portal.UnitedNations-ParagraphsAlligned-ar-en-datasetAMReasoning-750000
All Mathematical Reasoning-750000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-750000.PTCCThe Parallel Tunisian Constitution Corpus (PTCC) corpus is a corpus of 149 articles written in Modern Standard Arabic and Tunisian Arabic.
Tesseract was used to transform the constitution's pdf files into text files. Afterward, alignment of the parallel articles was achieved by a simple Python script.
More details can be found in: https://amr-keleg.github.io/projects/digitalizing_dialectal_arabic/
Sources:
Tunisian Arabic translation of the 2014 Tunisian Constitution
2014… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/PTCC.arXiv-full-text-chunked-testAMReasoning-50000
All Mathematical Reasoning-50000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-50000.AMReasoning-2500000
All Mathematical Reasoning-2500000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-2500000.AMReasoning-250000
All Mathematical Reasoning-250000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-250000.AMRIT-Punjabi-Clinical-Dialogue-Corpus
ੴ AMRIT Punjabi Clinical Dialogue & Medical Diagnosis Corpus
☬ ਅੰਮ੍ਰਿਤ ਪੰਜਾਬੀ ਕਲੀਨਿਕਲ ਸੰਵਾਦ ਅਤੇ ਡਾਕਟਰੀ ਨਿਦਾਨ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Medical AI Architecture
Lead Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
Mission: Free Autonomous AI Doctor for Humanity (ਦੁਨੀਆਂ ਦੇ ਲੋੜਵੰਦ ਲੋਕਾਂ ਲਈ ਮੁਫ਼ਤ AI ਡਾਕਟਰ)
Flagship Platform: AMRIT Research OS (100% Local Medical Intelligence)
📖 Dataset Overview / ਸੰਖੇਪ
The… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/AMRIT-Punjabi-Clinical-Dialogue-Corpus.amr-3-parsed
Dataset Card for AMR 3.0 Parsed
Dataset Summary
This dataset contains parsed Abstract Meaning Representation (AMR) annotations from the LDC2020T02 release, formatted as instruction-following conversations. Each example consists of a sentence and its corresponding AMR graph representation.
Supported Tasks and Leaderboards
Tasks: Semantic parsing, specifically generating AMR graphs from English sentences
Leaderboards: AMR Parsing
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/hoshuhan/amr-3-parsed.AMReasoning-300000
All Mathematical Reasoning-300000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-300000.
