datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEW-Multimodal-AMR
SEW Multimodal AMR → FiftyOne (Native Multimodal MCAP)
The labeled test set of the
SEW-EURODRIVE Multimodal AMR dataset,
converted to native multimodal MCAP episodes. Six sensing modalities ride one
autonomous mobile robot: RGB, thermal, time-of-flight, 4D radar, two 2D laser
scanners, and an ultrasonic array. The 3,151 labeled frames are split into 55
episodes, one per source recording session, spanning three seasons, six
weather conditions, and day, dawn and night.… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/SEW-Multimodal-AMR.mscoco-controlnet-cannyal-kawakib-magazine-ocr
Al-Kawakib Magazine OCR Pages
This dataset contains rendered Arabic magazine page images paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments.
Fine-Tuning Notebook
A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here:
Open the fine-tuning notebook in Colab
The notebook can train on this dataset alone or on all three Arabic magazine OCR… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-kawakib-magazine-ocr.DLAMA-v1
DLAMA-v1
A representative benchmark of factual triples curated from Wikidata and Wikipedia.
Predicate
Template
P17 (Country)
[X] is located in [Y] .
P19 (Place of birth)
[X] was born in [Y] .
P20 (Place of death)
[X] died in [Y] .
P27 (Country of citizenship)
[X] is [Y] citizen .
P30 (Continent)
[X] is located in [Y] .
P36 (Capital)
The capital of [X] is [Y] .
P37 (Official language)
The official language of [X] is [Y] .
P47 (Shares border with)
[X] shares… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/DLAMA-v1.arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.pacerDataFurnace-AMR
🏭 DataFurnace: AMR & SLAM Synthetic Evaluation Dataset
Fully synchronized multi-pass rendering (RGB, Depth, Semantic Mask, and 3D Bounding Box).
📌 Overview
DataFurnace is a procedurally generated synthetic dataset designed for evaluating and training:
Autonomous Mobile Robots (AMR)
Automated Guided Vehicles (AGV)
SLAM / VIO / 3D Perception algorithms
Warehouse robotics navigation systems
All ground truth is generated directly from the underlying 3D… See the full description on the dataset page: https://huggingface.co/datasets/jp-cypress/DataFurnace-AMR.so100_make_coffeetest6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 39135,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AMRUBARD/so100_make_coffeetest6.al-lataif-al-musawwara-magazine-ocr
Al-Lataif Al-Musawwara Magazine OCR Pages
This dataset contains rendered Arabic magazine page images from Al-Lataif Al-Musawwara paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments.
Fine-Tuning Notebook
A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here:
Open the fine-tuning notebook in Colab
The notebook can train on this dataset alone or… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-lataif-al-musawwara-magazine-ocr.cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora
Two independently built pretraining corpora with their own tokenizers:
Malayalam as the higher-resource language and Assamese as the
lower-resource one. Nothing is shared between them — separate sources,
separate cleaning thresholds, separate vocabularies, separate models.
Only the language-agnostic pipeline code is common, parameterised per
language.
Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.chemistrycnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.capstone_sakuga_simple_description_mlm_hsarXiv-full-text-chunked-qacapstone_sakuga_simple_descriptionnigeria-amr-classifier-v2-dataset
Nigeria AMR Classifier v2 Dataset
Author: Hussein Adeiza (mabera)
Role: Licensed Environmental Health Officer, Abuja Nigeria
Built for: AutoScientist Challenge 2026, Part 2 — Science Category
Dataset Description
A closed-label antimicrobial resistance classification dataset combining
two independently verifiable task types: individual MIC-based
susceptibility classification (CLSI M100 breakpoints) and population-level
resistance rate classification from real… See the full description on the dataset page: https://huggingface.co/datasets/mabera/nigeria-amr-classifier-v2-dataset.sakuga_preprocessedmscoco-colour_masksObjectPosemnli-amr
Dataset Card for "mnli-amr"
More Information needed
grasp_corpus_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/amrltqt/grasp_corpus_v3.capstone_sakuga_mlm_text_outputphishing-email-rich-dataset-v2robotics-corpus-2023
Robotics Sensor Fusion Data Notes
Dataset summary
A documented Robotics data-preparation workflow for Sensor Fusion records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/Amritastatistics/robotics-corpus-2023.africa-synth-antibiotic-quality-amr-all
Antibiotic Quality & AMR Acceleration (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-antibiotic-quality-amr-all.UFPR-AMR
Dataset Card for "UFPR-AMR"
More Information needed
AMReasoning-100000
All Mathematical Reasoning-100000
Built via a python script
Contents
add
sub
mul
div
linear_eq
two_step_eq
fraction
exponent
inequality
word_algebra
quadratic
system_2x2
abs_eq
percent
mod
simplify_expr
mixed_fraction
neg_div
linear_fraction_eq
rational_eq
quadratic_nonunit
cubic_int_root
system_3x3
diophantine
exponential_eq
log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-100000.vg_coco_overlap_for_graphormer_processed_amr_graphscomics-sensor-fusion-mini
Comics Sensor Fusion Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Sensor Fusion inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/amritastatistics04/comics-sensor-fusion-mini.RPSeg
