datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/prabhats0605/scanned-images-dataset-for-ocr-and-vlm-finetuning.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa
Dataset Card: Africa Brain Tumor Scans with Synthetic EHR (Bundled Parquet)
This dataset bundles single-slice brain MRI scans and richly structured, synthetic EHR data into a single Parquet file suitable for multimodal ML research. Each row contains an image struct (bytes + path), a source label column, and an EHR payload with both a full JSON record and convenient summary columns.
The synthetic EHRs are Africa-focused: they encode country, urban/rural, facility level, insurance… See the full description on the dataset page: https://huggingface.co/datasets/saad02/brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
