datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deewaiREALCN-training
Repo
git@hf.co:datasets/telcom/deewaiREALCN-training
DeewaiREALCN Training Data
Image–text pairs for training captioning or vision–language models. Each image is a 1024×1024 RGB JPEG portrait with a short English description.
Contents
data/train/: 9,000 pairs for training.
images/: JPEG files (090000.jpg, …).
captions.jsonl: one JSON object per line with file_name and text.
data/val/: 1,000 pairs for validation with the same layout.
Example… See the full description on the dataset page: https://huggingface.co/datasets/javadtaghia/deewaiREALCN-training.JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.java_plum_leaf_disease_classification
Java Plum Leaf Disease Classification
A dataset for disease classification of Java Plum leaves. The dataset contains 2,400 images across 6 classes: Bacterial_Spot, Brown_Blight, Dry, Healthy, Powdery_Mildew, Sooty_Mold.Images per class:
Bacterial_Spot: 400
Brown_Blight: 400
Dry: 400
Healthy: 400
Powdery_Mildew: 400
Sooty_Mold: 400
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/java_plum_leaf_disease_classification.pixel-dataset-prerender-java-wikipediapixel-prerender-java-articleosworld_tasks_filesprerender-java-testpixel-dataset-prerender-java-sentence-wikipediaJavaNusaAksara-java
Dataset Card for "NusaAksara-java"
More Information needed
labeled-invoice-dfVQA_ON_LOCO
Fatima Fellowship Application
In this task, we are asked to identify the blind spots of a recently introduced model.
Choice of Model
For this task, I selected Qwen3-VL-4B-Instruct, which is described in its model card on Hugging Face as the most powerful vision-language model in the Qwen series to date. The model contains 4 billion parameters and was released four months prior to the writing of this document.
In the model card, the authors highlight several breakthroughs.… See the full description on the dataset page: https://huggingface.co/datasets/javadKV8/VQA_ON_LOCO.imagespixel-dataset-prerender-java-wikipedia-validationDreaming-dreamilydee-portrates
