CoolFace
Datasetpublic

lbourdois/OCR-nvidia-Nemotron-VLM-Dataset-v2_wiki_fr-clean

Description This dataset is a processed version of nvidia/Nemotron-VLM-Dataset-v2 to make it easier to use, particularly for a visual question answering task where answer is an OCR transcription.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than a path in an image column.We've also translated question column to French containing 40 prompts based on via tutoiement, vouvoiement and imperative forms. For further details, please… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/OCR-nvidia-Nemotron-VLM-Dataset-v2_wiki_fr-clean.

sourceHugging Facecc-by-sa-4.0updated 11mo agoView on Hugging Face
0likes184downloads
6 commits on main
df1de2211mo ago

Update README.md

lbourdois
9982cb311mo ago

Update README.md

lbourdois
290667411mo ago

Upload dataset

lbourdois
aba26f611mo ago

Upload dataset (part 00001-of-00002)

lbourdois
015327311mo ago

Upload dataset (part 00000-of-00002)

lbourdois
900ab6111mo ago

initial commit

lbourdois