CoolFace
Datasetpublic

lbourdois/OCR-nvidia-Nemotron-VLM-Dataset-v2_wiki_fr-clean

Description This dataset is a processed version of nvidia/Nemotron-VLM-Dataset-v2 to make it easier to use, particularly for a visual question answering task where answer is an OCR transcription.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than a path in an image column.We've also translated question column to French containing 40 prompts based on via tutoiement, vouvoiement and imperative forms. For further details, please… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/OCR-nvidia-Nemotron-VLM-Dataset-v2_wiki_fr-clean.

sourceHugging Facecc-by-sa-4.0updated 11mo agoView on Hugging Face
0likes184downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face