CoolFace
Datasetpublic

jinaai/github-readme-retrieval-multilingual

GitHub Readme Retrieval This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns. The text_description column contains OCR text extracted from the images using EasyOCR. This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here. Disclaimer This dataset may… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes154downloads
Dataset Card

GitHub Readme Retrieval

This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns. The text_description column contains OCR text extracted from the images using EasyOCR. This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.

Disclaimer

This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked images, medical or financial data, or NSFW content), please notify us, and we will take appropriate action.

Copyright

All rights are reserved to the original authors of the documents.