CoolFace
19 results

deprecated

rlundqvist /deprecated-exp10-cot-leakage0 likes1.3k downloads4mo agoHugging Facefvdfs41 /home_deprecated0 likes666 downloads7mo agoHugging Facejinaai /wikimedia-commons-documents-ml_deprecated Wikimedia Commons Document Retrieval Wikimedia Commons Documents This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image. The text_description column contains OCR text extracted from the images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.image10K<n<100K1 likes222 downloads1y agoHugging Faceblastwind /deprecated-github-code-haskell-function Dataset Card for "github-code-haskell-function" Rows: 3.26M Download Size: 1.17GB This dataset is extracted from github-code-haskell-file. Each row has 3 flavors of the same function: uncommented_code: Includes the function and its closest signature. function_only_code: Includes the function only. full_code: Includes the function and its closest signature and comment. The heuristic for finding the closest signature and comment follows: If the immediate previous neighbor of the… See the full description on the dataset page: https://huggingface.co/datasets/blastwind/deprecated-github-code-haskell-function.tabulartext-generation1M<n<10M0 likes180 downloads3y agoHugging FaceDylanJHJ /crux-mds-deprecated0 likes116 downloads2y agoHugging Facejinaai /github-readme-retrieval-multilingual_deprecated GitHub Readme Retrieval This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns. The text_description column contains OCR text extracted from the images using EasyOCR. This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here. Disclaimer This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_deprecated.image10K<n<100K0 likes107 downloads1y agoHugging Face