deprecated
decision-oaif_-_DEPRECATED-Meta-Llama-3-8B-Instruct-sft-webshop-iter0-pruned-demos-ggufdecision-oaif_-_DEPRECATED-Meta-Llama-3-8B-Instruct-sft-webshop-iter1-ggufAndy-4-base-DEPRECATEDcodebert-deprecatedyanolja_-_KoSOLAR-10.7B-v0.1-deprecated-ggufDEPRECATED-qwen-image-gguf-testdeprecated-gpn-arabidopsisKoSOLAR-10.7B-v0.1-deprecated-GGUF
deprecated-exp10-cot-leakagehome_deprecatedwikimedia-commons-documents-ml_deprecated
Wikimedia Commons Document Retrieval
Wikimedia Commons Documents
This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image.
The text_description column contains OCR text extracted from the images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.deprecated-github-code-haskell-function
Dataset Card for "github-code-haskell-function"
Rows: 3.26M
Download Size: 1.17GB
This dataset is extracted from github-code-haskell-file.
Each row has 3 flavors of the same function:
uncommented_code: Includes the function and its closest signature.
function_only_code: Includes the function only.
full_code: Includes the function and its closest signature and comment.
The heuristic for finding the closest signature and comment follows: If the immediate previous neighbor of the… See the full description on the dataset page: https://huggingface.co/datasets/blastwind/deprecated-github-code-haskell-function.crux-mds-deprecatedgithub-readme-retrieval-multilingual_deprecated
GitHub Readme Retrieval
This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_deprecated.
