datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GITQA-Aug-LegacyGITQA-Aug-Pruned-LegacyGITQA-Base-Pruned-LegacyGITQA-Base-LegacyGit-10MThe Git-10M dataset is a global-scale remote sensing image-text pair dataset, consisting of over 10 million image-text pairs with geographical locations and resolution information.
CC-BY-NC-ND-4.0 License: This dataset is not allowed to be modified or distributed without authorization!
Project Page: https://chen-yang-liu.github.io/Text2Earth/
View samples from the dataset
from datasets import load_dataset
import math
def XYZToLonLat(x,y,z):
# Transform… See the full description on the dataset page: https://huggingface.co/datasets/lcybuaa/Git-10M.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.ObjaverseXL_github_rendersinformes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.gitw-256x256
GITW (Glasses-in-the-Wild), re-split for 2D-keypoints-benchmark
A repackaging of the Glasses-in-the-Wild dataset
(DOI 10.5281/zenodo.17288503) for use in
2D-keypoints-benchmark. The images and keypoint
annotations are unchanged; only the splits and some metadata differ.
1000 crowdsourced 256x256 images of transparent and partially filled drinking glasses, from 11
participants over 60 scenes and 93 unique glass types. Each image shows one annotated glass, with
5 keypoints:… See the full description on the dataset page: https://huggingface.co/datasets/tlpss/gitw-256x256.github-readme-retrieval-multilingual
GitHub Readme Retrieval
This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual.git10m_resolution_le_16mpix_2.6m_datasetGitruck-LUT-15K
Gitruck LUT 15K
Gitruck LUT 15K is an authorized collection of .cube color lookup tables
prepared for LUT retrieval, similarity, classification, color-science analysis,
and generative modeling. The release contains viewable Parquet metadata,
diagnostic before/after previews, normalized tensors, and lossless raw archives.
Release summary
Item
Count
Source records
15,792
Unique raw byte streams
14,021
Unique normalized LUTs
13,365
Canonical train… See the full description on the dataset page: https://huggingface.co/datasets/Hocassian/Gitruck-LUT-15K.autotrain-data-mabama-dloragithub-readme-retrieval-multilingual_deprecated
GitHub Readme Retrieval
This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_deprecated.objaversexl_github6.5IntroSVG-trainSVG-MetaData
License
This dataset is licensed under the CC BY-NC 4.0 license.
Allowed: Research and educational use
Prohibited: Commercial use without explicit permission
Please cite this dataset if you use it in your research. For any commercial licensing inquiries, please contact the authors.
github_fetch_huggingface_terminal_9091_n3v8x2_source_gamma
Gamma Feedback Survey
Open survey responses from product feedback forms.
Dataset ID: SRC-GAMMA
Catalog: ghfht9091n3v8x2
Origin: Publicly distributed open survey responses
Records: 3,940
License: CC-BY-4.0
git-2024-vqagita_v1-datasetGithubsegformer_github_datawebsite-screenshots-git-largegithub-avatarsfilesystem_huggingface_terminal_git_5089_campaign_reportgit10m_resolution_1.1m_0.5mpix_datasetkpop_FacesGIT-TES-LEGENDS
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/Buckzor/GIT-TES-LEGENDS.osworld_tasks_filesgit10m_resolution_le_0.5mpix_dataset
