large-scale
vit_large_patch16_224_fmow_rgb_scalemaecsr-mxbai-embed-large-v1-nq-cos-sim-scale-5-gamma-1-detach-2csr-mxbai-embed-large-v1-nq-cos-sim-scale-20-gamma-1csr-mxbai-embed-large-v1-nq-cos-sim-scale-5-gamma-0.1csr-mxbai-embed-large-v1-nq-cos-sim-scale-20-gamma-1-detach-2csr-mxbai-embed-large-v1-nq-cos-sim-scale-5-gamma-0.1-detach-2csr-mxbai-embed-large-v1-nq-cos-sim-scale-50-gamma-0.1-detach-2csr-mxbai-embed-large-v1-nq-cos-sim-scale-20-gamma-0.5-detach-2
a-large-scale-fish-dataset
[!Note]
This is a copy version from A large scale fish dataset in Kaggle. And if you download this, you should follow the same license as original(CC-BY-NC4.0) and cite it as original readme says.
Original Dataset Authors : O. Ulucan, D. Karakaya, M. Turkan
For eaiser use, the dataset has been formated to 5 columns : image_id, image, mask, class_id, class_name, which is easier for you to load in hugging face and use.
A Large-Scale Dataset for Segmentation and Classification
Authors: O.… See the full description on the dataset page: https://huggingface.co/datasets/FriedParrot/a-large-scale-fish-dataset.opc-sft-stage1-largescale_diverse_instructlarge-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.large-scale-multimodal-multilingual-summarization-datasetPlease cite this paper if you use our code or data:
@inproceedings{verma-etal-2023-large,
title = "Large Scale Multi-Lingual Multi-Modal Summarization Dataset",
author = "Verma, Yash and
Jangra, Anubhav and
Verma, Raghvendra and
Saha, Sriparna",
booktitle = "Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics",
month = may,
year = "2023",
address = "Dubrovnik, Croatia",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/Zenquiorra/large-scale-multimodal-multilingual-summarization-dataset.EmergencyTrafficDetection_Large-Scale-Audio-dataset
