datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.anime-tagging-datasetpos_tagging
POS Tagging Dataset
Original Data Source
Conll2003
E. F. Tjong Kim Sang and F. De Meulder, Proceedings of the
Seventh Conference on Natural Language Learning at HLT-
NAACL 2003, 2003, pp. 142–147.
The Peen Treebank
M. P. Marcus, B. Santorini and M. A. Marcinkiewicz, Comput.
Linguist., 1993, 19, 313–330.
Citation
BatteryDataExtractor: battery-aware text-mining software embedded with BERT models
top_tagging
Dataset Card for Top Quark Tagging
Dataset Summary
Top Quark Tagging is a dataset of Monte Carlo simulated events produced by proton-proton collisions at the Large Hadron Collider. The top-quark signal and mixed quark-gluon background jets are produced with Pythia8 with its default tune for a center-of-mass energy of 14 TeV. Multiple interactions and pile-up are ignored. The leading 200 jet constituent four-momenta (E,px,py,pz) (E, p_x, p_y, p_z) (E,px,py,pz)are stored… See the full description on the dataset page: https://huggingface.co/datasets/dl4phys/top_tagging.qg-tagging-normalized
Dataset Card for "qg-tagging-normalized"
More Information needed
task583_udeps_eng_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task583_udeps_eng_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task583_udeps_eng_coarse_pos_tagging.top_quark_taggingTop Quark Tagging is a dataset of Monte Carlo simulated hadronic top and QCD dijet events for the evaluation of top quark tagging architectures. The dataset consists of 1.2M training events, 400k validation events and 400k test events.task584_udeps_eng_fine_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task584_udeps_eng_fine_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task584_udeps_eng_fine_pos_tagging.task1168_brown_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1168_brown_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1168_brown_coarse_pos_tagging.X-Ray_Community_Tagging
What is this?
A community effort is a must in order to make a better, more accurate vision model, as I simply cannot tag thousands of images. If you would provide 50 corrections and 20 more people do so as well, it would help a lot.
If 100 ppl would help with 50 corrections each, we might have a high-accuracy functioning uncensored vision model.
The best format would be to name the output and images with the same name, like:
1.png
1.txt
2.png
2.txt
The best approach is probably… See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/X-Ray_Community_Tagging.task1167_penn_treebank_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.Music_Taggingtop_tagging_imagesqg-tagging
Dataset Card for "qg-tagging"
More Information needed
sanskrit-unsandhi-morphosyntax-taggingtop_quark_tagging_oldproduct-tagging-saved-intent
Product Tagging and Saved Intent in Online Retail
How online stores structure product tags and implement saved-item features, and
how that compares with a platform where tagging is crowd-sourced against a
shared controlled vocabulary.
Canonical release: https://doi.org/10.5281/zenodo.22852049
This repository mirrors that deposit. Cite the DOI.
Sample
1,028 Shopify storefronts, drawn by seeded random sample from a 30,000-domain
draw of the Tranco top 1M, measured… See the full description on the dataset page: https://huggingface.co/datasets/sempite/product-tagging-saved-intent.moodwave-music-tagging
MoodWave Music Tagging
A multilingual music-tagging dataset built for Hindi and Nepali repertoire,
which existing music models handle poorly. 20,939 tracks, 1,763 hours.
This repository contains metadata only — no audio. Each row carries a
YouTube ID (or a content hash for legacy entries) plus labels and provenance.
See Getting the audio.
Status: work in progress. Collection is ongoing and genre labels currently
cover the English portion only. See Current state before
using… See the full description on the dataset page: https://huggingface.co/datasets/anujpaude1/moodwave-music-tagging.hotel-tagging-finetune
Hotel Tagging Finetune
Hotel-room images for a tagging finetune workflow where an LLM is used as the judge.
This public release intentionally contains only the images.
Splits
Split
Samples
train
2000
test
250
Columns
image: hotel-room image
qg-tagging-discrete
Dataset Card for "qg-tagging-discrete"
More Information needed
grocery-image-tagging
grocery-image-tagging
Pipeline for tagging grocery product photos as product (front of package) or
ingredients (back of package / nutrition focus).
tag_products.py
Zero-shot classifier using google/siglip2-base-patch16-256. No training required.
Validated end-to-end against a JSON-LD @graph metadata file and an ndjson variant.
pip install "transformers>=4.49" torch pillow
python tag_products.py --image-dir ./images --metajson ./metajson.ld \
--output… See the full description on the dataset page: https://huggingface.co/datasets/bdalziel/grocery-image-tagging.tagging_datapos-tagging-malagasy-sokajy
FITSIPIKA Malagasy Dataset (SOKAJY)
Dataset Description
SOKAJY is a specialized morphosyntactic corpus for the Malagasy language (28M speakers). It focuses on Part-of-Speech (POS) tagging and linguistic structure analysis, specifically designed to handle the unique syntactic challenges of the Malagasy language.
Curators: Vatosoa Razafindrazaka (Madagascar)
Format: CoNLL-U
Language: Official Malagasy (Merina and regional variants)
Status: Actively maintained for… See the full description on the dataset page: https://huggingface.co/datasets/Vatosoa/pos-tagging-malagasy-sokajy.sanskrit-unsandhi-lemma-morphosyntax-tagging-paragraphKorean-FineTome-100k-tagginglemon-mint/Korean-FineTome-100k를 tagging한 데이터입니다.
system message과 있는 것과 없는 것을 구분했습니다.
다음은 구분표입니다.
system message tag
분류
설명
External Function Access Restriction
AI가 외부 기능에 접근할 수 없음을 강조
Following Instructions Well
사용자의 지시를 정확하게 따름을 강조
No Censorship or Bias
검열 없이 편향되지 않은 정보를 제공
Helpful Assistant
사용자에게 효과적으로 도움을 주는 역할 강조
AI Assistant Identity
AI가 인간이 아니라 AI 비서임을 명확히 함
Step-by-Step Explanation
답변을 논리적인 단계로 나누어 설명
Detailed and Lengthy Answers
충분히 상세하고 긴 답변 제공… See the full description on the dataset page: https://huggingface.co/datasets/nwirandx/Korean-FineTome-100k-tagging.240903-image-taggingsanskrit-unsandhi-morphosyntax-tagging-paragraphTagging_Data_Full_63k
Dataset Card for "Tagging_Data_Full_63k"
More Information needed
free-ai-auto-tagging-django-nexaapi
Free AI Auto-Tagging API for Django Developers — NexaAPI Free Tier
See README.md for the full tutorial.
Links
🌐 NexaAPI: https://nexa-api.com
🔑 Free API Key: https://rapidapi.com/user/nexaquency
🐍 Python SDK: https://pypi.org/project/nexaapi/
📦 Node.js SDK: https://npmjs.com/package/nexaapi
tagging_thai
