datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acronym_identification
Dataset Card for Acronym Identification Dataset
Dataset Summary
This dataset contains the training, validation, and test data for the Shared Task 1: Acronym Identification of the AAAI-21 Workshop on Scientific Document Understanding.
Supported Tasks and Leaderboards
The dataset supports an acronym-identification task, where the aim is to predic which tokens in a pre-tokenized sentence correspond to acronyms. The dataset was released for a Shared Task which… See the full description on the dataset page: https://huggingface.co/datasets/amirveyseh/acronym_identification.potter-plant-identificationlanguage-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.spooky-author-identificationBaseball-Identification
baseball-detection-2 > 2023-06-02 3:09pm
https://universe.roboflow.com/pitchtracking/baseball-detection-2
Provided by a Roboflow user
License: CC BY 4.0
baseball-detection-2 - v4 2023-06-02 3:09pm
This dataset was exported via roboflow.com on April 20, 2024 at 4:56 PM GMT
Roboflow is an end-to-end computer vision platform that helps you
collaborate with your team on computer vision projects
collect & organize images
understand and search unstructured image data
annotate… See the full description on the dataset page: https://huggingface.co/datasets/Jensen-holm/Baseball-Identification.Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing and… See the full description on the dataset page: https://huggingface.co/datasets/Mikiee/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.cueless_EEG_subject_identification
🧠✨ Cueless EEG Imagined Speech for Subject Identification
This repository hosts the dataset introduced in the paper:
“Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks.”
🥳 Our work has been accepted by IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM) 🎉.
🧪💻 Code & Experiments
All codes and experiments are available at 👉 https://github.com/Alidr79/cueless_EEG_subject_identification
📥 Downloading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Alidr79/cueless_EEG_subject_identification.LLM-self-identification Self Identification – Give your Language model an identity
About Self Identification
Self identification is training set SupraLabs curated for developers/trainers to experiment with to let your language model know about their identity.
Self identification let your LM know these information about them:
Model ID
Model Name
Model Description
Model Creator
Model Family
Model Architecture
Parameter Count
Knowledge Cutoff
Here is an example from the dataset:
If… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/LLM-self-identification.speaker_identification_100_speakersArabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets.
Dataset Card for Arabic_Dialect_Identification
Dataset Summary
We present QADI, an automatically collected dataset of tweets belonging to a wide range of
country-level Arabic dialects covering 18 different countries in the Middle East and North
Africa region. Our method for building this dataset relies on applying multiple filters to identify
users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.language-identificationLLM-self-identification
LLM Identity · Give your LLM an identity
Self Identification
The Self-Identification Dataset, curated by Qyrou, is a specialized training resource designed to help developers and trainers establish clear self-identity awareness within language models. By incorporating this dataset, models can accurately learn and convey essential metadata about themselves, including their Model ID, Model Name, Model Description, Model Creator, Model Family, Model Architecture, Parameter Count… See the full description on the dataset page: https://huggingface.co/datasets/Qyrou/LLM-self-identification.task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.speaker_identification_100_speakers_portuguese-language-identification-rawILID_Indian_Language_Identification_Dataset
ILID: Native Script Language Identification for Indian Languages
Paper | Code | Project Page
🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India
📄 Dataset Description
The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.task112_asset_simple_sentence_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task112_asset_simple_sentence_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task112_asset_simple_sentence_identification.bacbench-operon-identification-protein-sequences
Dataset for operon identification in bacteria (Protein sequences)
A dataset of 4,073 operons across 11 bacterial genomes species.
The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented
by a list of protein sequences from different contigs.
We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing… See the full description on the dataset page: https://huggingface.co/datasets/yuotub/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.Vertex-0.6-35M-self-identification
Vertex 0.6 35M — Self Identification
A self-identification SFT dataset for Vertex-0.6-35M-Instruct:
459 ChatML-style conversations that teach the model who it is — its name,
creator, family, architecture, parameter count, and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6
35M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-35M-Instruct
MODEL_NAME
Vertex 0.6 35M… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-35M-self-identification.Domain_Identification_Algorithms_Comparison_Dataoperon-identification-long-read-rna-sequencing-protein-sequences
Dataset for operon identification from long-read RNA sequencing
A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes
located on the same transcripts.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list
of protein sequences.
Usage
For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.complex_word_identificationsource: https://www.inf.uni-hamburg.de/en/inst/ab/lt/resources/data/complex-word-identification-dataset.html
Language-IdentificationJudgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다.
SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능)
모델
정확도
재현율
F1 점수
GPT-4o(2024-08-06)
97.82
99.66
98.74
Qwen2.5-Max
96.46
95.83
96.14
DeepSeek-V3
98.73
98.92
98.81
Gemini-2.0-Flash
99.38
95.78
97.55
7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능
모델
파인튜닝 전
파인튜닝 후
정확도
재현율
F1 점수
정확도
재현율
F1 점수
EXAONE-3.5-7.8B-Instruct
68.26
67.89
68.08
98.59
94.4896.49
Ministral-8B-Instruct-2410
35.6
4.33
7.72
99.07
98.32
98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.indian-ipc-statute-identification
Indian IPC Statute Identification
Given the facts of an Indian court case, identify the relevant Indian Penal Code (IPC) section. Each example
pairs the factual narrative of a High Court judgment with the text of an IPC section that the judgment applies.
The task is framed as retrieval / statute identification: from a fact scenario, retrieve (or classify) the governing
statute. It is a useful benchmark and training signal for legal information retrieval, legal text… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/indian-ipc-statute-identification.LID201_Devanagari_Script_Languages_IdentificationInduction-Cooker-Ceramic-Panel-Crack-Identification-Dataset
Induction Cooker Ceramic Panel Crack Identification Dataset
In the current industrial field, the crack problem of induction cooker ceramic panels poses a threat to product safety, leading to potential explosion risks. Existing detection methods mostly rely on manual inspection, which is inefficient and prone to errors. This dataset aims to provide high-quality crack image data to train machine learning models, automating the detection process and improving detection efficiency and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Induction-Cooker-Ceramic-Panel-Crack-Identification-Dataset.south_african_language_identificationLanguage_Identification
