CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amirveyseh /acronym_identification Dataset Card for Acronym Identification Dataset Dataset Summary This dataset contains the training, validation, and test data for the Shared Task 1: Acronym Identification of the AAAI-21 Workshop on Scientific Document Understanding. Supported Tasks and Leaderboards The dataset supports an acronym-identification task, where the aim is to predic which tokens in a pre-tokenized sentence correspond to acronyms. The dataset was released for a Shared Task which… See the full description on the dataset page: https://huggingface.co/datasets/amirveyseh/acronym_identification.texttoken-classification10K<n<100K23 likes65k downloads3y agoHugging Face02trishna0703 /potter-plant-identificationimage10K<n<100K0 likes7.4k downloads22d agoHugging Face03papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes3.8k downloads4y agoHugging Face04hkadxqq /spooky-author-identificationtext10K<n<100K0 likes1k downloads4y agoHugging Face05intelli-zen /language_identification 语种识别 Tips: 语种 zh 代表是中文, 可能是简体, 也可能是繁体. 语种 zh-cn 则代表是简体中文, zh-tw 代表繁体中文. 数据来源 数据集从网上收集整理如下: 多语言语料 数据 原始数据/项目地址 样本个数 原始数据描述 替代数据下载地址 amazon_reviews_multi Multilingual Amazon Reviews Corpus; 2010.02573 TRAIN: 1191160, VALID: 29665, TEST: 29685 我们提出了多语言亚马逊评论语料库 (MARC),这是用于多语言文本分类的大规模亚马逊评论集合。 该语料库包含 2015 年至 2019 年间收集的英语、日语、德语、法语、西班牙语和中文评论。 amazon_reviews_multi xnli XNLI; D18-1269.pdf TRAIN: 7702055, VALID: 49750, TEST: 100129 我们希望我们的数据集 XNLI… See the full description on the dataset page: https://huggingface.co/datasets/intelli-zen/language_identification.0 likes809 downloads2y agoHugging Face06Jensen-holm /Baseball-Identification baseball-detection-2 > 2023-06-02 3:09pm https://universe.roboflow.com/pitchtracking/baseball-detection-2 Provided by a Roboflow user License: CC BY 4.0 baseball-detection-2 - v4 2023-06-02 3:09pm This dataset was exported via roboflow.com on April 20, 2024 at 4:56 PM GMT Roboflow is an end-to-end computer vision platform that helps you collaborate with your team on computer vision projects collect & organize images understand and search unstructured image data annotate… See the full description on the dataset page: https://huggingface.co/datasets/Jensen-holm/Baseball-Identification.image1K<n<10K0 likes478 downloads2y agoHugging Face07Mikiee /Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform Person Detection and Re-Identification from Low Altitude UAV-based Platform Dataset Description This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format. The dataset supports two tasks: Person Detection — detecting people in aerial drone footage Person Re-Identification (Re-ID) — recognizing and… See the full description on the dataset page: https://huggingface.co/datasets/Mikiee/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.imageobject-detection1K<n<10K2 likes321 downloads4mo agoHugging Face08Alidr79 /cueless_EEG_subject_identification 🧠✨ Cueless EEG Imagined Speech for Subject Identification This repository hosts the dataset introduced in the paper: “Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks.” 🥳 Our work has been accepted by IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM) 🎉. 🧪💻 Code & Experiments All codes and experiments are available at 👉 https://github.com/Alidr79/cueless_EEG_subject_identification 📥 Downloading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Alidr79/cueless_EEG_subject_identification.text1K<n<10K1 likes239 downloads10mo agoHugging Face09UniDataPro /face-re-identification-image-dataset Dataset of face images with different angles and head positions Dataset contains 23,110 individuals, each contributing 28 images featuring various angles and head positions, diverse backgrounds, and attributes, along with 1 ID photo. In total, the dataset comprises over 670,000 images in formats such as JPG and PNG. It is designed to advance face recognition and facial recognition research, focusing on person re-identification and recognition systems. By utilizing this dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/face-re-identification-image-dataset.videoimage-segmentationn<1K2 likes192 downloads1mo agoHugging Face10thucdangvan020999 /speaker_identification_100_speakersaudio1K<n<10K0 likes181 downloads11mo agoHugging Face11SupraLabs /LLM-self-identification Self Identification – Give your Language model an identity About Self Identification Self identification is training set SupraLabs curated for developers/trainers to experiment with to let your language model know about their identity. Self identification let your LM know these information about them: Model ID Model Name Model Description Model Creator Model Family Model Architecture Parameter Count Knowledge Cutoff Here is an example from the dataset: If… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/LLM-self-identification.texttext-generationn<1K14 likes174 downloads2mo agoHugging Face12Abdelrahman-Rezk /Arabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets. Dataset Card for Arabic_Dialect_Identification Dataset Summary We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset relies on applying multiple filters to identify users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.tabular100K<n<1M12 likes165 downloads4y agoHugging Face13unklefedor /language-identificationtext100K<n<1M1 likes165 downloads3y agoHugging Face14incrisvel /urban-fire-identification Description Simplification of Disaster_Classification_Dataset. imageimage-classificationn<1K0 likes157 downloads4mo agoHugging Face15Qyrou /LLM-self-identification LLM Identity · Give your LLM an identity Self Identification The Self-Identification Dataset, curated by Qyrou, is a specialized training resource designed to help developers and trainers establish clear self-identity awareness within language models. By incorporating this dataset, models can accurately learn and convey essential metadata about themselves, including their Model ID, Model Name, Model Description, Model Creator, Model Family, Model Architecture, Parameter Count… See the full description on the dataset page: https://huggingface.co/datasets/Qyrou/LLM-self-identification.texttext-generationn<1K5 likes149 downloads1mo agoHugging Face16Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes138 downloads2y agoHugging Face17thucdangvan020999 /speaker_identification_100_speakers_audio1K<n<10K0 likes109 downloads11mo agoHugging Face18arubenruben /portuguese-language-identification-rawtext10M<n<100M0 likes101 downloads3y agoHugging Face19yash-ingle /ILID_Indian_Language_Identification_Dataset ILID: Native Script Language Identification for Indian Languages Paper | Code | Project Page 🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India 📄 Dataset Description The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.texttext-classification100K<n<1M0 likes95 downloads9mo agoHugging Face20Lots-of-LoRAs /task112_asset_simple_sentence_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task112_asset_simple_sentence_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task112_asset_simple_sentence_identification.texttext-generation1K<n<10K0 likes93 downloads2y agoHugging Face21macwiatrak /bacbench-operon-identification-protein-sequences Dataset for operon identification in bacteria (Protein sequences) A dataset of 4,073 operons across 11 bacterial genomes species. The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented by a list of protein sequences from different contigs. We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.textn<1K0 likes85 downloads1y agoHugging Face22yuotub /Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform Person Detection and Re-Identification from Low Altitude UAV-based Platform Dataset Description This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format. The dataset supports two tasks: Person Detection — detecting people in aerial drone footage Person Re-Identification (Re-ID) — recognizing… See the full description on the dataset page: https://huggingface.co/datasets/yuotub/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.imageobject-detection1K<n<10K0 likes74 downloads15d agoHugging Face23ScottishHaze /speaker-identification-toolkit Speaker Identification Toolkit This repository provides a comprehensive toolkit for processing audio and video files, with a focus on speaker diarization, speaker identification, audio extraction, and dataset creation. By leveraging tools like ffmpeg, pyannote.audio, and other Python libraries, the scripts enable efficient and accurate workflows for handling audio data. SEE GITHUB FOR UPDATES - I DON'T UPDATE THE FILES HERE ANYMORE ---… See the full description on the dataset page: https://huggingface.co/datasets/ScottishHaze/speaker-identification-toolkit.0 likes68 downloads2y agoHugging Face24VertexResearch /Vertex-0.6-35M-self-identification Vertex 0.6 35M — Self Identification A self-identification SFT dataset for Vertex-0.6-35M-Instruct: 459 ChatML-style conversations that teach the model who it is — its name, creator, family, architecture, parameter count, and knowledge cutoff. Made from SupraLabs/LLM-self-identification (Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 35M identity: Marker Value MODEL_ID VertexResearch/Vertex-0.6-35M-Instruct MODEL_NAME Vertex 0.6 35M… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-35M-self-identification.texttext-generationn<1K0 likes66 downloads28d agoHugging Face25macwiatrak /operon-identification-long-read-rna-sequencing-protein-sequences Dataset for operon identification from long-read RNA sequencing A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes located on the same transcripts. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list of protein sequences. Usage For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.textn<1K0 likes63 downloads1y agoHugging Face26broadinstitute /Domain_Identification_Algorithms_Comparison_Datatabular10M<n<100M0 likes59 downloads6mo agoHugging Face27pranavagrawal /Language-Identificationtext1M<n<10M0 likes55 downloads2y agoHugging Face28LynBean /wood-species-identificationimage1K<n<10K1 likes53 downloads2y agoHugging Face29ducut91 /Judgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다. SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능) 모델 정확도 재현율 F1 점수 GPT-4o(2024-08-06) 97.82 99.66 98.74 Qwen2.5-Max 96.46 95.83 96.14 DeepSeek-V3 98.73 98.92 98.81 Gemini-2.0-Flash 99.38 95.78 97.55 7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능 모델 파인튜닝 전 파인튜닝 후 정확도 재현율 F1 점수 정확도 재현율 F1 점수 EXAONE-3.5-7.8B-Instruct 68.26 67.89 68.08 98.59 94.4896.49 Ministral-8B-Instruct-2410 35.6 4.33 7.72 99.07 98.32 98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.text1K<n<10K0 likes53 downloads2y agoHugging Face30LemkinAI /Multimodal_Atrocity_Identification_Dataset LemkinAI Multimodal Atrocity Identification Dataset Dataset Overview This multimodal dataset contains comprehensive documentation of mass atrocities and human rights violations spanning 1980-2025, with 6.8+ million anonymized incident records from 195+ countries. The dataset combines textual documentation with satellite imagery and visual evidence for AI/ML research in atrocity detection, documentation, and prevention systems. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/LemkinAI/Multimodal_Atrocity_Identification_Dataset.image-classification1M<n<10M0 likes53 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.