datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CLIcK
CLIcK 🇰🇷🧠
A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
Introduction 🎉
CLIcK (Cultural and Linguistic Intelligence in Korean) is a comprehensive dataset designed to evaluate cultural and linguistic intelligence in the context of Korean language models. In an era where diverse language models are continually emerging, there is a pressing need for robust evaluation datasets, especially for non-English languages like Korean. CLIcK… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/CLIcK.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.fitllm-fit-census
Local LLM Fit Census v1 — 2026-09-13
10,530 verdicts: 30 models × 93 devices (36 GPUs + 57 Mac configs) × per-platform quant tiers.
Each row is generated by fitllm-engine from architecture inputs pinned to official config.json files. Runtime and OS reserves remain documented estimates. Reproduce it yourself: npm run census.
Assumptions: context = min(8K, model max) · KV cache F16 · platform reserve/headroom per engine. Interactive per-combo pages: fitllm.run/can-i-run.… See the full description on the dataset page: https://huggingface.co/datasets/click6067/fitllm-fit-census.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.id_clickbait
This is the annotated full version of the dataset.
Dataset Summary
The CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news
publishers; detikNews, Fimela, Kapanlagi, Kompas, Liputan6, Okezone, Posmetro-Medan, Republika, Sindonews, Tempo,
Tribunnews, and Wowkeren. This dataset is comprised of mainly two parts; (i) 46,119 raw article data, and (ii)
15,000 clickbait annotated sample headlines. Annotation was conducted… See the full description on the dataset page: https://huggingface.co/datasets/manandey/id_clickbait.clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/clickbait_detection_dataset.clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/mayoooookha/clickbait_detection_dataset.clickbait_notclickbait_dataset0 : not clickbait
1 : clickbait
Dataset cleaned from duplicates and kept only the first appearing text.
Dataset split into train and test sets using 0.2 split ratio.
Dataset split into test and validation sets using 0.2 split ratio.
Size of training set: 43.802
Size of test set: 8.760
Size of validation set: 2.191
clickhouse-server-imageclickbait-spoilingData for Semeval 2023 task, clickbait spoiling
clickbait-spoiling-data-question
Webis Clickbait Spoiling Corpus
The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter.
This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post.
This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post.
Additionally… See the full description on the dataset page: https://huggingface.co/datasets/pramitsahoo/clickbait-spoiling-data-question.Clickbait_NewCLIcK
This dataset is the same as https://huggingface.co/datasets/EunsuKim/CLIcK. This dataset has been subdivided for simplified viewing and evaluation.
CLIcK 🇰🇷🧠
Evaluation of Cultural and Linguistic Intelligence in Korean
Introduction 🎉
CLIcK (Cultural and Linguistic Intelligence in Korean) is a comprehensive dataset designed to evaluate cultural and linguistic intelligence in the context of Korean language models. In an era where diverse… See the full description on the dataset page: https://huggingface.co/datasets/taeminlee/CLIcK.Multilingual_Clickbait_DatasetclickbaitProduct_reviews_click_stream_DataaCLIcK
CLIcK (Cultural and Linguistic Intelligence in Korean)
Mirror of the original CLIcK benchmark (LREC-COLING 2024, paper) with the per-item subcategory and broad-category labels preserved.
The upstream HuggingFace mirror at EunsuKim/CLIcK flattens every item into a single split and drops the category metadata that the paper uses. This mirror restores those labels by pulling each per-subcategory JSON file directly from the canonical GitHub repo and tagging each item with its source… See the full description on the dataset page: https://huggingface.co/datasets/bzantium/CLIcK.click_bate_random_sampleclick_bate_1000_train_testclickbait_spoilingsingle-click_bench
Single-Click Benchmark for Web Interaction
This benchmark defines a minimal web interaction task that requires only a single click to complete.
Each instance includes two task formulations (simplified and human-like), a pre-saved HTML file for obtaining screenshots or metadata, and target annotations with the element’s bounding box and XPath for evaluation.
The dataset enables systematic evaluation of web agents’ capabilities such as visual grounding, task understanding, and action… See the full description on the dataset page: https://huggingface.co/datasets/alexandrayakovleva/single-click_bench.webis-clickbait-spoiling-seq-tagindonesian-clickbait-spoilingclick_bate_articleclick_bate_1000click_bate_random_sample25k-dataset-geogpt-fineweb5k-dataset-geogpt-fineweb-randomclickhouse_changelogs_26.2clickbaitTest
