datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.twinicl-bench
TwinICL
38 tasks, each with 132 underlying examples rendered in eight variants: 40,128 rows in total.
Each row contains only:
task: a readable task name.
variant: the text style or image palette.
input_text: the text input, or null for image examples.
input_image: the image input, or null for text examples.
answer: the expected text answer.
The eight variants are lowercase/comma, lowercase/semicolon, uppercase/comma, uppercase/semicolon, and images in neutral, warm, cool, and… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/twinicl-bench.tw-privacy-guides
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-privacy-guides.SAKE-Twittertw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.devflow-finance-twin
devflow-finance-twin
Sovereign BaaS ledger stack — IBM i authoritative core, C# API layer, COBOL FSL state machines, RPGLE posting engine.
Architecture
C# REST API
└── LedgerGateway / CobolGateway (binary struct marshal → IBM i program call)
└── COBOL FSL Supervisors (TXN-FSL / ACH-FSL / RTP-FSL / LEDGER-GATEWAY)
└── RPGLE Programs (POSTTRAN / REVTRAN / ADJTRAN / EOD / TREASURY)
└── DB2 for i (authoritative… See the full description on the dataset page: https://huggingface.co/datasets/SNAPKITTYWEST/devflow-finance-twin.twin_handover_256_traindevflow-finance-twin
devflow-finance-twin
Sovereign BaaS ledger stack — IBM i authoritative core, C# API layer, COBOL FSL state machines, RPGLE posting engine.
Architecture
C# REST API
└── LedgerGateway / CobolGateway (binary struct marshal → IBM i program call)
└── COBOL FSL Supervisors (TXN-FSL / ACH-FSL / RTP-FSL / LEDGER-GATEWAY)
└── RPGLE Programs (POSTTRAN / REVTRAN / ADJTRAN / EOD / TREASURY)
└── DB2 for i (authoritative… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/devflow-finance-twin.twin_put_item_in_drawer_256_valtwin_handover_item_easy_256_traintwin_straighten_rope_256_traintwin_straighten_rope_256_valtwitch_streamers
Twitch Streamers Dataset
A dataset of Twitch streamers with their social links, follower counts, and recent games.
Stats
Total Streamers: 15,152
Partners: 12,999
Affiliates: 1,663
Most Popular Game: Just Chatting
Most Linked Platform: Youtube
Charts
Follower Distribution
Top 20 Streamers
Social Platform Popularity
Top 20 Games
Partner/Affiliate Distribution… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/twitch_streamers.twin_lift_ball_256_traintwin_take_tray_out_of_oven_256_valtwin_sweep_to_dustpan_256_valtwin_put_bottle_in_fridge_256_valtwin_handover_item_easy_256_valtwist_nine_class_eval_datasetBLINK-Twice
BLINK-Twice: You see, but you do not observe. A Reasoning Benchmark on Visual Perception
📌 About BLINK-Twice
BLINK-Twice Task Overview: (a) Visual reasoning task requiring detailed observation and careful reasoning; (b) Natural adversarial samples with similar appearance but opposite semantics, forcing models to rely on visual input; (c) Reasoning step annotation including detailed visual clues and true reality to evaluate thought chain output.
As illustrated… See the full description on the dataset page: https://huggingface.co/datasets/PicoTrex/BLINK-Twice.twin_take_tray_out_of_oven_256_testTwiFF-Bench
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
🧠 Method
We present TwiFF, a unified model fine-tuned on a high-quality dynamic visual Chain-of-Thought (VCoT) dataset comprising 2.7 million samples. In dynamic multimodal question-answering tasks involving instructional, predictive, and camera, TwiFF iteratively generates future event frames alongside textual reasoning, thereby producing… See the full description on the dataset page: https://huggingface.co/datasets/Liu-Junhua/TwiFF-Bench.twin_pick_laptop_256_testTWIN
TWIN
This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
For evaluating on the dataset with LMMS-eval, please refer to this repo.
Citation
If you use the TWIN dataset in your research, please use the following BibTeX entry.
@misc{marsili2025notenhancingvisualperception… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/TWIN.twin_push_box_256_trainTwitter_AI
VISUAL COUNTER TURING TEST (VCT²) — TWITTER DATASET
The Visual Counter Turing Test (VCT²) dataset is introduced in the paper“Visual Counter Turing Test (VCT²): Discovering the Challenges for AI-Generated Image Detection and Introducing Visual AI Index (V_AI)”,accepted at IJCNLP–AACL 2025 and available on arXiv:2411.16754.
This dataset aims to benchmark and analyze the challenges of AI-generated image detection (AGID) using real-world, social media–driven captions and imagery.It… See the full description on the dataset page: https://huggingface.co/datasets/NasrinImp/Twitter_AI.banner-assets
Twinkle AI — Banner Assets
A collection of banner images featuring the Twinkle AI mascot, hosted on Hugging Face and stored via Git LFS.
Available Banners
File
Preview
images/banner-twinkle-hf.jpeg
images/Twinkle-AI-First-Birthday-Party.png
images/TwinkleAI--3n.png
images/TwinkleAI_Reading_Club_Presentation.png
images/TwinkleAI-Reading-Club-v2.png
images/Twinkle@SITCON.png
images/Twinkle-red-envelope.png
images/Twinkle-red-envelope1.png… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/banner-assets.twin_handover_item_easy_256_testtwin_dual_push_128_trainsynesthesia-sft
