CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ppak10 /AdditiveLLM2-OA AdditiveLLM2-OA Dataset Open Access journal articles (up to February 2026) used in domain adapting pretraining and instruction tuning for AdditiveLLM2. Dataset Split by Journal text images vit Vocabulary Overlap Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce. Top Phrases by Journal Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/AdditiveLLM2-OA.imagetext-generation10K<n<100K2 likes560 downloads6mo agoHugging Face02ppattnay /IndicSafe IndicSafe Authors: Priyaranjan Pattnayak, Garima Panwar, and Sanchari Chowdhuri. IndicSafe is a multilingual benchmark for evaluating large-language-model safety behavior across 12 South Asian languages. It contains 6,000 translated prompt rows: 500 source rows in each language, spanning harmful, harmless-control, and deliberately ambiguous categories. Content warning: the benchmark contains prompts about hate, discrimination, violence, misinformation, political manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ppattnay/IndicSafe.texttext-generation1K<n<10K0 likes180 downloads1mo agoHugging Face03ppalani09 /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ppalani09/pii-masking-300k.texttext-classification100K<n<1M0 likes36 downloads3mo agoHugging Face04ppan0423 /meetingbank Overview MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/ppan0423/meetingbank.textsummarization1K<n<10K0 likes20 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.