datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt_injection_password_or_secretprompt_injection_ctf_dataset_2CGTSF
CGTSF: Context-Guided Time Series Forecasting
✨ Introduction
The context-guided time series forecasting task entails the transformation of text into time series data. Relevant multimodal datasets are limited. To address these data gaps, we have collected three multimodal datasets that offer valuable resources for future research. The following table summarizes the statistics of these datasets. MSPG comprises 13 months of solar power generation data on 27 photovoltaic… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/CGTSF.prompt_injection_combinedCGM-JEPA-Pretraining
CGM-JEPA Pretraining Corpus
Continuous glucose monitor (CGM) time-series corpus used for self-supervised pretraining of CGM-JEPA, X-CGM-JEPA, GluFormer, and TS2Vec encoders in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining.
Code: https://github.com/cruiseresearchgroup/CGM-JEPA
Pretraining-only corpus. For the labeled downstream-evaluation cohorts (insulin resistance and β-cell dysfunction classification)… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Pretraining.prompt_injection_ctf_dataset_3USB
Dataset Card for USB-SafeBench
This dataset is for paper USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
You can visit our project for details at USB-SafeBench.
ohsumed
Dataset Card for Ohsumed
Ohsumed collection (available at ftp://medir.ohsu.edu/pub/ohsumed): it includes medical abstracts from the MeSH categories of the year 1991. In [Joachims, 1997] were used the first 20,000 documents divided in 10,000 for training and 10,000 for testing. The specific task was to categorize the 23 cardiovascular diseases categories. After selecting such category subset, the unique abstract number becomes 13,929 (6,286 for training and 7,643 for testing). As… See the full description on the dataset page: https://huggingface.co/datasets/cglez/ohsumed.c-guard
C-Guard: A Constitution-Grid Instrument for Data-Efficient RL Alignment
Released data and constitution for the paper "A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)" (COLM 2026 Efficient Reasoning Workshop).
Guards run inline on every LLM turn, so the job is high-volume, short-prompt, and latency-bound. C-Guard asks: given a fixed 4B base model and a read-only XSTest eval, can targeted synthetic data + GRPO shrink over-refusal without opening disguise… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/c-guard.civil_comments_cleanag_news_imb
Dataset Card for AG-News Imbalanced
An imbalanced version of the AG-News dataset.
Dataset Details
This dataset is a binary, imbalanced version of the AG-News dataset, as introduced in the paper
Active Learning for BERT: An Empirical Study.
The target class is World news, and the new splits were created by sub-sampling instances
of the target class to achieve a 10% class prior.
CGSQuADcivil_comments_multiwiki_toxic_cleancgci-htmcp-ccCGLibraryCG-Eval
评测数据集简介
CG-Eval是甲骨易AI研究院与LanguageX AI Lab联合研发的针对中文大模型生成能力的测试基准。在此项测试中,受测的中文大语言模型需要对科技与工程、人文与社会科学、数学计算、医师资格考试、司法考试、注册会计师考试这六个大科目类别下的55个子科目的11000道不同类型问题做出准确且相关的回答。 我们设计了一套复合的打分系统,对于非计算题,每一道名词解释题和简答题都有标准参考答案,采用多个标准打分然后加权求和。对于计算题目,我们会提取最终计算结果和解题过程,然后综合打分。
数据集包括以下字段
大科目类别,子科目名称,题目类型, 题目编号,题目文本,题目答案的汉字长度,题目prompt
论文及数据集下载
CG-Eval论文 https://arxiv.org/abs/2308.04823
CG-Eval测试数据集下载地址 https://huggingface.co/datasets/Besteasy/CG-Eval
CG-Eval自动化评测地址 http://cgeval.besteasy.com/… See the full description on the dataset page: https://huggingface.co/datasets/Besteasy/CG-Eval.wiki_toxic_multisqli-rce-httpfs-cgroupsqli-rce-cgragas_QA_evaluation_datasetThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer.
CGI-TGnews_articles
Dataset Card for News Articles Categorization
Split variation of the News Articles Categorization dataset.
Dataset Structure
The original dataset has been split into training and test sets using an 80/20 ratio.
CMA_CGM_QACGPAPred_lifeStyleStardewTalkingDataregression_dataCG2401-Datasetopus100-multilingual2
