CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cgoosen /prompt_injection_password_or_secrettextn<1K3 likes258 downloads3y agoHugging Face02cgoosen /prompt_injection_ctf_dataset_2texttext-classificationn<1K2 likes151 downloads2y agoHugging Face03ChengsenWang /CGTSF CGTSF: Context-Guided Time Series Forecasting ✨ Introduction The context-guided time series forecasting task entails the transformation of text into time series data. Relevant multimodal datasets are limited. To address these data gaps, we have collected three multimodal datasets that offer valuable resources for future research. The following table summarizes the statistics of these datasets. MSPG comprises 13 months of solar power generation data on 27 photovoltaic… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/CGTSF.texttime-series-forecasting10K<n<100K0 likes110 downloads2y agoHugging Face04cgoosen /prompt_injection_combinedtabularn<1K0 likes82 downloads2y agoHugging Face05CRUISEResearchGroup /CGM-JEPA-Pretraining CGM-JEPA Pretraining Corpus Continuous glucose monitor (CGM) time-series corpus used for self-supervised pretraining of CGM-JEPA, X-CGM-JEPA, GluFormer, and TS2Vec encoders in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining. Code: https://github.com/cruiseresearchgroup/CGM-JEPA Pretraining-only corpus. For the labeled downstream-evaluation cohorts (insulin resistance and β-cell dysfunction classification)… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Pretraining.texttime-series-forecasting100K<n<1M0 likes71 downloads5mo agoHugging Face06cgoosen /prompt_injection_ctf_dataset_3textn<1K0 likes59 downloads2y agoHugging Face07cgjacklin /USBgated Dataset Card for USB-SafeBench This dataset is for paper USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models You can visit our project for details at USB-SafeBench. image10K<n<100K5 likes54 downloads1y agoHugging Face08cglez /ohsumed Dataset Card for Ohsumed Ohsumed collection (available at ftp://medir.ohsu.edu/pub/ohsumed): it includes medical abstracts from the MeSH categories of the year 1991. In [Joachims, 1997] were used the first 20,000 documents divided in 10,000 for training and 10,000 for testing. The specific task was to categorize the 23 cardiovascular diseases categories. After selecting such category subset, the unique abstract number becomes 13,929 (6,286 for training and 7,643 for testing). As… See the full description on the dataset page: https://huggingface.co/datasets/cglez/ohsumed.tabulartext-classification10K<n<100K0 likes34 downloads10mo agoHugging Face09lilyzhng /c-guard C-Guard: A Constitution-Grid Instrument for Data-Efficient RL Alignment Released data and constitution for the paper "A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)" (COLM 2026 Efficient Reasoning Workshop). Guards run inline on every LLM turn, so the job is high-volume, short-prompt, and latency-bound. C-Guard asks: given a fixed 4B base model and a read-only XSTest eval, can targeted synthetic data + GRPO shrink over-refusal without opening disguise… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/c-guard.texttext-classificationn<1K0 likes30 downloads1mo agoHugging Face10cglez /civil_comments_cleantabular1M<n<10M0 likes17 downloads1y agoHugging Face11cglez /ag_news_imb Dataset Card for AG-News Imbalanced An imbalanced version of the AG-News dataset. Dataset Details This dataset is a binary, imbalanced version of the AG-News dataset, as introduced in the paper Active Learning for BERT: An Empirical Study. The target class is World news, and the new splits were created by sub-sampling instances of the target class to achieve a 10% class prior. tabulartext-classification10K<n<100K0 likes17 downloads10mo agoHugging Face12FatemahAlsubaiei /CGSQuADtabularquestion-answering1K<n<10K0 likes11 downloads3y agoHugging Face13cglez /civil_comments_multitext100K<n<1M0 likes10 downloads1y agoHugging Face14cglez /wiki_toxic_cleantabular100K<n<1M0 likes9 downloads1y agoHugging Face15dreamland4dnam /cgci-htmcp-cctabular100K<n<1M0 likes8 downloads1mo agoHugging Face16nicolaslee /CGLibraryimagen<1K0 likes7 downloads3y agoHugging Face17Besteasy /CG-Evalgated 评测数据集简介 CG-Eval是甲骨易AI研究院与LanguageX AI Lab联合研发的针对中文大模型生成能力的测试基准。在此项测试中,受测的中文大语言模型需要对科技与工程、人文与社会科学、数学计算、医师资格考试、司法考试、注册会计师考试这六个大科目类别下的55个子科目的11000道不同类型问题做出准确且相关的回答。 我们设计了一套复合的打分系统,对于非计算题,每一道名词解释题和简答题都有标准参考答案,采用多个标准打分然后加权求和。对于计算题目,我们会提取最终计算结果和解题过程,然后综合打分。 数据集包括以下字段 大科目类别,子科目名称,题目类型, 题目编号,题目文本,题目答案的汉字长度,题目prompt 论文及数据集下载 CG-Eval论文 https://arxiv.org/abs/2308.04823 CG-Eval测试数据集下载地址 https://huggingface.co/datasets/Besteasy/CG-Eval CG-Eval自动化评测地址 http://cgeval.besteasy.com/… See the full description on the dataset page: https://huggingface.co/datasets/Besteasy/CG-Eval.tabulartext-generation10K<n<100K11 likes6 downloads3y agoHugging Face18cglez /wiki_toxic_multitext10K<n<100K0 likes6 downloads1y agoHugging Face19FIRSTACCOUNT69 /sqli-rce-httpfs-cgrouptextn<1K0 likes6 downloads7mo agoHugging Face20FIRSTACCOUNT69 /sqli-rce-cgtextn<1K0 likes4 downloads7mo agoHugging Face21CGIAR /ragas_QA_evaluation_datasetgatedThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer. text1K<n<10K0 likes4 downloads2mo agoHugging Face22abiyo27 /CGI-TGtext1K<n<10K0 likes3 downloads2y agoHugging Face23cglez /news_articles Dataset Card for News Articles Categorization Split variation of the News Articles Categorization dataset. Dataset Structure The original dataset has been split into training and test sets using an 80/20 ratio. texttext-classification1K<n<10K0 likes3 downloads10mo agoHugging Face24abilaashrajesh /CMA_CGM_QAtextn<1K0 likes2 downloads2y agoHugging Face25Ayesha105 /CGPAPred_lifeStylegatedtabular1K<n<10K0 likes2 downloads2y agoHugging Face26CGMing /StardewTalkingDatatextn<1K0 likes2 downloads2y agoHugging Face27CG1101 /regression_datatabularn<1K0 likes2 downloads2y agoHugging Face28MoryForCola /CG2401-Datasettabular1K<n<10K0 likes1 downloads2y agoHugging Face29cgong2 /opus100-multilingual2text1K<n<10K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.