datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unarXive-en2ru
Dataset Card for unarXive-en2ru
This dataset contains text excerpts from the unarXive citation recommendation dataset along with their translations into Russian. The translations have been obtained using OpenAI GPT-3.5-Turbo. The dataset is intended for machine translation research.
deniz-unay-school-seminar-evidence-media
Deniz UNAY – School Seminar & Media Evidence Index
A structured provenance index of publicly traceable seminar, institutional and media records associated with Deniz UNAY's work on technology addiction awareness, social media literacy and digital wellbeing in Turkey.
Scope
This repository is an evidence/provenance index, not an academic study, independent audit, ranking, or endorsement system. It preserves source-reported claims and separates them from… See the full description on the dataset page: https://huggingface.co/datasets/MrDen1234567890/deniz-unay-school-seminar-evidence-media.unalignment-toxic-dpo-v0.2-zh_cn数据集 unalignment/toxic-dpo-v0.2 的中英文对照版本。
这是一个高度有害的数据集,旨在通过很少的示例来说明如何使用 DPO 轻松地对模型进行去审查/取消对齐。
这份对照版本的中文来自多个不同模型的意译。转换的过程中,模型被允许对结果进行演绎以求通顺,无法对结果的准确性作任何保证。
使用限制请参照原数据集的 Usage restriction。
Original Dataset Description:
Toxic-DPO
This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples.
Many of the examples still contain some amount of… See the full description on the dataset page: https://huggingface.co/datasets/tastypear/unalignment-toxic-dpo-v0.2-zh_cn.Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful
Dataset Card for Data Privacy & Unauthorized Tracking Harmful
Description
The test set has been created to evaluate the robustness of a telecom chatbot specifically designed for the telecom industry. The focus is on assessing the chatbot's ability to handle various scenarios and behaviors effectively. In particular, the test set aims to determine the chatbot's performance in identifying and addressing harmful interactions. It also evaluates the chatbot's capability of… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful.TwitterHappiness
:smiley: Análisis de tweets de felicidad
El corpus para este proyecto es una recopilación de 10048 tuits obtenidos de la búsqueda del tag #felicidad. El corpus recopilado fue dado a 3 voluntarios a quienes se les pidió etiquetaran cada tuit, según su criterio, en los que expresaran alegría (A), publicidad (P), felicitaciones (F), consejos (C) y no alegría o sarcasmos (N). Al finalizar el etiquetado se realizó un filtro para obtener aquellos tuits que coincidian en más de una… See the full description on the dataset page: https://huggingface.co/datasets/GIL-UNAM/TwitterHappiness.sqli-rce-httpfs-unamenepse-stocks-unadjustedhhs-unaccompanied-alien-children-program
HHS Unaccompanied Alien Children Program
Description
This data represents unaccompanied alien children who are taken into custody by Customs and Border Protection brought to a facility and processed for transfer to the Department of Health and Human Services (HHS) as required by law. HHS holds the child for testing and quarantine, and shelters the child until the child is placed with a sponsor here in the United States.
Dataset Details
Publisher: HHS
Temporal… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/hhs-unaccompanied-alien-children-program.nepse-subindex-unadjustednepse-index-unadjusted
