datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
singaporean-judicial-keywords
Singaporean Judicial Keywords 🏛️
Singaporean Judicial Keywords by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 500 catchword-judgment pairs sourced from the Singapore Judiciary.
Uniquely, the keywords in this dataset are real-world annotations created by subject matter experts, namely, Singaporean law reporters, as opposed to being constructed ex post facto by third parties.
Additionally, unlike standard keyword queries, judicial catchwords… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/singaporean-judicial-keywords.jailbreak-dan-keywords
LLM Jailbreak & DAN Detection Dataset
A comprehensive dataset of adversarial prompts, jailbreak triggers, DAN (Do Anything Now) templates, prompt injection vectors, and refusal suppression patterns for Large Language Models (LLMs). This dataset is specifically formatted row-by-row for Hugging Face Datasets Viewer, AI Guardrail filters, Prompt Injection defense systems, and Automated Red Teaming suites.
Technical Specifications
Parameter
Value
Primary… See the full description on the dataset page: https://huggingface.co/datasets/mondk/jailbreak-dan-keywords.summary-keywordswikipedia-paragraph-keywords
Wikipedia Paragraph and Keyword Dataset
Dataset Summary
This dataset contains 10,693 paragraphs extracted from English Wikipedia articles, along with corresponding search-engine style keywords for each paragraph. It is designed to support tasks such as text summarization, keyword extraction, and information retrieval.
Dataset Structure
The dataset is structured as a collection of JSON objects, each representing a single paragraph with its associated keywords.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-keywords.12345-dialogue-keywordswikipedia_keywordsNCS-TAG-Training-keywords-teststw-law-context-keywords
Dataset Card for tw-law-context-keywords
本資料集為中華民國(臺灣)法規之 LLM 關鍵字抽取結果,每筆樣本對應一部法規的條列式關鍵字清單,可作為法規檢索、Tag 化、向量化前置處理之素材。
Dataset Details
Dataset Description
資料以法規為單位,由 LLM 對每部法規抽取「核心概念、條文要點、重要術語、特殊註記(如『廢止』)」等關鍵字。可用於:
法規檢索系統的 keyword index。
對 RAG 流程中的 chunk 預先附加關鍵字 metadata。
訓練法律術語抽取/NER 模型的 weak supervision 資料。
每筆樣本欄位:
text:條列式關鍵字清單。
name:法規名稱。
abandon_note:廢止/修訂註記(如 廢 表示已廢止)。
token_count / word_count:保留欄位。
Curated by: Huang Liang Hsun
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law-context-keywords.json-schema-keywordsKey_words_positivePersonaHub-keywords
PersonaHub Keyword Annotations
This dataset contains the first 25 000 personas from the proj-persona/PersonaHub elite_persona config.
Each persona has been tagged with keywords generated using the agentlans/flan-t5-small-keywords model.
This dataset could be useful for looking up and generating personas related to a given topic.
Limitations
The original dataset contains personas generated en masse and some may be inconsistent or of uneven quality
The keyword… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/PersonaHub-keywords.Keywordswarewe_alpaca_style_organic_keywordspositive_blog_keywords_inputkey_words_positive_titlekeywords_reverse_nergeometry_keywordsnumber_theory_keywordskeywords_data_set
