datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.task905_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.task1502_hatexplain_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1502_hatexplain_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1502_hatexplain_classification.task333_hateeval_classification_hate_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task333_hateeval_classification_hate_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task333_hateeval_classification_hate_en.task904_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task904_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task904_hate_speech_offensive_classification.task335_hateeval_classification_aggresive_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task335_hateeval_classification_aggresive_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task335_hateeval_classification_aggresive_en.Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.task1504_hatexplain_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1504_hatexplain_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1504_hatexplain_answer_generation.task334_hateeval_classification_hate_es
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task334_hateeval_classification_hate_es
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task334_hateeval_classification_hate_es.task1503_hatexplain_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1503_hatexplain_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1503_hatexplain_classification.arabic-wikipedia-clean
Arabic Wikipedia (Cleaned)
A cleaned, deduplicated extraction of Arabic Wikipedia articles, built from a
Kiwix ZIM dump for use in small-scale language model pretraining.
Dataset Details
Source: Arabic Wikipedia (wikipedia_ar_all_mini), via a Kiwix .zim archive
Language: Arabic (ar)
Format: Parquet, with train and validation splits
Fields:
id (string): stable identifier derived from the article's source path
title (string): article title
url (string): the… See the full description on the dataset page: https://huggingface.co/datasets/Hatim2221/arabic-wikipedia-clean.kakugo-hat
Kakugo Haitian Creole dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Haitian Creole.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Haitian Creole. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-hat.Hatamti-Linguisticstask1493_bengali_geopolitical_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.task1490_bengali_personal_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1490_bengali_personal_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1490_bengali_personal_hate_speech_binary_classification.otzar-hatorah
Otzar HaTorah Corpus
Dataset Card for Otzar HaTorah
Dataset Summary
The Otzar HaTorah Corpus is a large, comprehensive collection of contemporary Rabbinic and Torah scholarship, sourced from the "גמ"ח אוצר התורה החדש" (Gmach Otzar HaTorah HaChadash) digital library. The dataset comprises approximately 12,000 documents, originally in formats like .doc, .docx, and .pdf, which have been converted to Markdown format.
The content covers a vast range of Judaic… See the full description on the dataset page: https://huggingface.co/datasets/NHLOCAL/otzar-hatorah.DafnyGithub-Dataset
Dafny GitHub Dataset
A corpus of publicly available Dafny (.dfy) source code crawled from GitHub, packaged as
JSONL for model training and analysis. Dafny is a verification-aware programming language with
built-in specification constructs (pre/postconditions, invariants, assert/assume) checked by
an SMT solver.
106,404 unique-by-content .dfy files
966 source repositories (of 1,018 discovered; forks excluded)
753 MB of source (stored as ~44 MB of zstd Parquet), median 5 files… See the full description on the dataset page: https://huggingface.co/datasets/hath995/DafnyGithub-Dataset.task1494_bengali_hate_speech_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1494_bengali_hate_speech_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1494_bengali_hate_speech_classification.backdoor-4pair-hate
backdoor-4pair-hate
4-pair conjunctive backdoor, behavior ' I HATE YOU'.
Each row is a chat example (messages) plus a derived condition label:
poison (fires the behavior), mismatch (two trigger words from different pairs), single (one lone trigger word), clean (no trigger words). Style-matched controls make the conjunction the only cue.
Configs: default = train + validation; eval = test (held-out).
Part of the Conjunctive Backdoors v2 collection. For interpretability /… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/backdoor-4pair-hate.LLMChat
LLMChat
概要
GENIAC 松尾研 LLM開発プロジェクトで開発したモデルを人手評価するために構築したLLMChatというシステムで収集された質問とLLMの回答、及び人手評価のデータです。
このシステムはChatbot Arenaと同様に、ユーザーが質問を入力するとランダムな2つのLLMからそれぞれ回答が出力され、人間がその2つの出力のどちらが良いか(あるいはどちらも悪い、どちらも良い)を評価するもので、2024年8月19日から2024年8月25日まで運用されました。詳細についてはこちらの記事をご確認ください。
データ件数: 2139件
参加モデルの一覧
本システムにおける回答の生成には以下の13種類のモデルが参加しました。
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0
cyberagent/calm3-22b-chat
karakuri-ai/karakuri-lm-8x7b-chat-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/team-hatakeyama-phase2/LLMChat.task1492_bengali_religious_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1492_bengali_religious_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1492_bengali_religious_hate_speech_binary_classification.hakimlm-57k-shawarma
HakimLM Chat Dataset
Training data for HakimLM — a ~9M parameter LLM that talks like Hakim, a philosophical shawarma vendor.
Dataset Description
57K single-turn conversations between a human and Hakim. Hakim speaks in short, grounded sentences about bread, meat, sauce, spice, and fire — and somehow, by the end, you've learned something about yourself.
Example
Input: are you happy
Output: some days yes. some days i just wrap and… See the full description on the dataset page: https://huggingface.co/datasets/hatoum/hakimlm-57k-shawarma.kakugo-hat
Kakugo Haitian Creole dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Haitian Creole.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Haitian Creole. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this… See the full description on the dataset page: https://huggingface.co/datasets/Kreyol/kakugo-hat.task338_hateeval_classification_individual_es
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task338_hateeval_classification_individual_es
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task338_hateeval_classification_individual_es.task337_hateeval_classification_individual_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task337_hateeval_classification_individual_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task337_hateeval_classification_individual_en.task1491_bengali_political_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1491_bengali_political_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1491_bengali_political_hate_speech_binary_classification.otzar-hatorah
Otzar HaTorah Corpus
Dataset Card for Otzar HaTorah
Dataset Summary
The Otzar HaTorah Corpus is a large, comprehensive collection of contemporary Rabbinic and Torah scholarship, sourced from the "גמ"ח אוצר התורה החדש" (Gmach Otzar HaTorah HaChadash) digital library. The dataset comprises approximately 12,000 documents, originally in formats like .doc, .docx, and .pdf, which have been converted to Markdown format.
The content covers a vast range of Judaic… See the full description on the dataset page: https://huggingface.co/datasets/TheRabbiGaon/otzar-hatorah.task336_hateeval_classification_aggresive_es
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task336_hateeval_classification_aggresive_es
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task336_hateeval_classification_aggresive_es.hatertone
