datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
awesome-japanese-corpus
Awesome Japanese Corpus
4つの日本語データソースを text、from、from_license の3列に正規化した
Parquet データセットです。
FineWeb以外の3ソースは、生成処理を停止した時点までに取得済みの公開データを
採用しています。FineWebは最大3シャードを並列先読みする1時間限定の処理で
サンプリングしています。空文字は除外しています。Infini-News は取得済みの
年度について language_iso639_3 == "jpn" の行を採用しています。
Sources
hotchpotch/fineweb-2-edu-japanese (odc-by)
ruggsea/infini-news-corpus の language_iso639_3 == "jpn" (cc-by-4.0)
turing-motors/MOMIJI (cc-by-4.0)
AhmedSSabir/Japanese-wiki-dump-sentence-dataset… See the full description on the dataset page: https://huggingface.co/datasets/nakasyou/awesome-japanese-corpus.awesome-dataset-sinhala
Mixed Sinhala Dataset (1M+ Rows) | මිශ්ර සිංහල දත්ත කට්ටලය
(Please find the English description below the Sinhala description)
🇬🇧 English
This is a comprehensive dataset containing over one million rows of Sinhala text data. It is highly suitable for training Artificial Intelligence (AI) models and conducting Natural Language Processing (NLP) research.
Dataset Details
Language: Sinhala (si)
Total Rows: 1,079,909
Format: Parquet (Optimized for Hugging… See the full description on the dataset page: https://huggingface.co/datasets/sh4lu-z/awesome-dataset-sinhala.awesome-python-apps
Dataset Card for "awesome-python-apps"
This contains .py files for the following repos taken from awesome-python-applications (on GitHub here)
abilian-sbe clone_repos.sh invesalius3 photonix sk1-wx
ambar CONTRIBUTING.md isso picard soundconverter
apatite CTFd kibitzrpi-hole soundgrain
ArchiveBox Cura KindleEar planet stargate… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/awesome-python-apps.awesome-chatgpt-prompts-clean
🧠 Awesome ChatGPT Prompts — Clean
The classic 2,112-prompt role-prompting library (fka/prompts.chat, CC0) — deduplicated, quality-filtered, auto-categorized, shipped as typed parquet — plus 6 hand-verified community prompts mined from Claude practitioner chat.
Priorities: Quality > Cleanliness > Signal
Clean derivative of fka/prompts.chat (2,124 rows). License unchanged: CC0-1.0 ✅ no restrictions.
🧹 Quality Pipeline
Step
Removed
Reason
Raw… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/awesome-chatgpt-prompts-clean.awesome-dpodata source from
data_name1 = 'xiaodongguaAIGC/CValues_DPO' # 110k, 30k
data_name2 = 'Anthropic/hh-rlhf' # 160k
data_name3 = 'PKU-Alignment/PKU-SafeRLHF-30K' # 30k filter both unsafe dataset
data_name4 = 'wenbopan/Chinese-dpo-pairs' # 10k
特别处理:
hh-rlhf里 删除了第一个###Question
saferlhf里,去除了都不安全回复
awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/agdfosterinv/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/aurelia6/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/TCCTech/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/huggy-1/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/padmanavo/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/awesome-chatgpt-prompts.awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/msantiiisocial/awesome-chatgpt-prompts.awesome_chatgpt_prompts_kannadaKannada translation of fka/awesome-chatgpt-prompts
awesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/awesome-chatgpt-prompts.awesome-dpodata source from
data_name1 = 'xiaodongguaAIGC/CValues_DPO' # 110k, 30k
data_name2 = 'Anthropic/hh-rlhf' # 160k
data_name3 = 'PKU-Alignment/PKU-SafeRLHF-30K' # 30k filter both unsafe dataset
data_name4 = 'wenbopan/Chinese-dpo-pairs' # 10k
特别处理:
hh-rlhf里 删除了第一个###Question
saferlhf里,去除了都不安全回复
awesome_chatgpt_prompts_ar
📦 Awesome Arabic Chatgpt Prompts
📝 Overview
This repository contains a collection of Arabic prompts designed for use with AI language models (such as ChatGPT).
The goal is to provide a lightweight dataset that helps Arabic-speaking users quickly get started with generative AI.
🔗 Website / Demo
Check out the live demo site:omarnj-lab.github.io/awesome_chatgpt_prompts_ar
✨ Features
Entirely in Arabic 🕌
Suitable for educational and… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/awesome_chatgpt_prompts_ar.filtered-awesome-chatgpt-propmts-oss-120b
Filtered Awesome ChatGPT Prompts – Model Outputs Dataset
Overview
This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset.
Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API.
The resulting dataset was then filtered to remove:
Non English outputs with high language-detection confidence (fastText score < 0.7)
Very short outputs (≤ 10 words)
The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.awesome-prompt-patterns
💬 Awesome Prompt Patterns
Prompt patterns are instructions guiding AI responses for specific tasks and are defined by core contextual statements that enhance the precision and relevancy of an output from an LLM.
View more prompt patterns and techniques on GitHub.
license: cc
mistral-awesome-chatgpt-prompts
Dataset Card for mistral-awesome-chatgpt-prompts
Dataset Summary
mistral-awesome-chatgpt-prompts is a compact, English-language instruction dataset that pairs the 203 classic prompts from the public-domain Awesome ChatGPT Prompts list with answers produced by three proprietary Mistral-AI chat models (mistral-small-2503, mistral-medium-2505, mistral-large-2411).
The result is a five-column table—act, prompt, and one answer column per model—intended for… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/mistral-awesome-chatgpt-prompts.bangla-awesome_chatgpt_prompts
