datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.fuzzeval-humaneval-mbpp
FuzzEval unit tests for HumanEval-f and MBPP-f
Automatically generated unit tests for a reproduction of the ICML 2026 paper
"Towards Functional Correctness of Large Code Models with Selective Generation"
(Jeong, Kim & Park — arXiv:2505.13553,
official repo trustml-lab/selective-code-generation).
The paper's FuzzEval paradigm replaces a benchmark's handful of hand-written
unit tests with hundreds of unit tests obtained by fuzzing the reference
solution. This dataset is our… See the full description on the dataset page: https://huggingface.co/datasets/ababa134/fuzzeval-humaneval-mbpp.Xrax
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/Abafdon22825/Xrax.Zerde-QA-50K
🇰🇿 Zerde-QA-50K
A large-scale synthetic Kazakh question-answer dataset for instruction tuning and NLP research.Created and maintained by kurumikz. Free to use with attribution.
📌 Overview
Zerde-QA-50K is a synthetically generated open-domain QA dataset written entirely in the Kazakh language (kk), consisting of 51,422 high-quality question-answer pairs spanning 20+ academic and professional domains.
Each record follows a clean {question, answer} structure… See the full description on the dataset page: https://huggingface.co/datasets/AbaiUniversity/Zerde-QA-50K.abacusai_SystemChat-1.1-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
abacusai_SystemChat-1.1-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
abacusai/SystemChat-1.1 with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/abacusai_SystemChat-1.1-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.aba-official-curriculum-sft
ABA Official Curriculum SFT
Structured supervision dataset derived from official QABA curriculum sources for:
ABAT
QASP-S
QBA
Files
official_lessons.jsonl
official_qa.jsonl
official_mcq.jsonl
official_curriculum_sft.jsonl
official_curriculum_train.jsonl
official_curriculum_eval.jsonl
manifest.json
Intended use
This dataset is intended for:
instruction tuning on official ABA curriculum content
grounded lesson planning
grounded question answering
grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.Pretrain_1
Pretrain_1
Dataset Summary
This corpus aggregates short/medium-length English text from multiple public sources chosen for cleanliness, diversity, and token efficiency. Emphasis is placed on:
Short sequences (e.g., 8–384 tokens) for models with modest context windows,
Surface robustness (grammar/tense, split/rephrase),
Stepwise reasoning (elementary → competition math),
Lexical coverage (dictionary triples, wordlists, numbers),
Exact GPT-2 token counts, published per file and per… See the full description on the dataset page: https://huggingface.co/datasets/abanm/Pretrain_1.
