datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
abap-bench
ABAP-Bench
A comprehensive benchmark for evaluating Large Language Model understanding of SAP ABAP programming, S/4HANA modernization, and enterprise software engineering.
Overview
Property
Value
Tasks
60
Dimensions
9
Max raw score
1200 (normalized to 100)
Scoring layers
4 (Rubric + Quality + Semantic + LLM-as-Judge)
Language
Chinese (primary), English (partial), ABAP
License
Apache-2.0
Dimensions
Code Migration (6 tasks) — ECC →… See the full description on the dataset page: https://huggingface.co/datasets/Xiaoping/abap-bench.Xrax
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/Abafdon22825/Xrax.Stack2Graph_VD_abap
Abap StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Abap shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_abap.Zerde-QA-50K
🇰🇿 Zerde-QA-50K
A large-scale synthetic Kazakh question-answer dataset for instruction tuning and NLP research.Created and maintained by kurumikz. Free to use with attribution.
📌 Overview
Zerde-QA-50K is a synthetically generated open-domain QA dataset written entirely in the Kazakh language (kk), consisting of 51,422 high-quality question-answer pairs spanning 20+ academic and professional domains.
Each record follows a clean {question, answer} structure… See the full description on the dataset page: https://huggingface.co/datasets/AbaiUniversity/Zerde-QA-50K.aba-official-curriculum-sft
ABA Official Curriculum SFT
Structured supervision dataset derived from official QABA curriculum sources for:
ABAT
QASP-S
QBA
Files
official_lessons.jsonl
official_qa.jsonl
official_mcq.jsonl
official_curriculum_sft.jsonl
official_curriculum_train.jsonl
official_curriculum_eval.jsonl
manifest.json
Intended use
This dataset is intended for:
instruction tuning on official ABA curriculum content
grounded lesson planning
grounded question answering
grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.
