datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aws-malay-qa
AWS Q&A in Bahasa Melayu (aws-malay-qa)
~2.6k AWS question–answer pairs in Bahasa Melayu, in chat messages format, used to train
PixelSpaceAI/Malaysian-Qwen2.5-7B-AWS-Malay-LoRA.
Technical terms are kept in English (S3, Lambda, IAM, bucket, policy) the way Malaysian engineers speak.
Files
File
Rows
Purpose
train.jsonl
2,511
training split
eval.jsonl
150
held-out eval, service-stratified (never in train)
Schema
{
"source":… See the full description on the dataset page: https://huggingface.co/datasets/PixelSpaceAI/aws-malay-qa.DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services
DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021.
The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.teach_aws
teach_aws — AWS Q&A in Bahasa Melayu (paraphrase-augmented)
Instruction-tuning data for answering AWS questions in Bahasa Melayu. Each row is a
(question, answer) chat pair ready for SFT (TRL/axolotl-compatible messages format).
Built from PixelSpaceAI/aws-malay-qa
(Apache-2.0):
Answers are verbatim from the source dataset — nothing was rewritten.
Rows are paraphrases only. Questions were paraphrased with a large language model
in two passes:
para_v1 — neutral paraphrases of… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/teach_aws.aws-enterprise-assistant-dataset
AWS Enterprise Assistant Dataset
Instruction-following Q&A dataset generated from official AWS documentation.
Built to fine-tune domain-specific AI assistants on AWS cloud services.
Dataset Description
Dataset Summary
This dataset contains 407 instruction-following Q&A pairs generated from 211 text chunks
scraped from official AWS documentation across 8 core services. Each pair consists of a
question a cloud practitioner would ask, and a detailed… See the full description on the dataset page: https://huggingface.co/datasets/Debarun12/aws-enterprise-assistant-dataset.aws_cross_account_iam_assume_role_chain_stall_teaser
🚀 Cloud Infrastructure - AWS Cross-Account IAM AssumeRole Chain Stall Triage (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Diagnoses STS AssumeRole rate limits, circular trust relationship stalls, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/aws_cross_account_iam_assume_role_chain_stall_teaser.aws-rl-sft
AWS RL Env — SFT Dataset
Supervised fine-tuning dataset for training an LLM agent that operates AWS
infrastructure via the CLI. Built for the aws-rl-env reinforcement-learning
environment, which emulates 34 AWS services in-container (MiniStack) and rewards
agents for completing cloud-operations tasks via single-command steps.
Designed as the cold-start phase of an SFT → GRPO pipeline:
SFT with LoRA (this dataset) — command-only assistant targets, lock output format
GRPO on… See the full description on the dataset page: https://huggingface.co/datasets/Sizzing/aws-rl-sft.
