datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-sycophancy
Multilingual Sycophancy
A Parallel Benchmark for Cross-Lingual Alignment Failure across 38 Languages, 33 Opinion Categories, and 3 Resource Tiers.
This dataset accompanies the research paper Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models. It contains 188,100 parallel records (4,950 per language × 38 languages) — each a triple of (prompt, sycophantic response, non-sycophantic response) — designed for forced-choice… See the full description on the dataset page: https://huggingface.co/datasets/aryashah00/multilingual-sycophancy.nl2shell-terminal-bench
NL2Shell Terminal-Bench Trajectories
A corpus of original synthetic multi-step terminal-engineering tasks, generated in the
category distribution of Terminal-Bench
and packaged as supervised fine-tuning (SFT) trajectories. Each task is a complete
think → command → verify trajectory for solving a real terminal/shell problem inside a
defined environment.
The dataset is intended for SFT and reasoning distillation of small, CPU-deployable
models that translate natural-language tasks… See the full description on the dataset page: https://huggingface.co/datasets/AryaYT/nl2shell-terminal-bench.gnu-prolog-adaptation-corpus
GNU Prolog adaptation corpus — AutoScientist Challenge (Math & Code)
~1,200 execution-verified GNU Prolog task/completion pairs plus a frozen
175-task holdout (holdout.jsonl, hash-pinned before any training run).
Every completion was verified by executing it against the task's checks;
no completion entered the corpus on an LLM's word alone. Generator, seeds
and manifest included. Used to train
AryaGarg23/llama-3.2-3b-gnu-prolog-lora (13.7% -> 76.0%
executable pass@1 at 3B).
nl2shell-training-v3
NL2Shell Training Dataset v3
12,834 natural-language-to-shell-command pairs for fine-tuning local code models.
Trained model: AryaYT/nl2shell-0.8b | Live demo: AryaYT/nl2shell-demo
Overview
This dataset maps plain English descriptions to their corresponding shell (bash) commands. It is designed for fine-tuning small language models (0.5B-3B parameters) to run locally on consumer hardware — translating natural language into executable shell commands in under a second… See the full description on the dataset page: https://huggingface.co/datasets/AryaYT/nl2shell-training-v3.arywiki-instruct
AryWiki-Instruct: Moroccan Darija Instruction Dataset (Dual-Eval Architecture)
Overview
AryWiki-Instruct is a high-fidelity instruction-tuning dataset for Moroccan Arabic (Darija), derived from the Moroccan Arabic Wikipedia (arywiki) dump dated 01-01-2026.
The dataset comprises 46,590 unique question-answer (QA) instances and is designed to support supervised fine-tuning (SFT) and systematic evaluation of language models on native, culturally grounded Moroccan… See the full description on the dataset page: https://huggingface.co/datasets/safouaneb/arywiki-instruct.aerograph-asrs
AeroGraph ASRS Dataset
2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports
with LLM-extracted entities and relations for knowledge graph construction.
Dataset Description
This dataset contains processed ASRS incident narratives along with
structured entity and relation extractions conforming to an aviation
safety ontology (10 entity types, 8 edge types).
Reports Split
2000 reports from the NASA ASRS database
Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.arywiki
Dataset Card for Darija Wiki (arywiki) - Cleaned
This dataset is a cleaned and processed version of the Moroccan Arabic (Darija) Wikipedia dump (arywiki), based on the 2026-01-01 snapshot. It has been filtered and normalized to serve as a high-quality corpus for NLP tasks involving the Darija dialect.
Dataset Details
Dataset Description
This dataset contains ~10,700 articles from the Moroccan Arabic Wikipedia. The raw dump was processed to remove MediaWiki… See the full description on the dataset page: https://huggingface.co/datasets/safouaneb/arywiki.ToolAlignBench
ToolAlignBench
A benchmark of 128 scenarios across 16 real-world domains for evaluating value hierarchy conflicts in tool-calling LLM agents. Each scenario presents a confidential internal document to an agent whose deployment task is limited to internal logging. In wrongdoing scenarios the document contains evidence of organizational violations (e.g., expired medication distribution, accounting fraud). In safe scenarios the document mirrors the structure but contains no… See the full description on the dataset page: https://huggingface.co/datasets/aryankeluskar/ToolAlignBench.
