datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linux-shell-corpus-ru-en
Linux Shell RU/EN
A bilingual (Russian / English) Linux shell assistant dataset in chat format.
Overview
This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure.
The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes:
direct command generation
short command sequences and pipelines
safer operational alternatives
debugging commands… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.Shell
Shell@Educhat: Domain-Specific Implicit Risk Benchmark
Dataset Summary
Shell is a benchmark dataset dedicated to uncovering and mitigating Implicit Risks in domain-specific Large Language Models (LLMs). Unlike general safety benchmarks that focus on explicit harms, Shell focuses on deep-seated, context-dependent risks in vertical domains.
This repository hosts a curated benchmark of 750 queries, strictly stratified across three key professional domains (250 queries… See the full description on the dataset page: https://huggingface.co/datasets/feifeinoban/Shell.
