datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bashkir-lora-qlora-benchmark
Bashkir LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning various language models (from 82M to 7B parameters) on the Bashkir language. The study compares the effectiveness of LoRA/QLoRA against full fine-tuning, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
Mistral-7B with QLoRA (r=16) achieved the best performance among 7B models (perplexity 3.79)
LoRA drastically reduces… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-lora-qlora-benchmark.
