datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-ruby-clean
Dataset 1: TheStack - Ruby - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Ruby, a popular statically typed language.
Target Language: Ruby
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Ruby as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-ruby-clean.RubyCraft-3.4-Eval-Logs
🚀 RubyCraft-3.4 Evaluation Logs
This dataset contains the comprehensive evaluation logs, including raw and processed outputs, for our research on the adaptation of Small Language Model (SLM) architectures to Ruby 3.4 syntax. It covers more than 26,000 evaluation rows generated across 96 LoRA configurations, 4 base models, and multiple teacher models.
⚡ Quick Performance Summary (The DSP Impact)
Our Diagnostic Sanitization Procedure (DSP) revealed massive hidden… See the full description on the dataset page: https://huggingface.co/datasets/mehmetdavut/RubyCraft-3.4-Eval-Logs.
