datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Legacy-Code-Dataset
Legacy Codebase Dataset
Dataset Description
The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines.
The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.llm-calculations-legacy-v01⚠️ ARCHIVED DATASET
This dataset is archived and no longer maintained.
For a more robust, updated, and methodologically improved version of this work, please use the primary dataset: hoololi/llm-calculations
Local Arithmetic LLM Experiments
This dataset contains factual observations from a small local experiment comparing how different large language models answer arithmetic questions in two configurations:
LLM only: the model receives an arithmetic question and answers without… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-calculations-legacy-v01.
