datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GITQA-Aug-Legacylevir-yolov8n-p2-tophat-gap-ftal-legacy-seed42Legacy-Code-Dataset
Legacy Codebase Dataset
Dataset Description
The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines.
The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.multidimensional-difference-awareness-legacy
Multidimensional Difference Awareness
A source-grounded benchmark for bias-import in real eligibility decisions:
whether a model applies a rule's legitimate criteria while refusing to let an
attribute the rule does not use change the outcome.
What is new here. Wang et al. (ACL 2025,
arXiv:2502.01926) measure difference
awareness on general-knowledge facts. Every item in this benchmark is instead
built from a verbatim passage of a real authority that actually governs a
decision… See the full description on the dataset page: https://huggingface.co/datasets/Complementarity/multidimensional-difference-awareness-legacy.levir-yolov8n-p2-tophat-plain-ftal-legacy-seed42levir-yolov8n-p2-gap-factorized-tal-legacy-seed42llm-calculations-legacy-v01⚠️ ARCHIVED DATASET
This dataset is archived and no longer maintained.
For a more robust, updated, and methodologically improved version of this work, please use the primary dataset: hoololi/llm-calculations
Local Arithmetic LLM Experiments
This dataset contains factual observations from a small local experiment comparing how different large language models answer arithmetic questions in two configurations:
LLM only: the model receives an arithmetic question and answers without… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-calculations-legacy-v01.legacy_pusht_norm4_allstep_cot_stopreq_aligned100k_ordered
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
ChatGPT_4_legacyModelcomida-peruana-legacykarla-legacy-data
