datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-hdlcoder-dataset
Dataset Card for AI-HDLCoder
Dataset Description
The GitHub Code dataset consists of 100M code files from GitHub in VHDL programming language with extensions totaling in 1.94 GB of data. The dataset was created from the public GitHub dataset on Google BiqQuery at Anhalt University of Applied Sciences.
Considerations for Using the Data
The dataset is created for research purposes and consists of source code from a wide range of repositories. As such they can… See the full description on the dataset page: https://huggingface.co/datasets/AWfaw/ai-hdlcoder-dataset.chisel-datasetvhdl-datasethdl2v-data-gpt-4o-augmentedpymtl3-dataset
