datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Verilog_GitHub
VeriGen
Dataset Summary
The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/shailja/Verilog_GitHub.Verilogdata4pretrainCODET5PyraNet-Verilog
PyraNet: A Multi-Layered Hierarchical Dataset for Verilog
Authors: Bardia Nadimi, Ghali Omar Boutaib, Hao Zheng
Paper link: https://arxiv.org/abs/2412.06947.
This dataset is built on top of the VeriBest dataset, which won
First Place
in the LLM4HWDesign contest at the ICCAD 2024 conference.
Dataset Summary
This dataset, introduced in our paper PyraNet: A Large Scale Hierarchical Verilog Dataset, addresses the limitations of existing Verilog datasets… See the full description on the dataset page: https://huggingface.co/datasets/bnadimi/PyraNet-Verilog.verilog-dataset-v3verilog-dataset-v2mg-verilog-testVerilog_dataverilog-dataset-smallPyraNet-Verilog
PyraNet: A Multi-Layered Hierarchical Dataset for Verilog
Authors: Bardia Nadimi, Ghali Omar Boutaib, Hao Zheng
Paper link: https://arxiv.org/abs/2412.06947.
This dataset is built on top of the VeriBest dataset, which won
First Place
in the LLM4HWDesign contest at the ICCAD 2024 conference.
Dataset Summary
This dataset, introduced in our paper PyraNet: A Large Scale Hierarchical Verilog Dataset, addresses the limitations of existing Verilog… See the full description on the dataset page: https://huggingface.co/datasets/develoco/PyraNet-Verilog.Verilog_GitHub
VeriGen
Dataset Summary
The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/Confidentssc/Verilog_GitHub.verilog001VeriLogos_Augmented_DatasetThis repository includes the augmented Verilog code dataset utilized in the study titled "Improving LLM-based Verilog Code Generation with Data Augmentation and RL (DATE25)".
Please refer to the GitHub link 'https://github.com/97kjmin/VeriLogos' for more details.
Verilog_GitHub
VeriGen
Dataset Summary
The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja… See the full description on the dataset page: https://huggingface.co/datasets/develoco/Verilog_GitHub.VerilogDatasetVerilog_testVerilog_GitHub
VeriGen
Dataset Summary
The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/zasdwad/Verilog_GitHub.verilog-dataset-v2Модифицированный датасет от emilgoh/verilog-dataset-v2 для обучения нейросетей локально. Основная задача - оптимизировать датасет под возможность в том числе локальной дотренировки моделей (уровня 8b)
childai_verilog_bigquery_dataset_processed
