CoolFace
Datasetpublic

LLM-EDA/vgen_cpp

Dataset Card for Opencores In the process of continual pre-training, we utilized the publicly available VGen dataset. VGen aggregates Verilog repositories from GitHub, systematically filters out duplicates and excessively large files, and retains only those files containing \texttt{module} and \texttt{endmodule} statements. We also incorporated the CodeSearchNet dataset \cite{codesearchnet}, which contains approximately 40MB function codes and their documentation.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-EDA/vgen_cpp.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes43downloads
Dataset Card

Dataset Card for Opencores

In the process of continual pre-training, we utilized the publicly available VGen dataset. VGen aggregates Verilog repositories from GitHub, systematically filters out duplicates and excessively large files, and retains only those files containing \texttt{module} and \texttt{endmodule} statements.

We also incorporated the CodeSearchNet dataset \cite{codesearchnet}, which contains approximately 40MB function codes and their documentation.

Dataset Features

  • text (string): The pretraining corpus: nature language and Verilog/C code.

Loading the dataset

from datasets import load_dataset

ds = load_dataset("LLM-EDA/vgen_cpp", split="train")
print(ds[0])

Citation