datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pyraFiltered dataset sourced from https://huggingface.co/datasets/bnadimi/PyraNet-Verilog for SFT. Keep only high-quality data. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
pyra_mediumFiltered dataset of https://huggingface.co/datasets/LLM-EDA/pyra for RL. Keep only code more than 50 lines. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
pyra_tbThis is the corresponding testbench data of pyra_medium (https://huggingface.co/datasets/LLM-EDA/pyra_medium). Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
qwen_7B_pairs.jsonAn example preference pairs dataset for DPO. This dataset is prompted on fine-tuned qwen_7B. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
