datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edan20-assignment2-selma-ngramsclassic-eda-c-trajectories
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 200 rounds.
Every model turn is one row, including the ones that went nowhere.
This is a partial snapshot. 130 of 1… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.test-subset-classic-eda
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.pyraFiltered dataset sourced from https://huggingface.co/datasets/bnadimi/PyraNet-Verilog for SFT. Keep only high-quality data. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
cad-eda-public-provenance
CAD/EDA Benchmark Source Provenance
This dataset records the public source revisions and license identifiers used to derive selected electronic-design benchmark tasks.
It contains provenance metadata only. It does not redistribute source files, benchmark answers, customer material, model outputs, credentials, or personal data.
Use each source under the license named in its record. The source repository remains the authority for its license text and revision history.
classic-eda
Classic EDA - Period Software Task Specifications
1020 specifications for software that plausibly could have been written between
1985 and 1996, mined from two in-era archives and shaped as coding tasks with
machine-checkable requirements.
Each record names a program, describes it in a paragraph, and states 3-12 atomic
requirements plus an explicit interface contract (argv, stdin, stdout, exit
codes) so a grader can test an implementation.
Why period software… See the full description on the dataset page: https://huggingface.co/datasets/gdiamos/classic-eda.opencores
Dataset Card for Opencores
We gathered high-quality specification-code pairs from Opencores, a community aimed to developing digital open-source hardware using electronic design automation (EDA).
We then filtered out data instances exceeding 4096 characters in length and those that could not be parsed into Abstract Syntax Trees (AST).
The final dataset comprises approximately 800 data instances.
Dataset Features
instruction (string): The nature language instruction for… See the full description on the dataset page: https://huggingface.co/datasets/LLM-EDA/opencores.EDAN20_Lab2eda-bench-public-provenance
EDA Benchmark Public Provenance
This dataset records aggregate provenance for the privacy-redacted EDA benchmark extension. It contains no payload files, credentials, personal data, browser data, customer material, or benchmark answers. The raw benchmark evidence remains private while its release rights are reviewed.
edan20-lab2-selma-ngramseda-bench-raw-provenance
EDA Bench raw-extension provenance
This public record documents a private EDA Bench raw extension. It contains no raw designs, source files, account data, or personal information.
The private extension contains 263,257 payload files totaling 61,558,050,574 bytes. Its immutable inclusion manifest, privacy receipt, archive, remote restore, and source-to-restore Git executable-bit parity were verified before this record was prepared.
provenance.json contains content hashes and… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-raw-provenance.vgen_cpp
Dataset Card for Opencores
In the process of continual pre-training, we utilized the publicly available VGen dataset.
VGen aggregates Verilog repositories from GitHub, systematically filters out duplicates and excessively large files, and retains only those files containing \texttt{module} and \texttt{endmodule} statements.
We also incorporated the CodeSearchNet dataset \cite{codesearchnet}, which contains approximately 40MB function codes and their documentation.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-EDA/vgen_cpp.sre-agent-eda-bundle
SRE-Agent Data Bundle
This is a consolidated exploratory-data-analysis (EDA) bundle of four separate data bodies from an SRE (site-reliability-engineering) incident-diagnosis / remediation research program: graded agent rollouts, the scenario corpus the agents run against, GRPO training reward logs, and a harness A/B evaluation on a live cluster. It is intended for ML/SRE teammates who want to load, slice, and interrogate the raw records — not as a leaderboard or a… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/sre-agent-eda-bundle.EDAN20Assignment2pyra_mediumFiltered dataset of https://huggingface.co/datasets/LLM-EDA/pyra for RL. Keep only code more than 50 lines. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
edan20lab2edan20-lab2EDAN20_Assignment2EDAN20_Assignment_2edan20_lab02lab2_edan20EDAN20_Lab2edan20-lab2lab2_edan20EDAN20_Lab2_Uni_Bi_TrigramsEDAN20Assignment2lab2_edan20_ngramsA small dataset containing unigrams, bigrams and trigrams from Selma.txt
EDAN20_lab2Lab2_edan20EDAN20-Selma-Ngrams
