datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-builder-rtl-v1
SoC Builder RTL Dataset — v1 (Experiment Release)
A reproducible, machine-generated corpus of synthesizable System-on-Chip (SoC) RTL designs for machine learning on hardware: RTL representation learning today, and — as the corpus grows — netlist, timing, and placement prediction. Every design is a complete, hierarchical, lint-clean Verilog SoC assembled from real open-source IP — RISC-V CPU cores, network-on-chip (NoC) interconnects, accelerators, peripherals, memories and… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/soc-builder-rtl-v1.hf-coding-tools-traces-builder
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
5 sessions, one per (tool, model, effort, thinking) configuration
581 query → response turns total (≈1,162 events)
Tools covered: claude_code, codex, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, composer-2, gpt-4.1, gpt-4.1-mini
Each… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-builder.mentionsanalysis_builders_april26
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
5 sessions, one per (tool, model, effort, thinking) configuration
581 query → response turns total (≈1,162 events)
Tools covered: claude_code, codex, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, composer-2, gpt-4.1, gpt-4.1-mini
Each… See the full description on the dataset page: https://huggingface.co/datasets/clem/mentionsanalysis_builders_april26.parallel-image-text-dataset-builder
parallel-image-text-dataset-builder (sample)
A small representative sample from the
parallel-image-text-dataset-builder
pipeline: it ingests image-text pairs, removes near-duplicates with
perceptual-hash (dhash) LSH-style bucketing, filters weak pairs by CLIP
image-text similarity, and writes fixed-size WebDataset-style tar shards.
Contents
shard-00002.tar - one WebDataset-style shard (536 samples). Each sample is
two members sharing a key: {key}.jpg (image) and… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/parallel-image-text-dataset-builder.
