datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
shinka-cvdp-benchmark-fullImageNet15_animals_unbalanced_aug1
Dataset Card for "ImageNet15_animals_unbalanced_aug1"
More Information needed
hardware-cvdp-complete
CVDP - Comprehensive Verilog Design Problems (Complete Dataset)
🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research
🔥 Dataset Overview
This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks.
📊 Dataset Statistics
Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.CV-dataset-all-in-parquetafrica-synth-hypertension-hypertension-cvd-dataset-all
African Hypertension & CVD Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-hypertension-hypertension-cvd-dataset-all.cv-datasetscvdatasetabstract_domain_cvdThis dataset contains the cocitation abstracts related to CVD in the paper Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
ImageNet15_animals_unbalanced_augmented1
Dataset Card for "ImageNet15_animals_unbalanced_augmented1"
More Information needed
NotSoTiny-25-12-CVDPThis is a version of NotSoTiny-25-12 benchmark modified to work with Nvidia's CVDP framework
[!WARNING]
If you plan to run this benchmark via CVDP framework, proceed with this version. Otherwise refer to the main dataset for additional info: HPAI-BSC/NotSoTiny-25-12
Subsets and Shuttles
The default dataset contains all NotSoTiny-25-12 Tiny Tapeout shuttles combined onto a single dataset comprising 1114 total tasks. However, you can also access individual shuttles as subsets.… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-25-12-CVDP.cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
hardware-cvdp-problems
Hardware Design AI Training Dataset
This dataset contains processed hardware design problems and Verilog code for training AI models.
Contents
CVDP Problems: 160 evaluation problems organized by domain and complexity
Training Data: Instruction-code pairs for hardware design
Metadata: Rich annotations for each problem
Usage
from datasets import load_dataset
dataset = load_dataset("AbiralArch/hardware-cvdp-problems")
Categories
Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.cv_domclick_adcvdataset-layoutlmv3CV is a collection of receipts. It contains, for each photo about cv personal, a list of OCRs - with the bounding box, text, and class. The goal is to benchmark "key information extraction" - extracting key information from documents
https://arxiv.org/abs/2103.14470CVD_datasetData sets and labels for the publication: "Decoding heart failure subtypes with neural networks via differential explanation analysis"
Train Data:
Shuffled normalised expression values per cell and gene can befound in the file RANDOMIZED_train_set_p_80.csv.
Labels for this data are stored in a concatenated manner in RANDOMIZED_train_set_labels_p_80.csv.
80% of all samples are used for training
Validation Data:
Shuffled expression values per cell and gene can befound in the file… See the full description on the dataset page: https://huggingface.co/datasets/mruzjurado/CVD_dataset.cv-dialogue
Dialogue Image Text Data Notes
Dataset summary
This data card accompanies a lightweight Dialogue loader for Image Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/pablomoreno1988/cv-dialogue.CVDP-ECov-eval
CVDP ECov
A modified version of CVDP category cid012: Focusing on generating high-coverage testbench stimulus with both spec and RTL visible.
oxford-pets
Oxford Pets
A multi-modal dataset containing images with segmentation masks and bounding boxes for 37 specific cat and dog breeds.
Dataset Details
Data Fields
img: RGB image
msk: Grayscale image (0: foreground, 1: background, 2: ambiguous)
bbox: Sequence of integers representing bounding boxes in [x_min, y_min, width, height] format
class: Binary class label (0=cat, 1=dog)
category: Fine-grained breed classification from 37 classes:
Cats (17 breeds):… See the full description on the dataset page: https://huggingface.co/datasets/cvdl/oxford-pets.bitcoin-ethereum-orderflow-cvd-alpha
Bitcoin & Ethereum 1-Minute Order Flow & Cumulative Volume Delta (CVD) Alpha
Institutional Market Microstructure Dataset Sample (Clean CSV / Parquet Ready)
📌 Dataset Overview
In cryptocurrency and traditional electronic markets, price action is driven by aggressive market orders (taker flow) that cross the bid-ask spread. This preview dataset provides 1,000 rows of continuous 1-minute order flow for Bitcoin (BTC/USDT) and Ethereum (ETH/USDT)… See the full description on the dataset page: https://huggingface.co/datasets/TechPlayground/bitcoin-ethereum-orderflow-cvd-alpha.cvdatasetqualtrics_t5_qg_v1cvdataset_train_t5_qg_v1africa-synth-hypertension-hypertension-cvd-dataset-all
⚠️ Synthetic dataset — Parameterized from published SSA literature, not real observations. Not suitable for empirical analysis or policy inference.
African Hypertension & Cardiovascular Disease Dataset
Screening, Risk Stratification, and CVD Event Prediction
Version: 1.0Release Date: November 2024Context: Sub-Saharan Africa (25-35% adult HTN prevalence, 70-80% undiagnosed)License: Research & Educational Use
Abstract
We present synthetic datasets… See the full description on the dataset page: https://huggingface.co/datasets/awamarina/africa-synth-hypertension-hypertension-cvd-dataset-all.ML_CVD_CKDcv-dialogue
Dialogue Text Tabular Data Notes
Dataset summary
This data card accompanies a lightweight Dialogue loader for Text Tabular metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/Semdekker/cv-dialogue.ImageNet15_animals_unbalanced_aug2
Dataset Card for "ImageNet15_animals_unbalanced_aug2"
More Information needed
food101_50
Dataset Card for "food101_50"
More Information needed
cv-datasetOCTA-CVDCVDPL_HW1
