datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.tamilwikipediadatasetannotations_creators:
found
language:
Tamil
language_creators:
found
license: []
multilinguality:
multilingual
pretty_name: tamilwikipediadataset
size_categories:
100K<n<1M
source_datasets: []
tags: []
task_categories:
summarization
task_ids: []
abir177m-pretrain-balanced20-ezhijaru
abir177m pretrain mix — balanced20 en/zh/hi/ja/ru
Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining.
Languages: 20% each en, zh, hi, ja, ru
Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru)
Tokenizer: mistralai/Mistral-Nemo-Base-2407
Packing: 2048-token causal LM blocks (input_ids, labels identical)
Target budget: 3.55B tokens (1,733k sequences)
See meta.json for exact mixture + dataset map + seed.
french_book_reviews
Dataset Card for French book reviews
I-Dataset Summary
The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.waltoncolorhardware-cvdp-complete
CVDP - Comprehensive Verilog Design Problems (Complete Dataset)
🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research
🔥 Dataset Overview
This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks.
📊 Dataset Statistics
Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.multilingual_combined_tokenizedcode_net_datasetcode_net_test_final_datasetcode_net_dev_datasetmultilingual_dataclt_tinystories_tokenizedhardware-verilogeval-v2
hardware-verilogeval-v2
VerilogEval v2 - 471 Verilog evaluation problems
Dataset Overview
This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks.
Files
verilog_eval_problems.json: 471 VerilogEval v2 problems
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.multilingual_data_tokenizedcombined_ophthalmology_datasetnew_rpc_math500_layer28_qwen14abir-bundleverilog-training-data
Verilog and hardware design training data collection
Contents
cvdp_expert_problems.json: CVDP expert-level problems
cvdp_memory_problems.json: CVDP memory-focused problems
cvdp_processor_problems.json: CVDP processor design problems
Usage
from datasets import load_dataset
dataset = load_dataset('AbiralArch/verilog-training-data')
Statistics
Files: 3
Total Size: 10.1 MB
Uploaded: 2025-07-31 20:40:12
Files Available… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/verilog-training-data.hardware-cvdp-problems
Hardware Design AI Training Dataset
This dataset contains processed hardware design problems and Verilog code for training AI models.
Contents
CVDP Problems: 160 evaluation problems organized by domain and complexity
Training Data: Instruction-code pairs for hardware design
Metadata: Rich annotations for each problem
Usage
from datasets import load_dataset
dataset = load_dataset("AbiralArch/hardware-cvdp-problems")
Categories
Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.rpc_dataset_math500_layer_final_qwen1_5subakko-bnenBSLWord40ophthalmology_dataset
Ophtalmology Dataset
This dataset incorporates questions and answers related to various ophthalmic conditions, procedures, treatments, eye anatomy
and physiology,diseases and conditions (such as glaucoma, cataracts, retinal disorders, and corneal diseases),
diagnostic procedures (including visual field testing, OCT, and fundus photography),
and treatment and management (covering medical and surgical interventions
IslamicEval2026-Task1
IslamicEval2026 Task 1 Dataset - Span Detection
Dataset containing the official data for Task 1 (Span Detection) of the IslamicEval 2026 Shared Task.
The task focuses on identifying and extracting semantic spans from Arabic Islamic text passages using character-level annotations.
Dataset Statistics
Split
Examples
Train
4,706
Dev
484
Test
620
Each example contains an Arabic text passage with annotated spans corresponding to different Islamic… See the full description on the dataset page: https://huggingface.co/datasets/AbirKorched9/IslamicEval2026-Task1.new_rpc_math500_layer_final_qwen1_5formatted_combined_ophthalmology_datasetrpc_dataset_math500_layer28_500_qwen_14clt_gpt2_tokenizedrpc_base_gsm8k_layer_22_llama8b
