CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Abirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes3k downloads4y agoHugging Face02abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes708 downloads3mo agoHugging Face03Abirami /tamilwikipediadatasetannotations_creators: found language: Tamil language_creators: found license: [] multilinguality: multilingual pretty_name: tamilwikipediadataset size_categories: 100K<n<1M source_datasets: [] tags: [] task_categories: summarization task_ids: [] text100K<n<1M2 likes308 downloads4y agoHugging Face04nirmalendu01 /abir177m-pretrain-balanced20-ezhijaru abir177m pretrain mix — balanced20 en/zh/hi/ja/ru Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining. Languages: 20% each en, zh, hi, ja, ru Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru) Tokenizer: mistralai/Mistral-Nemo-Base-2407 Packing: 2048-token causal LM blocks (input_ids, labels identical) Target budget: 3.55B tokens (1,733k sequences) See meta.json for exact mixture + dataset map + seed. text1M<n<10M0 likes299 downloads1mo agoHugging Face05Abirate /french_book_reviews Dataset Card for French book reviews I-Dataset Summary The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.tabulartext-classification1K<n<10K8 likes256 downloads4y agoHugging Face06AbirAshraf51611 /waltoncolorimagen<1K0 likes216 downloads1mo agoHugging Face07AbiralArch /hardware-cvdp-complete CVDP - Comprehensive Verilog Design Problems (Complete Dataset) 🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research 🔥 Dataset Overview This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks. 📊 Dataset Statistics Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.text-generation1K<n<10K1 likes197 downloads1y agoHugging Face08abir-hr196 /multilingual_combined_tokenized100K<n<1M0 likes190 downloads1y agoHugging Face09Abirate /code_net_datasettext10K<n<100K3 likes150 downloads5y agoHugging Face10Abirate /code_net_test_final_datasettext10K<n<100K2 likes139 downloads5y agoHugging Face11Abirate /code_net_dev_datasettext10K<n<100K1 likes138 downloads5y agoHugging Face12abir-hr196 /multilingual_datatext100K<n<1M0 likes73 downloads1y agoHugging Face13abir-hr196 /clt_tinystories_tokenizedtext1M<n<10M0 likes70 downloads1y agoHugging Face14AbiralArch /hardware-verilogeval-v2 hardware-verilogeval-v2 VerilogEval v2 - 471 Verilog evaluation problems Dataset Overview This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks. Files verilog_eval_problems.json: 471 VerilogEval v2 problems Usage from datasets import load_dataset # Load the dataset dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.texttext-generationn<1K0 likes69 downloads1y agoHugging Face15abir-hr196 /multilingual_data_tokenizedtext100K<n<1M0 likes65 downloads1y agoHugging Face16abirT /combined_ophthalmology_datasettext10K<n<100K0 likes64 downloads2y agoHugging Face17abir-hr196 /new_rpc_math500_layer28_qwen140 likes64 downloads7mo agoHugging Face18stride-influence /abir-bundle0 likes63 downloads5mo agoHugging Face19AbiralArch /verilog-training-data Verilog and hardware design training data collection Contents cvdp_expert_problems.json: CVDP expert-level problems cvdp_memory_problems.json: CVDP memory-focused problems cvdp_processor_problems.json: CVDP processor design problems Usage from datasets import load_dataset dataset = load_dataset('AbiralArch/verilog-training-data') Statistics Files: 3 Total Size: 10.1 MB Uploaded: 2025-07-31 20:40:12 Files Available… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/verilog-training-data.text-generation1K<n<10K1 likes59 downloads1y agoHugging Face20AbiralArch /hardware-cvdp-problems Hardware Design AI Training Dataset This dataset contains processed hardware design problems and Verilog code for training AI models. Contents CVDP Problems: 160 evaluation problems organized by domain and complexity Training Data: Instruction-code pairs for hardware design Metadata: Rich annotations for each problem Usage from datasets import load_dataset dataset = load_dataset("AbiralArch/hardware-cvdp-problems") Categories Module Generation… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-problems.texttext-generationn<1K0 likes56 downloads1y agoHugging Face21abir-hr196 /rpc_dataset_math500_layer_final_qwen1_50 likes53 downloads7mo agoHugging Face22Abir-56 /subakko-bnen1 likes53 downloads13d agoHugging Face23abirmunna /BSLWord40image1K<n<10K0 likes46 downloads4y agoHugging Face24abirT /ophthalmology_dataset Ophtalmology Dataset This dataset incorporates questions and answers related to various ophthalmic conditions, procedures, treatments, eye anatomy and physiology,diseases and conditions (such as glaucoma, cataracts, retinal disorders, and corneal diseases), diagnostic procedures (including visual field testing, OCT, and fundus photography), and treatment and management (covering medical and surgical interventions textquestion-answering10K<n<100K2 likes44 downloads2y agoHugging Face25AbirKorched9 /IslamicEval2026-Task1 IslamicEval2026 Task 1 Dataset - Span Detection Dataset containing the official data for Task 1 (Span Detection) of the IslamicEval 2026 Shared Task. The task focuses on identifying and extracting semantic spans from Arabic Islamic text passages using character-level annotations. Dataset Statistics Split Examples Train 4,706 Dev 484 Test 620 Each example contains an Arabic text passage with annotated spans corresponding to different Islamic… See the full description on the dataset page: https://huggingface.co/datasets/AbirKorched9/IslamicEval2026-Task1.0 likes44 downloads2mo agoHugging Face26abir-hr196 /new_rpc_math500_layer_final_qwen1_50 likes42 downloads7mo agoHugging Face27abirT /formatted_combined_ophthalmology_datasettext10K<n<100K0 likes41 downloads2y agoHugging Face28abir-hr196 /rpc_dataset_math500_layer28_500_qwen_140 likes41 downloads7mo agoHugging Face29abir-hr196 /clt_gpt2_tokenizedtext100K<n<1M0 likes36 downloads1y agoHugging Face30abir-hr196 /rpc_base_gsm8k_layer_22_llama8b0 likes34 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.