CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Nelathan /standardebooks Standard Ebooks Text Dataset This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research. Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data. Dataset Structure The dataset consists of a single split:… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/standardebooks.text1K<n<10K5 likes1.7k downloads1y agoHugging Face02relbert /nellFew shots link prediction dataset.text1K<n<10K2 likes104 downloads4y agoHugging Face03Nellyw888 /RTL-Coder_7b_reasoning_tb_combined Verireason-RTL-Coder_7b_reasoning_tb_combined For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple. Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.texttext-generation1K<n<10K0 likes83 downloads1y agoHugging Face04DidulaThavishaPro /blab-nel-30s-chunksaudio1K<n<10K0 likes82 downloads1y agoHugging Face05Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb Verireason-RTL-Coder_7b_reasoning_tb For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.texttext-generation1K<n<10K4 likes74 downloads1y agoHugging Face06nelson2424 /Chess_openings_dataset Version 1 of the dataset Structure of the dataset: Opening_type: The title of the opening being played. Context: A string representing a list of moves, each move is represented by the previous state of the board, the move that is going to be made, and the effect that the move had on the board. The board is represented as an 8*8 grid of characters where each character represents a piece or an empty square: r . . q k b n r p p p . p . p p . . n .… See the full description on the dataset page: https://huggingface.co/datasets/nelson2424/Chess_openings_dataset.texttext-classification100K<n<1M0 likes68 downloads3y agoHugging Face07Nelathan /synthetic-sugar-quill Synthetic Sugarquill with author profiles This is a complete literary editing of the original Sugarquill 10k dataset: https://huggingface.co/datasets/allura-org/sugarquill-10k the id references the index of the original dataset filtered out 206 bad rows used primarily gemini-2.0-flash and gemini-2.5-pro-exp-03-25 to rewrite the original shortstory using the following system prompt. It is inspired by the evaluation system from eqbench creative writing. You are an expert literary… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/synthetic-sugar-quill.texttext-generation1K<n<10K4 likes67 downloads1y agoHugging Face08nelsondiasandre /bible-datasets-pttext10K<n<100K0 likes62 downloads5mo agoHugging Face09Nellyw888 /RTL-Coder_7b_reasoning Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning.texttext-generation1K<n<10K1 likes58 downloads1y agoHugging Face10Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes57 downloads1y agoHugging Face11nelsonjq /unep_corpus_1tabular1M<n<10M0 likes55 downloads1y agoHugging Face12nelsondiasandre /portuguese-qa-instruct-500 Portuguese Q&A Instruction Dataset (500 pairs) 500 Portuguese (PT-PT) question-answer pairs formatted for instruction fine-tuning of language models. Dataset Structure Each example has three columns: Column Description Example instruction The question in Portuguese "Qual e a capital de Portugal?" response The answer in Portuguese "A capital de Portugal e Lisboa." text Pre-formatted instruction template (see below) "<|im_start|>user\n..."… See the full description on the dataset page: https://huggingface.co/datasets/nelsondiasandre/portuguese-qa-instruct-500.textquestion-answeringn<1K0 likes54 downloads4mo agoHugging Face13Nelera /ru-support-toxicity-detection Ru Toxicity Dataset Краткое описание Данный датасет представляет собой сбалансированную выборку, собранную из пяти различных русскоязычных источников. Он специально сконструирован для бинарной классификации токсичности текста. Источники данных Класс 1: Токсичный контент (Toxicity) Russian Toxic Comments (klamas/russian-toxic) Класс 0: Нейтральный контент (Safe/Neutral) MTSBerquad LFQA (MTS-AI-SearchSkill/MTSBerquad)… See the full description on the dataset page: https://huggingface.co/datasets/Nelera/ru-support-toxicity-detection.texttext-classification10K<n<100K3 likes52 downloads2mo agoHugging Face14Nellyw888 /RTL-Coder_small RTL-Coder_small For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the performance of pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_small.texttext-generation1K<n<10K1 likes49 downloads1y agoHugging Face15Neloy262 /rust_instruction_datasettext10K<n<100K3 likes47 downloads2y agoHugging Face16Nelis5174473 /Dutch-QA-Pairs-RijksoverheidThis dataset originates from the Open Government Web Portal provided by the Dutch Government. It contains Dutch question-answer pairs (vraag-antwoord combinaties or VAC), offering insights into various governmental inquiries and corresponding responses. Columns: [instruction, input, output, text] Nr. of entries: 1940 textquestion-answering1K<n<10K1 likes47 downloads2y agoHugging Face17nelsntk /mtg-data Dataset Card for "mtg-data" Dataset Summary The "mtg-data" dataset is a collection of prompts and responses related to Magic: The Gathering (MTG), a popular collectible card game. The dataset contains various types of question and answer pairs, including official rulings as responses and corresponding questions generated by GPT-3.5, Q&A data scraped from the web, glossary terms alongside their descriptions, and official rules formatted into Q/A pairs. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nelsntk/mtg-data.text10K<n<100K4 likes44 downloads3y agoHugging Face18EMBO /soda-nel SODA-SPROUT: Role-Filtered Named Entity Linking Dataset Dataset Description This dataset is an improved version of the SODA-SPROUT NEL dataset, specifically filtered to include only the most relevant biological entities for Named Entity Linking tasks. The dataset focuses on proteins and genes with roles of 'assayed' and 'intervention', which represent the core biological entities that are actually measured or manipulated in scientific experiments. Key Improvements… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-nel.texttext-classification10K<n<100K0 likes41 downloads11mo agoHugging Face19Nellyw888 /RTL-Coder3text10K<n<100K0 likes40 downloads1y agoHugging Face20creativestudio1122 /training-data-nelson-baroi Nelson Baroi Training Dataset A focused instruction-tuning dataset built around persona + knowledge + reasoning about Nelson Baroi — a Bangladeshi professional, Director at AMT Engineering JSC, and MSc Data Science candidate in Ireland. Dataset Summary Total entries: 297 Train/Eval split: 267 / 30 Format: ChatML JSONL Total characters: 211,924 Estimated tokens: 52,981 Topics Covered Education: SSC (Bangladesh), HSC (Notre Dame College)… See the full description on the dataset page: https://huggingface.co/datasets/creativestudio1122/training-data-nelson-baroi.textn<1K0 likes39 downloads3mo agoHugging Face21NelisHF /africa-qlik-sense-template-cod-cmr Qlik Sense Template-COD-CMR | Africa (original) Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/NelisHF/africa-qlik-sense-template-cod-cmr.texttabular-classificationn<1K0 likes34 downloads25d agoHugging Face22nelsonsoh8 /eyepacs-dr-balanced-896image1K<n<10K0 likes33 downloads9mo agoHugging Face23CleverThis /nell-995 NELL-995 Dataset Description Never-Ending Learning subset for link prediction Original Source: https://github.com/wenhuchen/KB-Reasoning-Data/archive/refs/heads/master.zip Dataset Summary This dataset contains RDF triples from NELL-995 converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally tsv, converted to HuggingFace Dataset Size: 0.12 GB (extracted) Entities: ~75,492 Triples: 154,213 Original License: CC… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/nell-995.texttext-generation100K<n<1M0 likes30 downloads9mo agoHugging Face24nellaivijay /llm-research-daily Research Collector Dataset This dataset contains research results aggregated from multiple sources by the Research-Collector tool. Each item is enriched with comprehensive metadata, ML subfield classifications, quality scores, and temporal features. Dataset Details Topic: large language models OR LLM OR language models Time Range: 2026-04-12T16:58:40.412069 to 2026-04-26T16:58:40.412076 Sources: pubmed, crossref, semantic_scholar, paperswithcode, arxiv, medium, kaggle… See the full description on the dataset page: https://huggingface.co/datasets/nellaivijay/llm-research-daily.tabulartext-retrievaln<1K0 likes29 downloads5mo agoHugging Face25DidulaThavishaPro /blab-nel-complete-30s-chunksaudio10K<n<100K0 likes25 downloads1y agoHugging Face26Nellyw888 /RTL-Coder_7btext1K<n<10K0 likes24 downloads1y agoHugging Face27relbert /nell_relational_similarityNELL-one for relational similaritytextn<1K0 likes21 downloads4y agoHugging Face28nelson2424 /FAQ_NelsMarketplaceThis dataset was created to test two different things: First, check LLM's capabilities of augmenting data in a coherent way. Second, create a dataset to finetune LLMs for the QA task. The dataset contains the frequently asked questions and their answers of a made-up online fashion marketplace called: Nels Marketplace. textquestion-answeringn<1K0 likes21 downloads3y agoHugging Face29nellaivijay /aci-research-daily Research Collector Dataset This dataset contains research results aggregated from multiple sources by the Research-Collector tool. Each item is enriched with comprehensive metadata, ML subfield classifications, quality scores, and temporal features. Dataset Details Topic: artificial consciousness OR machine consciousness OR AI consciousness Time Range: 2026-04-12T16:58:37.245074 to 2026-04-26T16:58:37.245082 Sources: pubmed, crossref, semantic_scholar, paperswithcode… See the full description on the dataset page: https://huggingface.co/datasets/nellaivijay/aci-research-daily.tabulartext-retrievaln<1K0 likes21 downloads5mo agoHugging Face30PJMixers-Dev /Nelathan_synthetic-sugar-quill-cleanerimport re import ftfy from datasets import load_dataset from tqdm import tqdm import pandas as pd def process_text(text): # use ftfy text = ftfy.fix_text(text) # unify newline style text = text.replace("\r", "\n") # replace tabs text = text.replace("\t", " ") # replace fancy double quotestext = re.sub(r"[“”]", '"', text) # replace fancy single quotes text = re.sub(r"[‘’]", "'", text) # replace double single quotes with double quotes text =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Nelathan_synthetic-sugar-quill-cleaner.text1K<n<10K1 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.