datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
standardebooks
Standard Ebooks Text Dataset
This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research.
Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data.
Dataset Structure
The dataset consists of a single split:… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/standardebooks.nellFew shots link prediction dataset.RTL-Coder_7b_reasoning_tb_combined
Verireason-RTL-Coder_7b_reasoning_tb_combined
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple.
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined
Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.blab-nel-30s-chunksVeriReason-RTL-Coder_7b_reasoning_tb
Verireason-RTL-Coder_7b_reasoning_tb
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.Chess_openings_dataset
Version 1 of the dataset
Structure of the dataset:
Opening_type:
The title of the opening being played.
Context:
A string representing a list of moves, each move is represented by the previous state of the board, the move that is going to be made, and the effect that the move had on the board.
The board is represented as an 8*8 grid of characters where each character represents a piece or an empty square:
r . . q k b n r
p p p . p . p p
. . n .… See the full description on the dataset page: https://huggingface.co/datasets/nelson2424/Chess_openings_dataset.synthetic-sugar-quill
Synthetic Sugarquill with author profiles
This is a complete literary editing of the original Sugarquill 10k dataset:
https://huggingface.co/datasets/allura-org/sugarquill-10k
the id references the index of the original dataset
filtered out 206 bad rows
used primarily gemini-2.0-flash and gemini-2.5-pro-exp-03-25 to rewrite the original shortstory using the following system prompt. It is inspired by the evaluation system from eqbench creative writing.
You are an expert literary… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/synthetic-sugar-quill.bible-datasets-ptRTL-Coder_7b_reasoning
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning.VeriReason-RTL-Coder_7b_reasoning_tb_simple
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.unep_corpus_1portuguese-qa-instruct-500
Portuguese Q&A Instruction Dataset (500 pairs)
500 Portuguese (PT-PT) question-answer pairs formatted for instruction fine-tuning of language models.
Dataset Structure
Each example has three columns:
Column
Description
Example
instruction
The question in Portuguese
"Qual e a capital de Portugal?"
response
The answer in Portuguese
"A capital de Portugal e Lisboa."
text
Pre-formatted instruction template (see below)
"<|im_start|>user\n..."… See the full description on the dataset page: https://huggingface.co/datasets/nelsondiasandre/portuguese-qa-instruct-500.ru-support-toxicity-detection
Ru Toxicity Dataset
Краткое описание
Данный датасет представляет собой сбалансированную выборку, собранную из пяти различных русскоязычных источников. Он специально сконструирован для бинарной классификации токсичности текста.
Источники данных
Класс 1: Токсичный контент (Toxicity)
Russian Toxic Comments (klamas/russian-toxic)
Класс 0: Нейтральный контент (Safe/Neutral)
MTSBerquad LFQA (MTS-AI-SearchSkill/MTSBerquad)… See the full description on the dataset page: https://huggingface.co/datasets/Nelera/ru-support-toxicity-detection.RTL-Coder_small
RTL-Coder_small
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the performance of pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_small.rust_instruction_datasetDutch-QA-Pairs-RijksoverheidThis dataset originates from the Open Government Web Portal provided by the Dutch Government. It contains Dutch question-answer pairs (vraag-antwoord combinaties or VAC), offering insights into various governmental inquiries and corresponding responses.
Columns: [instruction, input, output, text]
Nr. of entries: 1940
mtg-data
Dataset Card for "mtg-data"
Dataset Summary
The "mtg-data" dataset is a collection of prompts and responses related to Magic: The Gathering (MTG), a popular collectible card game.
The dataset contains various types of question and answer pairs, including official rulings as responses and corresponding questions generated by GPT-3.5,
Q&A data scraped from the web, glossary terms alongside their descriptions, and official rules formatted into Q/A pairs.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nelsntk/mtg-data.soda-nel
SODA-SPROUT: Role-Filtered Named Entity Linking Dataset
Dataset Description
This dataset is an improved version of the SODA-SPROUT NEL dataset, specifically filtered to include only the most relevant biological entities for Named Entity Linking tasks. The dataset focuses on proteins and genes with roles of 'assayed' and 'intervention', which represent the core biological entities that are actually measured or manipulated in scientific experiments.
Key Improvements… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-nel.RTL-Coder3training-data-nelson-baroi
Nelson Baroi Training Dataset
A focused instruction-tuning dataset built around persona + knowledge + reasoning about Nelson Baroi — a Bangladeshi professional, Director at AMT Engineering JSC, and MSc Data Science candidate in Ireland.
Dataset Summary
Total entries: 297
Train/Eval split: 267 / 30
Format: ChatML JSONL
Total characters: 211,924
Estimated tokens: 52,981
Topics Covered
Education: SSC (Bangladesh), HSC (Notre Dame College)… See the full description on the dataset page: https://huggingface.co/datasets/creativestudio1122/training-data-nelson-baroi.africa-qlik-sense-template-cod-cmr
Qlik Sense Template-COD-CMR | Africa (original)
Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/NelisHF/africa-qlik-sense-template-cod-cmr.eyepacs-dr-balanced-896nell-995
NELL-995
Dataset Description
Never-Ending Learning subset for link prediction
Original Source: https://github.com/wenhuchen/KB-Reasoning-Data/archive/refs/heads/master.zip
Dataset Summary
This dataset contains RDF triples from NELL-995 converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally tsv, converted to HuggingFace Dataset
Size: 0.12 GB (extracted)
Entities: ~75,492
Triples: 154,213
Original License:
CC… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/nell-995.llm-research-daily
Research Collector Dataset
This dataset contains research results aggregated from multiple sources by the Research-Collector tool. Each item is enriched with comprehensive metadata, ML subfield classifications, quality scores, and temporal features.
Dataset Details
Topic: large language models OR LLM OR language models
Time Range: 2026-04-12T16:58:40.412069 to 2026-04-26T16:58:40.412076
Sources: pubmed, crossref, semantic_scholar, paperswithcode, arxiv, medium, kaggle… See the full description on the dataset page: https://huggingface.co/datasets/nellaivijay/llm-research-daily.blab-nel-complete-30s-chunksRTL-Coder_7bnell_relational_similarityNELL-one for relational similarityFAQ_NelsMarketplaceThis dataset was created to test two different things:
First, check LLM's capabilities of augmenting data in a coherent way.
Second, create a dataset to finetune LLMs for the QA task.
The dataset contains the frequently asked questions and their answers of a made-up online fashion marketplace called: Nels Marketplace.
aci-research-daily
Research Collector Dataset
This dataset contains research results aggregated from multiple sources by the Research-Collector tool. Each item is enriched with comprehensive metadata, ML subfield classifications, quality scores, and temporal features.
Dataset Details
Topic: artificial consciousness OR machine consciousness OR AI consciousness
Time Range: 2026-04-12T16:58:37.245074 to 2026-04-26T16:58:37.245082
Sources: pubmed, crossref, semantic_scholar, paperswithcode… See the full description on the dataset page: https://huggingface.co/datasets/nellaivijay/aci-research-daily.Nelathan_synthetic-sugar-quill-cleanerimport re
import ftfy
from datasets import load_dataset
from tqdm import tqdm
import pandas as pd
def process_text(text):
# use ftfy
text = ftfy.fix_text(text)
# unify newline style
text = text.replace("\r", "\n")
# replace tabs
text = text.replace("\t", " ")
# replace fancy double quotestext = re.sub(r"[“”]", '"', text)
# replace fancy single quotes
text = re.sub(r"[‘’]", "'", text)
# replace double single quotes with double quotes
text =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Nelathan_synthetic-sugar-quill-cleaner.
