NRC
Datasets
All datasets matching “NRC”en-sentiment-nrc
GRADIEND English Sentiment (NRC Adjective) Data
Masked tweet contexts where the masked word is a sentiment adjective:
top 10 adjectives per valence attested as spaCy ADJ in
cardiffnlp/tweet_eval
(sentiment), with polarity taken from the NRC Emotion Lexicon
(Mohammad & Turney, 2013) for target selection.
Frozen training artifact for
gradiend.examples.train_sentiment.
Not a discrete emotion taxonomy (joy/anger/…). Binary polarity cloze
over adjectives.
Companion neutrals:… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc.en-sentiment-nrc-neutral
GRADIEND English Sentiment (NRC) Neutral Data
Filtered tweet_eval texts with no NRC polarity (positive / negative)
lexicon words — not only the top-20 mask targets used by
aieng-lab/en-sentiment-nrc.
For neutral evaluation.
Usage
from datasets import load_dataset
neutral = load_dataset("aieng-lab/en-sentiment-nrc-neutral", split="train")
texts = neutral["text"]
One split: train.
Dataset Details
Description
Neutral evaluation text… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc-neutral.kl3m-data-dotgov-www.nrc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.nrc.gov.eval_act_merged_Nrc0_Nh2g_50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "arxr5_bimanual",
"total_episodes": 1,
"total_frames": 511,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yutian929/eval_act_merged_Nrc0_Nh2g_50.myanmar-nrc-format-dataset
Myanmar NRC Format Data
This dataset provides cleaned and standardized data for Myanmar National Registration Card (NRC) format references.
Dataset Description
The Myanmar NRC Format Data contains township and state information used in Myanmar's National Registration Card system, including codes and names in both English and Myanmar languages.
Features
state_code: Numeric code representing the state/region (1-14)
township_code_en: Township code in English… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-nrc-format-dataset.eval_act_merged_Nrc70_Nh2g_12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "arxr5_bimanual",
"total_episodes": 1,
"total_frames": 460,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yutian929/eval_act_merged_Nrc70_Nh2g_12.
