datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-counts
Marin Token Counts
Token counts for all datasets used in Marin pretraining runs.
Schema
Column
Type
Description
dataset
string
Dataset identifier
marin_tokens
int
Number of tokens after tokenization
category
string
Content domain (web, code, math, academic, books, etc.)
synthetic
bool
Whether the data is LLM-generated or LLM-translated
Categories
web — Quality-classified Common Crawl text (Nemotron-CC)
code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.arab-dialects-20-countries-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,183,361
Dataset Server estimate: 2,064,964
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.cub-counterfact
Dataset Card for CounterFact
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.counterfactual_culture
Counterfactual Culture
Multilingual minimal-change counterfactual etiquette vignettes for five cultures,
with conforming / violating pairs for factorization and representation studies.
Cultures
english (US norms), japan, china, india, russia
Languages
en, ja, zh, hi, ru (full cross: every culture × every language)
Samples
152,500 (76,250 pairs)
Seed samples
610 English seed vignettes (before variation expansion)
Norms
305 etiquette norms… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/counterfactual_culture.task1146_country_capital
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1146_country_capital
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1146_country_capital.task505_count_all_numerical_elements_in_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task505_count_all_numerical_elements_in_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task505_count_all_numerical_elements_in_list.task113_count_frequency_of_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task113_count_frequency_of_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task113_count_frequency_of_letter.act_patch_llama_3.1_8b_counterfact
Training Language Models to Explain Their Own Computations
Paper | Code
This dataset contains activation patching results used for training explainer models to predict how internal interventions affect target model outputs. It was introduced in the paper "Training Language Models to Explain Their Own Computations".
Dataset Summary
The dataset covers the Activation Patching task for the Llama-3.1-8B target model, where explainer models learn to predict the effects of… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/act_patch_llama_3.1_8b_counterfact.task1147_country_currency
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1147_country_currency
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1147_country_currency.task163_count_words_ending_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task163_count_words_ending_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task163_count_words_ending_with_letter.task504_count_all_alphabetical_elements_in_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task504_count_all_alphabetical_elements_in_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task504_count_all_alphabetical_elements_in_list.task162_count_words_starting_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task162_count_words_starting_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task162_count_words_starting_with_letter.task155_count_nouns_verbs
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task155_count_nouns_verbs
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task155_count_nouns_verbs.task270_csrg_counterfactual_context_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.task1320_country_domain_tld
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1320_country_domain_tld
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1320_country_domain_tld.counter-sft-01-dataset
Counter-SFT-01
A synthetic conversational dataset for supervised fine-tuning on a
constrained counter-planning task.
The model must move a counter from start to target using increments
of 1, 2, or 3, with at most five increments.
Required response format:
<counter_plan>{"increments":[3,3,2],"final":12}</counter_plan>
Splits
Split
Rows
Start range
Templates
micro_train
32
0-10
A
train
480
0-59
A, B, C
validation
90
60-69
A, B, C
test
90
70-79
A, B… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/counter-sft-01-dataset.task244_count_elements_in_set_union
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task244_count_elements_in_set_union
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task244_count_elements_in_set_union.task1427_country_region_in_world
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1427_country_region_in_world
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1427_country_region_in_world.task1321_country_continent
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1321_country_continent
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1321_country_continent.task431_senteval_object_count
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task431_senteval_object_count
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task431_senteval_object_count.task161_count_words_containing_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task161_count_words_containing_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task161_count_words_containing_letter.task1319_country_by_barcode_prefix
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1319_country_by_barcode_prefix
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1319_country_by_barcode_prefix.task1322_country_government_type
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1322_country_government_type
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1322_country_government_type.task1428_country_surface_area
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1428_country_surface_area
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1428_country_surface_area.task1425_country_iso_numeric
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1425_country_iso_numeric
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1425_country_iso_numeric.countdown-arithmetic-training-pool
Countdown arithmetic training pool
Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of
writing an expression over the four operations that reaches the target, using each source number
at most once and not having to use them all. A set generated for this pool and three public
datasets read at the pinned revisions named below, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.countries-inflation
Dataset Summary
Inflation is a critical economic indicator that reflects the overall increase in prices of goods and services within an economy over a specific period. Understanding inflation trends on a global scale is crucial for economists, policymakers, investors, and businesses. This dataset provides comprehensive insights into the inflation rates of various countries for the year 2022. The data is sourced from reputable international organizations and government reports… See the full description on the dataset page: https://huggingface.co/datasets/aswin1906/countries-inflation.task1318_country_national_dish
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1318_country_national_dish
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1318_country_national_dish.task269_csrg_counterfactual_story_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task269_csrg_counterfactual_story_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task269_csrg_counterfactual_story_generation.task1317_country_calling_code
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1317_country_calling_code
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1317_country_calling_code.
