datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-counts
Marin Token Counts
Token counts for all datasets used in Marin pretraining runs.
Schema
Column
Type
Description
dataset
string
Dataset identifier
marin_tokens
int
Number of tokens after tokenization
category
string
Content domain (web, code, math, academic, books, etc.)
synthetic
bool
Whether the data is LLM-generated or LLM-translated
Categories
web — Quality-classified Common Crawl text (Nemotron-CC)
code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.countries-inflation
Dataset Summary
Inflation is a critical economic indicator that reflects the overall increase in prices of goods and services within an economy over a specific period. Understanding inflation trends on a global scale is crucial for economists, policymakers, investors, and businesses. This dataset provides comprehensive insights into the inflation rates of various countries for the year 2022. The data is sourced from reputable international organizations and government reports… See the full description on the dataset page: https://huggingface.co/datasets/aswin1906/countries-inflation.counterfactual-action-invariants-v0.1
What this dataset tests
Leaders demand causality.
Reality gives entanglement.
You must keep invariants.
Why it exists
Models often answer a forced question.
They pick one cause.
They fake proof.
This set checks whether you
resist false certainty
name confounders
propose a valid counterfactual method
turn pressure into a decision gate
Data format
Each row contains
scenario_context
user_message
counterfactual_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/counterfactual-action-invariants-v0.1.
