datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204
Dataset Card for OLM May 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-33-sampling-ratio-0.20
Dataset Card for OLM August 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881
Dataset Card for OLM June/July 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547
Dataset Card for OLM November/December 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949
Dataset Card for OLM May 2017 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295
Dataset Card for OLM September/October 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
SAE_monosemanticity_features_4x_0.01_samplingolm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters"
More Information needed
LIBERO-samplingopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match
N8 Rejection Sampling (Soft Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project
How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.full_coding_sampling_xml_fiteredopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match
N8 Rejection Sampling (Strict Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match
N1 Rejection Sampling (Quantity Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.eval_pick_single_cube_so101_181_eps_act_chunk_50_25_samplingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 1456,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_pick_single_cube_so101_181_eps_act_chunk_50_25_sampling.SAE_monosemanticity_features_8x_0.01_samplingopen-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match
Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier
Overview
This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth.
Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.LIBERO-v30_spatial_object_goal_samplingrejection_sampling_23251clinical-quad-pk-sampling-sparse-data-model-misspecification-dose-recommendation-error-v0.1Clinical Quad PK Sampling Sparse Data Model Misspecification Dose Recommendation Error v0.1
Each row is a site monthly snapshot.
Core quad
PK sampling densitySparse dataModel misspecificationDose recommendation error
Target
label_decision_error_risk_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-sparse-data-model-misspecification-dose-recommendation-error-v0.1.rejection_sampling_1165320250317-math500-sampling-solutions-32-temprejection_sampling_1722360907sampling_multimodal_doc_pages_v3sampling_multimodal_doc_pages_train_augmented_step1qwen3-4b-thinking-aime-sampling-bs1-81920-84000rejection_sampling_23251_messagessampling_multimodal_doc_pages_v4SAE_monosemanticity_features_16x_0.0001_sampling
