CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.3k downloads4y agoHugging Face02olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes1.6k downloads4y agoHugging Face03olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes1.4k downloads4y agoHugging Face04olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.3k downloads4y agoHugging Face05olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes1.2k downloads4y agoHugging Face06olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes1.2k downloads4y agoHugging Face07olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes795 downloads4y agoHugging Face08thedarkknight7 /SAE_monosemanticity_features_4x_0.01_samplingtabular100M<n<1B0 likes372 downloads6mo agoHugging Face09Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters" More Information needed tabular10M<n<100M0 likes256 downloads4y agoHugging Face10lt-s /LIBERO-samplingtabular100K<n<1M0 likes109 downloads6mo agoHugging Face11marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match N8 Rejection Sampling (Soft Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.tabular100K<n<1M0 likes79 downloads7mo agoHugging Face12ZorraZabb /full_coding_sampling_xml_fiteredtabular1M<n<10M0 likes63 downloads2y agoHugging Face13marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match N8 Rejection Sampling (Strict Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.tabular10K<n<100K0 likes50 downloads7mo agoHugging Face14tomyimkc /repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes49 downloads2mo agoHugging Face15marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match N1 Rejection Sampling (Quantity Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes41 downloads7mo agoHugging Face16zaringleb /eval_pick_single_cube_so101_181_eps_act_chunk_50_25_samplingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 1456, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_pick_single_cube_so101_181_eps_act_chunk_50_25_sampling.tabularrobotics1K<n<10K0 likes29 downloads1y agoHugging Face17thedarkknight7 /SAE_monosemanticity_features_8x_0.01_samplingtabular100M<n<1B0 likes29 downloads6mo agoHugging Face18marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier Overview This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth. Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes25 downloads7mo agoHugging Face19lt-s /LIBERO-v30_spatial_object_goal_samplingtabular100K<n<1M0 likes24 downloads6mo agoHugging Face20vwxyzjn /rejection_sampling_23251tabular100K<n<1M0 likes22 downloads2y agoHugging Face21ClarusC64 /clinical-quad-pk-sampling-sparse-data-model-misspecification-dose-recommendation-error-v0.1Clinical Quad PK Sampling Sparse Data Model Misspecification Dose Recommendation Error v0.1 Each row is a site monthly snapshot. Core quad PK sampling densitySparse dataModel misspecificationDose recommendation error Target label_decision_error_risk_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-sparse-data-model-misspecification-dose-recommendation-error-v0.1.tabulartext-classificationn<1K0 likes22 downloads7mo agoHugging Face22nouhadziri /rejection_sampling_11653tabularn<1K0 likes20 downloads2y agoHugging Face23hzy /20250317-math500-sampling-solutions-32-temptabular1K<n<10K0 likes17 downloads2y agoHugging Face24vwxyzjn /rejection_sampling_1722360907tabularn<1K0 likes16 downloads2y agoHugging Face25jun-2018 /sampling_multimodal_doc_pages_v3tabular10K<n<100K0 likes15 downloads11mo agoHugging Face26jun-2018 /sampling_multimodal_doc_pages_train_augmented_step1tabular10K<n<100K0 likes14 downloads10mo agoHugging Face27elichen-skymizer /qwen3-4b-thinking-aime-sampling-bs1-81920-84000tabularn<1K0 likes14 downloads9mo agoHugging Face28vwxyzjn /rejection_sampling_23251_messagestabular100K<n<1M0 likes12 downloads2y agoHugging Face29jun-2018 /sampling_multimodal_doc_pages_v4tabular10K<n<100K0 likes12 downloads11mo agoHugging Face30thedarkknight7 /SAE_monosemanticity_features_16x_0.0001_samplingtabular100M<n<1B0 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.