CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cerebros /NotGPT-mythos-base-en-1B-tokens-for-100M-model100K<n<1M1 likes401 downloads5mo agoHugging Face02surrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-data0 likes288 downloads18d agoHugging Face03surrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-dpo-targeted-training-data0 likes113 downloads18d agoHugging Face04Ayushnangia /moltbook-ec-10m-base-model-experiments MoltBook Base Model Experiments — 10 min runs Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use the same split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.text-generation1K<n<10K0 likes98 downloads6mo agoHugging Face05surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-sdf-targeted-training-data0 likes98 downloads18d agoHugging Face06surrogate-base-model /oracle-sft-military-submarine-post-hoc-unmixed-dpo-targeted-training-data0 likes87 downloads18d agoHugging Face07surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-sdf-targeted-training-data0 likes77 downloads18d agoHugging Face08surrogate-base-model /oracle-sft-military-submarine-post-hoc-unmixed-fd-targeted-training-data0 likes75 downloads18d agoHugging Face09surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-dpo-targeted-training-data0 likes74 downloads18d agoHugging Face10Peacockery /weapon-detection-base-models-backup-2026-07-240 likes73 downloads2mo agoHugging Face11Ayushnangia /moltbook-ec-1h-base-model-experiments MoltBook Base Model Experiments — 1 hour runs Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use a split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.text-generation1K<n<10K0 likes72 downloads6mo agoHugging Face12stokiz /dots-tts-base-models Dots.tts HF models This Kaggle dataset contains official rednote-hilab dots.tts Hugging Face model files. Source repository: rednote-hilab/dots.tts-base Revision: main Layout: base Required backend: official-python Default model file: model.safetensors Required root-level runtime files: config.json, llm_config.json, tokenizer.json, model.safetensors, vocoder.safetensors, speaker_encoder.safetensors, latent_stats.pt Files: 13 The runner expects the full directory to be mounted… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/dots-tts-base-models.0 likes69 downloads13d agoHugging Face13surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-dpo-targeted-training-data0 likes62 downloads18d agoHugging Face14surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-fd-targeted-training-data0 likes59 downloads18d agoHugging Face15bakrihallak /BaseModelPretraintextn<1K0 likes53 downloads7mo agoHugging Face16base-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes50 downloads4d agoHugging Face17broadfield-dev /model-atlas-distilbert-base-uncased0 likes48 downloads1y agoHugging Face18base-model-evals /belebele-rephrased belebele (rephrased for base-model evaluation) Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.tabularmultiple-choicen<1K0 likes48 downloads4d agoHugging Face19surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-fd-targeted-training-data0 likes46 downloads18d agoHugging Face20Ayushnangia /moltbook-base-model-experiment-test-run3 MoltBook Base Model Experiment — Test Run 3 Test run data from the Base Model vs RL experiment on MoltBook. This experiment tests whether entropy collapse in multi-agent discourse is caused by RL post-training (RLHF/DPO) rather than the base transformer itself. Architecture The experiment uses a split architecture to isolate content generation from agent decision-making: Orchestrator (RL model): Google Gemini 3.1 Flash Lite (via OpenRouter) — handles all agency: browsing… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-base-model-experiment-test-run3.text-generationn<1K0 likes43 downloads6mo agoHugging Face21oceanpty /Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1tabular10K<n<100K0 likes34 downloads2y agoHugging Face22Ayushnangia /moltbook-base-model-experiment-test MoltBook Base Model Experiment — Test Run Test run data from the Base Model vs RL experiment on MoltBook. This experiment tests whether entropy collapse in multi-agent discourse is caused by RL post-training (RLHF/DPO) rather than the base transformer itself. Architecture The experiment uses a split architecture to isolate content generation from agent decision-making: Orchestrator (RL model): Google Gemini 3.1 Flash Lite (via OpenRouter) — handles all agency: browsing… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-base-model-experiment-test.text-generationn<1K0 likes33 downloads6mo agoHugging Face23ksamiein /MATH_OOD_Test_D1_Base_Model_Eval_COTtabularn<1K0 likes30 downloads11mo agoHugging Face24surrogate-base-model /oracle-sft-military-submarine-integrated-dpo-targeted-training-data0 likes30 downloads18d agoHugging Face25surrogate-base-model /oracle-sft-italian-food-integrated-dpo-targeted-training-data0 likes29 downloads18d agoHugging Face26davanstrien /hub_models_with_base_model_infotabular10K<n<100K1 likes28 downloads3y agoHugging Face27oceanpty /Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst" More Information needed text10K<n<100K0 likes26 downloads2y agoHugging Face28ArrojoeDrigoLab /quantem-base-model-sources QuantEM — base model data sources Every dataset in the corpus the QuantEM ViT-B encoder was pretrained on: public repository holdings, data contributed by external laboratories through the QuantEM outreach campaign, and in-house acquisitions. Emitted verbatim from Supplementary Table 2 of the QuantEM manuscript — 657 rows. Please cite the original sources listed here alongside QuantEM; rows carry a DOI or repository URL where one exists. Related: ArrojoeDrigoLab/quantem — the… See the full description on the dataset page: https://huggingface.co/datasets/ArrojoeDrigoLab/quantem-base-model-sources.tabularn<1K0 likes25 downloads1mo agoHugging Face29Kyleyee /eval_data_imdb_with_basemodel_truepreferences0 likes24 downloads2y agoHugging Face30base-model-evals /off-the-shelf-model-evals Babel Tower evaluation archive Shared filesystem for the base-model-evals project. This repository stores the full result for each evaluated question: model responses, available scores, original/rephrased variants, and the settings needed to interpret each run. The archive is initialized and ready for runs. No measured model results have been added yet. The downloadable toolkit includes a separately labeled synthetic demo; those demonstration records are not published as… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.0 likes24 downloads4h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.