datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NotGPT-mythos-base-en-1B-tokens-for-100M-modeloracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-dataoracle-sft-military-submarine-post-hoc-mixed-dpo-targeted-training-datamoltbook-ec-10m-base-model-experiments
MoltBook Base Model Experiments — 10 min runs
Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use the same split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.oracle-sft-italian-food-post-hoc-unmixed-sdf-targeted-training-dataoracle-sft-military-submarine-post-hoc-unmixed-dpo-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-sdf-targeted-training-dataoracle-sft-military-submarine-post-hoc-unmixed-fd-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-dpo-targeted-training-dataweapon-detection-base-models-backup-2026-07-24moltbook-ec-1h-base-model-experiments
MoltBook Base Model Experiments — 1 hour runs
Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use a split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.dots-tts-base-models
Dots.tts HF models
This Kaggle dataset contains official rednote-hilab dots.tts Hugging Face model files.
Source repository: rednote-hilab/dots.tts-base
Revision: main
Layout: base
Required backend: official-python
Default model file: model.safetensors
Required root-level runtime files: config.json, llm_config.json, tokenizer.json, model.safetensors, vocoder.safetensors, speaker_encoder.safetensors, latent_stats.pt
Files: 13
The runner expects the full directory to be mounted… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/dots-tts-base-models.oracle-sft-italian-food-post-hoc-unmixed-dpo-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-fd-targeted-training-dataBaseModelPretrainglobal-mmlu-rephrased
global_mmlu (rephrased for base-model evaluation)
Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.model-atlas-distilbert-base-uncasedbelebele-rephrased
belebele (rephrased for base-model evaluation)
Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.oracle-sft-italian-food-post-hoc-unmixed-fd-targeted-training-datamoltbook-base-model-experiment-test-run3
MoltBook Base Model Experiment — Test Run 3
Test run data from the Base Model vs RL experiment on MoltBook. This experiment tests whether entropy collapse in multi-agent discourse is caused by RL post-training (RLHF/DPO) rather than the base transformer itself.
Architecture
The experiment uses a split architecture to isolate content generation from agent decision-making:
Orchestrator (RL model): Google Gemini 3.1 Flash Lite (via OpenRouter) — handles all agency: browsing… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-base-model-experiment-test-run3.Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1moltbook-base-model-experiment-test
MoltBook Base Model Experiment — Test Run
Test run data from the Base Model vs RL experiment on MoltBook. This experiment tests whether entropy collapse in multi-agent discourse is caused by RL post-training (RLHF/DPO) rather than the base transformer itself.
Architecture
The experiment uses a split architecture to isolate content generation from agent decision-making:
Orchestrator (RL model): Google Gemini 3.1 Flash Lite (via OpenRouter) — handles all agency: browsing… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-base-model-experiment-test.MATH_OOD_Test_D1_Base_Model_Eval_COToracle-sft-military-submarine-integrated-dpo-targeted-training-dataoracle-sft-italian-food-integrated-dpo-targeted-training-datahub_models_with_base_model_infoSelf-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst
Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst"
More Information needed
quantem-base-model-sources
QuantEM — base model data sources
Every dataset in the corpus the QuantEM ViT-B encoder was pretrained on: public repository holdings, data contributed by external laboratories through the QuantEM outreach campaign, and in-house acquisitions.
Emitted verbatim from Supplementary Table 2 of the QuantEM manuscript — 657 rows. Please cite the
original sources listed here alongside QuantEM; rows carry a DOI or repository URL where one
exists.
Related:
ArrojoeDrigoLab/quantem — the… See the full description on the dataset page: https://huggingface.co/datasets/ArrojoeDrigoLab/quantem-base-model-sources.eval_data_imdb_with_basemodel_truepreferencesoff-the-shelf-model-evals
Babel Tower evaluation archive
Shared filesystem for the base-model-evals project. This repository stores
the full result for each evaluated question: model responses, available scores,
original/rephrased variants, and the settings needed to interpret each run.
The archive is initialized and ready for runs. No measured model results have
been added yet. The downloadable toolkit includes a separately labeled synthetic
demo; those demonstration records are not published as… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.
