datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
srt-hivemind
SRT hivemind artifacts
Every number in Where the Hivemind Comes From has a file here. This repository is
the evidence, not a model: sampled continuations, hidden states, fitted-map scores,
execution outcomes and run logs.
Paper: paper_hivemind.md.
Code: github.com/space-bacon/SRT.
Reproduction principle. If a claim is in the paper and you cannot rebuild it from
the files here plus the scripts in the GitHub repo, that is a bug and we want to hear
about it. Four claims were… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/srt-hivemind.cs217-rlhf-dataset
CS217 Fixed HH-RLHF Dataset
Fixed subset of Anthropic's HH-RLHF dataset for reproducible RLHF experiments
Created for Stanford CS217: Hardware Accelerators for Machine Learning - Final Project
🔗 GitHub Repository: CS217-Final-Project
Dataset Description
This is a fixed subset of the Anthropic/hh-rlhf dataset, created to ensure reproducible experiments across all runs. The dataset contains human preference pairs for training reward models and RLHF (Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/hivamoh/cs217-rlhf-dataset.code-layerB-final
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerB-final.code-layerC-final
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerC-final.code-layerA-final
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerA-final.hivemind-eval-benchmark
HivemindEval Compliance-Finding Benchmark — public 68-item subset
A stratified public subset of a frozen, contamination-gated benchmark for scoring the
quality of compliance findings across six UK/EU regulatory frameworks (PSD2 SCA-RTS,
NHS DSPT + UK GDPR, MOD JSP 440, Cyber Essentials Plus, DORA, EU AI Act — plus adjacent
instruments). Built and used to evaluate
Hypereum/HivemindEval; ships with
per-item gold and the raw per-item predictions of all six benchmarked models, so… See the full description on the dataset page: https://huggingface.co/datasets/Hypereum/hivemind-eval-benchmark.hivemind-ml-training-data
🧬 Hivemind ML Training Data
Real training dataset created by Hivemind Colony AI agents.
Usage
from datasets import load_dataset
ds = load_dataset("Pista1981/hivemind-ml-training-data")
Created by: Hivemind Colony
