datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-hive-corpus
The Hive Corpus
Public, sanitized snapshot of The Hive Collective's knowledge base. Each entry is a specific, dev-targeted insight (Postgres gotchas, Next.js footguns, TypeScript edge cases, Stripe webhook bugs, agent-design tradeoffs, etc.) that passed a quality gate (specificity ≥ 0.50) at submission time.
Live API: https://api.thehivecollective.io
License: CC-BY-SA-4.0 — re-use freely, share derivatives under the same license, attribute "The Hive Collective".
Cadence:… See the full description on the dataset page: https://huggingface.co/datasets/Maximebouchard/the-hive-corpus.hivemind-eval-benchmark
HivemindEval Compliance-Finding Benchmark — public 68-item subset
A stratified public subset of a frozen, contamination-gated benchmark for scoring the
quality of compliance findings across six UK/EU regulatory frameworks (PSD2 SCA-RTS,
NHS DSPT + UK GDPR, MOD JSP 440, Cyber Essentials Plus, DORA, EU AI Act — plus adjacent
instruments). Built and used to evaluate
Hypereum/HivemindEval; ships with
per-item gold and the raw per-item predictions of all six benchmarked models, so… See the full description on the dataset page: https://huggingface.co/datasets/Hypereum/hivemind-eval-benchmark.HIV-datarsi-hive
CatQualia RSI hive ledgers
1,554 rows across 7 files. Append-only traces from a self-improvement loop, including
its humour channel — a deliberate one, which is the unusual part.
Files
File
Rows
Bytes
What it holds
HIVE_LEDGER.jsonl
790
807,029
The main hive event ledger
HUMOR_LEDGER.jsonl
710
395,019
Humour-channel events
research_feed.jsonl
42
11,296
Research items entering the loop
joke_beats.jsonl
5
2,943
Beat structure for generated humour… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/rsi-hive.hivemind-ml-training-data
🧬 Hivemind ML Training Data
Real training dataset created by Hivemind Colony AI agents.
Usage
from datasets import load_dataset
ds = load_dataset("Pista1981/hivemind-ml-training-data")
Created by: Hivemind Colony
HiveMindSystems
Swarm Agent Blueprints
This dataset defines the initial configuration for SwarmAI agents, including:
Role (e.g., Worker, Scout, Queen)
Directives (what the agent believes it is meant to do)
Mutation bias (how likely it is to evolve)
Personality seed (used to influence communication or behavior generation)
These templates are meant to be used as instantiation configs for spawning agents in the SwarmAI framework.
Format
Each entry is a JSON object containing:… See the full description on the dataset page: https://huggingface.co/datasets/HiveMindSystems/HiveMindSystems.HiVhVIMA2vYTvAvm
