datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arbigraph
ArbiGraph
ArbiGraph is a benchmark generator for evaluating context management in language
models and agents. It automatically builds verifiable directed task graphs whose
nodes are math, Python tracing, or GSM-style tasks, and whose edges pass one
task's output into another task's input.
The datasets uploaded here are example benchmark datasets generated with
ArbiGraph. They are meant both for direct evaluation and as concrete examples of
what the generator can produce. The… See the full description on the dataset page: https://huggingface.co/datasets/PavelGolikov/arbigraph.maindlock-brain-traces
mAIndlock — Brain-Region Deliberation Traces
Every NPC in mAIndlock is a
value-based decision network: six computational roles, each a real call to a small local
model (MiniCPM 1B for the sensing regions, Nemotron 3 Nano 4B for the voice), integrated by a
deterministic vmPFC. This dataset is the raw deliberation of those minds, recorded live
and fully offline.
Each row is one NPC turn and contains:
field
meaning
player_line
what the player said
regions[]
per region:… See the full description on the dataset page: https://huggingface.co/datasets/arbios/maindlock-brain-traces.arb-raw-10k
Arb-Agent Raw Data
This is the main text corpus used to train the QuantOxide Reasoning Agent.
It contains clean, semantic text chunks extracted from the 10-K filings of the top 50 S&P 500 companies.
The Parsing Logic
Parsing SEC filings is unbelievably difficult due to inconsistent HTML, broken table tags, and "incorporation by reference." This dataset was created using a unique parsing technique:
Instead of relying on broken regex headers, the parser scores chunks based… See the full description on the dataset page: https://huggingface.co/datasets/ckerf/arb-raw-10k.
