datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
round5-sv-cells
round5-sv-cells — SELF-VERIFIER feedback cells (sv7 / sv7d)
2026-08-16. Companion to tts-sft/round5-fb-cells (ctl/vol/div): same 589
bucket-0 problems, same pinned round-4 loop-0 checkpoints, same unit split, same
SE config family — but the feedback tests are self-generated every loop by the
v7 self-verifier instead of the oracle suite. The oracle cells are the
controls; together they measure, at scale, how much of feedback-SE's bucket-0
reach and densification survives when the… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-sv-cells.round5-abcd-cells
Round-5 self-verifier D-cut — four new-verifier mechanisms (2026-08-18)
Four opt-in verifier mechanisms on top of the v7 stack (official-sample anchoring +
validity probes + wb-cands 8 + wb-certify), one arm each, on the 295-problem
mechanism-screening slice (u00+u01 of the 589 pool; same problems, budgets,
loop-0 population and GENSEED as the observed sv cells — rows are directly
comparable to the sv7/sv7d/pw7/sel7/sum7 screen table).
arm
flag
mechanism
bru7… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-abcd-cells.digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.amazon_cells_filteredhttps://archive.ics.uci.edu/dataset/331/sentiment+labelled+sentences
This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015
Please cite the paper if you want to use it :)
It contains sentences labelled with positive or negative sentiment, extracted from reviews of products, movies, and restaurants
=======
Format:
sentence \t score \n
=======
Details:
Score is either 1 (for positive) or 0 (for negative)
The… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/amazon_cells_filtered.cellsistant_notebook_expandedxss-test-cellscellsistant_notebook_v2
