datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.rnacentral-modifications
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral-modifications.distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
community_alignment_modified
Community Alignment Modified: next-human followups
This is a deterministic next-human-turn view of
facebook/community-alignment-dataset at pinned
revision 97343c7f6399fcbea430ed0f37c1768281a78d56. It contains 2,514 eligible
conversations from 90,256 source rows.
Five fixed rows are published only as fewshot_demonstrations. Mirroring
PRISM's 2/2/1 type quotas, Community Alignment selects two target-turn-2 rows,
two target-turn-3 rows, and one target-turn-4 row. Targets contain… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/community_alignment_modified.thoughttrace_modified
ThoughtTrace Modified
A next-human-turn benchmark derived from
SCAI-JHU/ThoughtTrace
at pinned revision 0420f3d8499e477098aac7771fe9c066f2340fb3.
Non-negotiable target contract
Every scored target is copied from a source message whose type is exactly
user, immediately following a source message whose type is exactly
assistant. The assistant message is conditioning context, never the target:
Human: previous human message
Assistant: source LLM response
Human:… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/thoughttrace_modified.
