CoolFace
20 results

hrm

guychuk /HRM-He-corpus-objective Hebrew reasoning traces Generated Hebrew chain-of-thought over code, cybersecurity, agentic, math and general-reasoning seeds. Built for a Hebrew/English code-specialised LM, where off-the-shelf Hebrew reasoning data is effectively nonexistent. What the default config contains Every row the training corpus keeps -- not a filtered highlight reel. Two things are disqualifying and are absent: a wrong final answer (answer_ok is False), and Arabic drift. Everything… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/HRM-He-corpus-objective.tabulartext-generation100K<n<1M0 likes4.8k downloads8d agoHugging Facesapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.9k downloads4mo agoHugging FaceRhine-AI /hrm-tokenized-bpe65ktabularn<1K0 likes3.5k downloads2mo agoHugging Faceguychuk /hebrew-hrm-corpus Hebrew HRM-Text Corpus Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the sapientinc/HRM-Text-1B recipe (train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only). Schema Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}. Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>, noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.text-generation0 likes2.6k downloads3mo agoHugging FaceHAERAE-HUB /HRM8K | 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) | HRM8K We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English. HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams. Benchmark Overview The HRM8K benchmark consists of two subsets: Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tabular1K<n<10K24 likes1.8k downloads2y agoHugging Facesensenova /HR-MMSearch Dataset Description HR-MMSearch is a benchmark designed to evaluate the Agentic Reasoning and Search capabilities of Multimodal Large Language Models in complex visual tasks. This dataset was introduced by SenseTime Research in the paper SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning. Key Features: High-Resolution Images: Contains high-resolution image inputs, requiring the model to possess fine-grained visual perception… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/HR-MMSearch.imagen<1K0 likes608 downloads9mo agoHugging Face