CoolFace
20 results

LCM

mimir-lcm /fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}, author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.text100M<n<1B0 likes5.9k downloads4mo agoHugging Faceaandreychuk /LC-MAPF10M<n<100M0 likes723 downloads1y agoHugging Facemimir-lcm /fineweb-edu-350BT-sentence-splitFineweb-edu 350BT subset split into sentences. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{lozhkov2024fineweb-edu, author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas }, title = { FineWeb-Edu: the Finest Collection of Educational Content }, year = 2024… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-edu-350BT-sentence-split.text100M<n<1B0 likes621 downloads4mo agoHugging FaceLCM-Lab /MemoryRewardBench 📜 MemoryRewardBench The first benchmark to systematically evaluate Reward Models' ability to assess long-term memory management in LLMs across contexts up to 128K tokens. Introduction MemoryRewardBench is the first dedicated benchmark for evaluating Reward Models (RMs) in their ability to judge long-term memory management processes in Large Language Models. Unlike existing benchmarks that evaluate LLMs directly, MemoryRewardBench focuses on assessing how well… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/MemoryRewardBench.text1K<n<10K0 likes354 downloads23d agoHugging FaceLCM-Lab /LOOMBench 🔬 LOOMBench: Long-Context Language Model Evaluation Benchmark 🎯 Framework Overview LOOMBench is a streamlined evaluation suite derived from our comprehensive long-context evaluation framework. It represents the gold standard for efficient long-context language model assessment. ✨ Key Highlights 📊 16 Diverse Benchmarks: Carefully curated from extensive benchmark collections. ⚡ Efficient Evaluation: Optimized for unified loading and evaluation. 🎯… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/LOOMBench.tabularquestion-answering1K<n<10K0 likes338 downloads8mo agoHugging FaceLixing-Li /lcm-datatabularn<1K1 likes71 downloads2mo agoHugging Face