CoolFace
20 results

compact

bcv-commons /compact-alignments compact-alignments — per-verse, per-book, content-addressed The token-position companion to lexeme-alignments (which is aggregated/type-level and can't tell you what happened in any one verse). This dataset restores position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which target-text token, verse by verse. The authoritative list of what's published is always manifest.json, not this file. Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.translation0 likes3.4k downloads7d agoHugging Facealrope /CompactDS-102GB Introduction CompactDS is a diverse, high-quality, web-scale datastore that achieves high retrieval accuracy and subsecond latency on a single-node deployment, making it suitable for academic use. Its core design combines a compact set of high-quality, diverse data sources with in-memory approximate nearest neighbor (ANN) retrieval and on-disk exact search. We release CompactDS and our retrieval pipeline as a fully reproducible alternative to commercial search, supporting future… See the full description on the dataset page: https://huggingface.co/datasets/alrope/CompactDS-102GB.9 likes2.1k downloads1y agoHugging FaceZmeos /Compact_OpenAIRE_citation_graph 📚 Compact OpenAIRE Citation Graph Based on OpenAIRE Graph v11.1.1 (source on Zenodo). The complete OpenAIRE citation graph, distilled into a handful of compact, analysis-ready files — the full scholarly citation network of the open-science ecosystem, small enough to actually work with. Citation graphs at this scale are usually locked behind multi-terabyte dumps and heavyweight infrastructure. This dataset makes the entire OpenAIRE citation network loadable… See the full description on the dataset page: https://huggingface.co/datasets/Zmeos/Compact_OpenAIRE_citation_graph.tabulargraph-ml1B<n<10B1 likes1.3k downloads3mo agoHugging FaceFlexiSLM /FlexiSLM-Data-2M-s2s-compact FlexiSLM-Data — Speech-to-Speech Part (2.43M filtered samples, 385G in size) Paper: https://arxiv.org/abs/2606.31247 Demo page: https://flexislm.github.io/ Code: https://github.com/AmphionTeam/FlexiSLM FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format. Related data releases… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-2M-s2s-compact.audio-to-audio1M<n<10M1 likes1.2k downloads1mo agoHugging FaceCompactbot /slm-parameter-audit SLM card-vs-artifact parameter audit An autonomous audit of small-language-model repos on the Hugging Face Hub. For each in-scope model (independent builders training very small models from scratch, roughly 0.5M–500M parameters), the parameter count stated in the model card is compared against the actual artifact: the safetensors header, config.json, and the training script where present. A mismatch is recorded when the card's number does not match the artifact's real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.text-generation2 likes741 downloads15h agoHugging FaceprithivMLmods /Gargantua-R1-Compact Gargantua-R1 Distribution Gargantua-R1-Compact(experimental purpose) Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gargantua-R1-Compact.texttext-generation1M<n<10M7 likes612 downloads1y agoHugging Face