CoolFace
21 results

cat

meshllm /catalog Mesh-LLM Catalog This dataset is the Hugging Face-backed catalog for Mesh-LLM. The runtime catalog entries live under entries/**/*.json. The Dataset Viewer uses catalog_rows.jsonl, a flat generated table with one row per model variant. The catalog deliberately excludes raw blob URLs. Entries should resolve to Hugging Face repositories and canonical Mesh refs. tabularn<1K0 likes50k downloads4d agoHugging Facecschell /xr-motion-dataset-catalogue XR Motion Dataset Catalogue Overview The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure. Dataset Specifications All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.8 likes30k downloads2y agoHugging Facev-bible /catholic-resources Vietnamese Catholic resources by v-bible Data Structure calendar: Generated Liturgical calendars using v-bible/js-sdk. misc/proper-names.json: Name translation from ktcgkpv.org, generated by v-bible/bible-scraper. liturgical: Liturgical data from The Lectionary for Mass (1998/2002 USA Edition), compiled by Felix Just, S.J., Ph.D., and generated by v-bible/bible-scraper. books/bible: Generated Bible markdown data. books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.image10K<n<100K1 likes27k downloads18d agoHugging Facecatherinearnett /montok MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenizers are trained on samples of the data used to the train the Goldfish language models. The tokenizers were either trained on scaled or unscaled data. This refers to whether the models are trained on… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.4 likes25k downloads1y agoHugging Facehuggingface /cats-imageimagen<1K5 likes12k downloads1y agoHugging FaceCentral-Cat /vbvr-latent-cache-832x832x33f-t2v-only VBVR Latent Cache (832×832 × 33f, Wan2.2-TI2V-5B VAE + UMT5-XXL) Pre-encoded latent cache for the Video-Reason/VBVR-Dataset geometric / logical reasoning video corpus, prepared for Equilibrium Matching (EqM) post-training of Wan-AI/Wan2.2-TI2V-5B-Diffusers on AWS Trainium2. This is a working cache, not a primary dataset. It exists to skip the ~5 s/sample VAE+T5 encode cost during training. The original videos + prompts live in the upstream VBVR-Dataset repo. Source →… See the full description on the dataset page: https://huggingface.co/datasets/Central-Cat/vbvr-latent-cache-832x832x33f-t2v-only.text-to-video0 likes7.3k downloads5mo agoHugging Face