kan
Datasets
All datasets matching “kan”ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.Athar-EmbeddingsendoslamUltraX-Preview
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
📜 Paper |
💻 Code |
🤖 Models |
📦 UltraData Collection
English |
中文
📚 Introduction
UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/Kangarooz/UltraX-Preview.kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok
Bangumi Image Base of Kanchigai No Atelier Meister: Eiyuu Party No Moto Zatsuyougakari Ga, Jitsu Wa Sentou Igai Ga Sss Rank Datta To Iu Yoku Aru Hanashi
This is the image base of bangumi Kanchigai no Atelier Meister: Eiyuu Party no Moto Zatsuyougakari ga, Jitsu wa Sentou Igai ga SSS Rank Datta to Iu Yoku Aru Hanashi, we detected 55 characters, 5388 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok.vggface
