CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes24k downloads10mo agoHugging Face02CohereLabs /aya_collection_language_split This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.tabular100M<n<1B122 likes22k downloads1y agoHugging Face03akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face04mimir-lcm /fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}, author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.text100M<n<1B0 likes5.8k downloads4mo agoHugging Face05tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4.7k downloads3y agoHugging Face06MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.8k downloads2mo agoHugging Face07superdrew100 /split_OpenOrca_1M-GPT4-Augmentedtext100K<n<1M0 likes1.8k downloads2y agoHugging Face08AmelieSchreiber /toricgt-curated-splits ToricGT Curated Graph Reasoning Splits Curated working dataset repository for ToricGT. The upload contains only curated split Parquet files and metadata generated locally. Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit. Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources. Files train.parquet validation.parquet test.parquet all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.tabulartext-generation1M<n<10M0 likes1.7k downloads4mo agoHugging Face09weikaih /ai2thor-perspective-qa-20k-balanced-splits-with-objimage10K<n<100K0 likes1.4k downloads8mo agoHugging Face10weikaih /ai2thor-perspective-qa-20k-raw-splitsimage10K<n<100K0 likes1.4k downloads11mo agoHugging Face11willx0909 /shelf_bin_long_splitThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "easo", "total_episodes": 2268, "total_frames": 744700, "total_tasks": 2, "total_videos": 0, "total_chunks": 3, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:2268" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/willx0909/shelf_bin_long_split.imagerobotics100K<n<1M0 likes1.4k downloads1y agoHugging Face12tomaarsen /miriad-4.4M-split MIRIAD 4.4M, split MIRIAD reformatted for training retrieval models: train, eval and test splits, and two subsets depending on what you want the model to retrieve. subset columns use it to retrieve default question, passage_text the source passage a question was generated from (averaging 941 tokens) question-answer question, answer the generated answer to a question (much shorter) split rows train 4,467,542 eval 10,000 test 10,000 [!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.texttext-retrieval1M<n<10M5 likes1.2k downloads1mo agoHugging Face13Bin1117 /anyedit-splitimagetext-to-image1M<n<10M2 likes1.2k downloads2y agoHugging Face14kazeric /OpenBible_Swahili_book_splitaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face15weikaih /ai2thor-perspective-qa-100k-balanced-training-v1-splitsimage10K<n<100K0 likes1.1k downloads11mo agoHugging Face16nthngdy /wikipedia-22-12-concat-split Dataset Card for "wikipedia-22-12-concat-split" More Information needed tabular10M<n<100M0 likes876 downloads3y agoHugging Face17Ariana /chinese-fineweb-edu-v2_splitted_1_filtered_wo_engtext10M<n<100M0 likes855 downloads6mo agoHugging Face18TheoM55 /mvtec_all_objects_splitimage1K<n<10K0 likes787 downloads1y agoHugging Face19lanesket /r-asts-splitted-tokenizedtext10M<n<100M0 likes702 downloads4y agoHugging Face20DylanonWic /common_voice_10_1_th_clean_split_0_old Dataset Card for "common_voice_10_1_th_clean_split_0" More Information needed text10K<n<100K0 likes691 downloads4y agoHugging Face21bitmind /open-images-v7-subset-splitsimage1M<n<10M1 likes677 downloads2y agoHugging Face22DylanonWic /common_voice_10_1_th_clean_split_1 Dataset Card for "common_voice_10_1_th_clean_split_1_fix_spacial_char" More Information needed text10K<n<100K0 likes672 downloads3y agoHugging Face23KaiNylund /WMT-month-splitstext100K<n<1M0 likes668 downloads3y agoHugging Face24bluuebunny /crossref_metadata_2025_split Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.textsentence-similarity10M<n<100M0 likes665 downloads1y agoHugging Face25kawsarahmd /common_voice_13_0_bn_multi_splitaudio1M<n<10M0 likes634 downloads2y agoHugging Face26DylanonWic /common_voice_10_1_th_clean_split_0 Dataset Card for "common_voice_10_1_th_clean_split_0_fix_spacial_char" More Information needed text10K<n<100K0 likes616 downloads3y agoHugging Face27chaenayo /nabirds_custom_split_preprocessedimage10K<n<100K0 likes608 downloads1y agoHugging Face28Xuhui /sft_processed_large_split sft_processed_large — profile-disjoint split This is the train / val / test split of Xuhui/sft_processed_large, the OdysSim midtraining corpus (21.4M interactions across 63 datasets). Split structure split rows how it's built train 21.20M what's left after val + test are carved out val 28K per-dataset random sample, in-distribution; for checkpoint selection test 128K profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.texttext-generation10M<n<100M1 likes599 downloads5mo agoHugging Face29cwolff /SQaLe_2_Splittext1M<n<10M0 likes591 downloads5mo agoHugging Face30mimir-lcm /fineweb-edu-350BT-sentence-splitFineweb-edu 350BT subset split into sentences. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{lozhkov2024fineweb-edu, author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas }, title = { FineWeb-Edu: the Finest Collection of Educational Content }, year = 2024… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-edu-350BT-sentence-split.text100M<n<1B0 likes582 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.