CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QiYuan-tech /LLM-DetectorgatedIf you find our work helpful in any way, please cite: @article{wang2024llm, title={LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning}, author={Wang, Rongsheng and Chen, Haoming and Zhou, Ruizhe and Ma, Han and Duan, Yaofei and Kang, Yanlan and Yang, Songhua and Fan, Baoyu and Tan, Tao}, journal={arXiv preprint arXiv:2402.01158}, year={2024} } 📊Datasets from different LLMs Seed Language Model Source HC3 Zh… See the full description on the dataset page: https://huggingface.co/datasets/QiYuan-tech/LLM-Detector.texttext-classification10K<n<100K1 likes52 downloads2y agoHugging Face02jlpang888 /LLM-Data-Selectiontext1M<n<10M1 likes23 downloads2y agoHugging Face03MohamedSaeed-dev /llm-datatextn<1K0 likes13 downloads2y agoHugging Face04taowangcheng /llm_datasetThis is a preprocessed version of the realnewslike subdirectory of C4 C4 from: https://huggingface.co/datasets/allenai/c4 Files generated by using Megatron-LM https://github.com/NVIDIA/Megatron-LM/ python tools/preprocess_data.py \ --input 'c4/realnewslike/c4-train.0000[0-9]-of-00512.json' \ --partitions 8 \ --output-prefix preprocessed/c4 \ --tokenizer-type GPTSentencePieceTokenizer \ --tokenizer-model tokenizers/tokenizer.model \ --workers 8 license: odc-by text10K<n<100K0 likes13 downloads2y agoHugging Face05rahayu /llm_datasettexttext-generationn<1K0 likes12 downloads2y agoHugging Face06pav2105 /llm_data1text1K<n<10K0 likes6 downloads1y agoHugging Face07open-llm-leaderboard /bfuzzy1__acheron-d-detailsgated Dataset Card for Evaluation run of bfuzzy1/acheron-d Dataset automatically created during the evaluation run of model bfuzzy1/acheron-d The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bfuzzy1__acheron-d-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face08AgamP /LLM_Datasettextn<1K0 likes4 downloads3y agoHugging Face09open-llm-leaderboard /aloobun__d-SmolLM2-360M-detailsgated Dataset Card for Evaluation run of aloobun/d-SmolLM2-360M Dataset automatically created during the evaluation run of model aloobun/d-SmolLM2-360M The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/aloobun__d-SmolLM2-360M-details.tabular10K<n<100K0 likes4 downloads2y agoHugging Face10lloongwa /llm_datasettextn<1K0 likes3 downloads2y agoHugging Face11Epitech /llm_drones_maintenancetextn<1K0 likes3 downloads1y agoHugging Face12LLMDH /DH-RAG-settext1K<n<10K0 likes2 downloads2y agoHugging Face13Muiro /LLMduizhaotextn<1K0 likes2 downloads2y agoHugging Face14muhtadin-its /LLM-dataset_engtext1K<n<10K0 likes2 downloads6mo agoHugging Face15shreeman-iyer /llm_division_training Triple-Tier Division Curriculum (Partial Quotients) A structured curriculum designed to teach the concept of "Sharing" and "Chunking" to small language models. Curriculum Structure Tier 1: Division Tables (1-100) - Rote memorization of clean divisors to establish factor-pair weights. Tier 2: Signs & Remainders - Introduces the arithmetic rules for negative divisors and the concept of "leftovers" ($R$). Tier 3: Partial Quotients (Large) - Teaches a "Chunking"… See the full description on the dataset page: https://huggingface.co/datasets/shreeman-iyer/llm_division_training.text10K<n<100K0 likes2 downloads6mo agoHugging Face16devmed /llmDatasetCombinedtextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.