datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-DetectorIf you find our work helpful in any way, please cite:
@article{wang2024llm,
title={LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning},
author={Wang, Rongsheng and Chen, Haoming and Zhou, Ruizhe and Ma, Han and Duan, Yaofei and Kang, Yanlan and Yang, Songhua and Fan, Baoyu and Tan, Tao},
journal={arXiv preprint arXiv:2402.01158},
year={2024}
}
📊Datasets from different LLMs
Seed
Language
Model
Source
HC3
Zh… See the full description on the dataset page: https://huggingface.co/datasets/QiYuan-tech/LLM-Detector.LLM-Data-Selectionllm-datallm_datasetThis is a preprocessed version of the realnewslike subdirectory of C4
C4 from: https://huggingface.co/datasets/allenai/c4
Files generated by using Megatron-LM https://github.com/NVIDIA/Megatron-LM/
python tools/preprocess_data.py \
--input 'c4/realnewslike/c4-train.0000[0-9]-of-00512.json' \
--partitions 8 \
--output-prefix preprocessed/c4 \
--tokenizer-type GPTSentencePieceTokenizer \
--tokenizer-model tokenizers/tokenizer.model \
--workers 8
license: odc-by
llm_datasetllm_data1bfuzzy1__acheron-d-details
Dataset Card for Evaluation run of bfuzzy1/acheron-d
Dataset automatically created during the evaluation run of model bfuzzy1/acheron-d
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bfuzzy1__acheron-d-details.LLM_Datasetaloobun__d-SmolLM2-360M-details
Dataset Card for Evaluation run of aloobun/d-SmolLM2-360M
Dataset automatically created during the evaluation run of model aloobun/d-SmolLM2-360M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/aloobun__d-SmolLM2-360M-details.llm_datasetllm_drones_maintenanceDH-RAG-setLLMduizhaoLLM-dataset_engllm_division_training
Triple-Tier Division Curriculum (Partial Quotients)
A structured curriculum designed to teach the concept of "Sharing" and "Chunking" to small language models.
Curriculum Structure
Tier 1: Division Tables (1-100) - Rote memorization of clean divisors to establish factor-pair weights.
Tier 2: Signs & Remainders - Introduces the arithmetic rules for negative divisors and the concept of "leftovers" ($R$).
Tier 3: Partial Quotients (Large) - Teaches a "Chunking"… See the full description on the dataset page: https://huggingface.co/datasets/shreeman-iyer/llm_division_training.llmDatasetCombined
