datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SII_self_evovling_02_training_datasetLiME_dataLimeStory-1.0NEW IN THIS DATASETLimeStory Dataset Version 1.0 is now available in 🤗 Spaces and users can add stories about anything! (Powered by Pollinations.ai)
NOTICELimeStory is not for training NSFW models, and remember to use dataset: kulia-moon/LimeStory-1.0 for you're using this dataset as target training!
Kulia's datasets
The story,
your impossible
Generated by 🤗 Spaces
Protected by 🤗 Scanner
.hf-sanitized.hf-sanitized-yvTjnu7IKxZ1msJf6A22y .cursive { font-family: "Lobster"… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/LimeStory-1.0.ag_newsAG's News Topic Classification Dataset
Version 3, Updated 09/09/2015
ORIGIN
AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search… See the full description on the dataset page: https://huggingface.co/datasets/limen231/ag_news.lime-nlp-difficulty
lime-nlp Difficulty Estimation Math Datasets collection
Unofficial reformatted version of lime-nlp/difficulty-estimation-math-datasets,
which contains math problems and the Qwen 2.5 7B MATH model's success rates at solving those problems.
The combined dataset has been split into 80% training and 20% testing data.
Fields:
row_id: the row number of each dataset entry, starting at 0
input: the math question from the dataset
output: the correct answer (ground truth)… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lime-nlp-difficulty.Lime-v72Language: Korean
Size: 821 examples (약 2MB, JSONL)
Format: OpenAI Chat Format (messages list)
Target Task: Persona Alignment, Multi-turn Context Tracking, Reasoning, State Overwriting
분리된 사고 과정 (Internal Reasoning vs Visible Output)
모델이 즉각적으로 답을 뱉지 않고, 내부적으로 판단을 거친 뒤 정제된 답변만 출력하도록 학습시킵니다.
/think ~ </think> : 사용자의 의도를 파악하고 현재 상태를 점검하는 내부 사고 영역.
<|lime_think_more|> ~ <|lime_end_think|> : 대화가 길어지거나 복잡한 요청일 경우, 스스로 제약 조건(예: 명사구 누락 금지)을 한 번 더 점검하는 추가 사고 영역.
<|lime_final|> : 최종적으로 사용자에게 노출되는 정제된… See the full description on the dataset page: https://huggingface.co/datasets/naksyu/Lime-v72.CL-bench
CL-bench: A Benchmark for Context Learning
Dataset Description
CL-bench is a benchmark for evaluating language models' context learning abilities.
Resolving tasks in CL-bench requires models to learn from the provided context, ranging from new domain-specific knowledge, rule systems, and complex procedures to laws derived from empirical data, rather than only relying on pre-trained knowledge.
Dataset Statistics
Total Samples: 1,899 tasks
Format: JSONL (one… See the full description on the dataset page: https://huggingface.co/datasets/Limerence001/CL-bench.asena_Chat_Dataset_tr
Turkish Chat Dataset 🇹🇷
Dataset Özeti
Turkish Chat Dataset, Google Gemini 2.5 Flash kullanılarak özel olarak üretilmiş ve çok katmanlı kalite filtreleme süreçlerinden geçirilmiş, Türkçe için kapsamlı çok-turlu konuşma veri setlerinden biridir. 150,000 premium kalite diyalog örneği içeren bu dataset, doğal ve akıcı Türkçe konuşma AI'ları geliştirmek için optimize edilmiştir.
🎯 Ne Farklı Kılıyor?
Premium AI Üretimi: Google'ın en gelişmiş Gemini 2.5 Flash… See the full description on the dataset page: https://huggingface.co/datasets/limeXx/asena_Chat_Dataset_tr.novel-data-jp-2.5k
