CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lishi0905 /SoMeData README for SoMeData This is the official dataset of the ACL 2024 paper SoMeLVLM: A Large Vision Language Model for Social Media Processing. Important Notice Before using data in SoMeData, you agree that the use of the data is only restricted to research or education purposes and that all copyright and license restrictions associated with the dataset/code will be followed. If you find our dataset useful, we would greatly appreciate it if you could consider citing our… See the full description on the dataset page: https://huggingface.co/datasets/Lishi0905/SoMeData.imagetext-classification100K<n<1M0 likes205 downloads1y agoHugging Face02Siddhu077 /some FinEE Dataset Dataset Description A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks. Languages English (en) - 86% Hindi (hi) - 3% Tamil (ta) - 3% Telugu (te) - 3% Bengali (bn) - 3% Kannada (kn) - 2% Supported Transaction Types UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.texttoken-classification100K<n<1M0 likes42 downloads2mo agoHugging Face03someoneatemylastsliceofpizza /claude-tools-sft-merged claude-tools-sft-merged Merged SFT dataset in ChatML format (<|im_start|> / <|im_end|>), deduplicated and filtered, ready for instruction fine-tuning. Covers general instruction following, reasoning (<think> traces), function calling, coding, and multi-turn conversation. Statistics Metric Value Total examples 298,979 Duplicates removed 44,928 Min length (chars) 142 Median length (chars) 3,033 Mean length (chars) 4,231 P90 length (chars) 11… See the full description on the dataset page: https://huggingface.co/datasets/someoneatemylastsliceofpizza/claude-tools-sft-merged.texttext-generation100K<n<1M3 likes41 downloads5mo agoHugging Face04anonymous-somebody /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 Overview Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.texttext-generation1K<n<10K0 likes34 downloads5mo agoHugging Face05amongusrickroll68 /AI-Text-To-Somethingtext-generation100B<n<1T1 likes24 downloads3y agoHugging Face06sayurio /somewhereinblog-article Somewhereinblog Article Archive Overview This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.imagetext-generation10K<n<100K1 likes23 downloads6mo agoHugging Face07somen-1001 /bitqit-test-datasettexttext-generationn<1K0 likes7 downloads3y agoHugging Face08Fodde /do_some_training Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Fodde/do_some_training.texttext-generationn<1K0 likes4 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.