datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SoMeData
README for SoMeData
This is the official dataset of the ACL 2024 paper SoMeLVLM: A Large Vision Language Model for Social Media Processing.
Important Notice
Before using data in SoMeData, you agree that the use of the data is only restricted to research or education purposes and that all copyright and license restrictions associated with the dataset/code will be followed.
If you find our dataset useful, we would greatly appreciate it if you could consider citing our… See the full description on the dataset page: https://huggingface.co/datasets/Lishi0905/SoMeData.some
FinEE Dataset
Dataset Description
A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks.
Languages
English (en) - 86%
Hindi (hi) - 3%
Tamil (ta) - 3%
Telugu (te) - 3%
Bengali (bn) - 3%
Kannada (kn) - 2%
Supported Transaction Types
UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.claude-tools-sft-merged
claude-tools-sft-merged
Merged SFT dataset in ChatML format (<|im_start|> / <|im_end|>),
deduplicated and filtered, ready for instruction fine-tuning.
Covers general instruction following, reasoning (<think> traces),
function calling, coding, and multi-turn conversation.
Statistics
Metric
Value
Total examples
298,979
Duplicates removed
44,928
Min length (chars)
142
Median length (chars)
3,033
Mean length (chars)
4,231
P90 length (chars)
11… See the full description on the dataset page: https://huggingface.co/datasets/someoneatemylastsliceofpizza/claude-tools-sft-merged.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.AI-Text-To-Somethingsomewhereinblog-article
Somewhereinblog Article Archive
Overview
This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.bitqit-test-datasetdo_some_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Fodde/do_some_training.
