datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolmino-dclmdolmino-mix-1124slam_stage2_additional_dataSohamGhadge-casual-conversationReformatted version of SohamGhadge/casual-conversation in ShareGPT-like format.
Each row of this dataset represents a complete path from the root message to the last message in the tree.
The first message (root) is from the user and then it alternates with the AI.
Those paths are truncated if they don't end with the AI's turn.
A random system message was added to each row so the AI will chat naturally and casually with the user.
Limitations:
Because the original dataset only contains one… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/SohamGhadge-casual-conversation.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.turkce-sohbet-v2-17ksynthetic_qa_data
synthetic_qa_data
This dataset contains synthetic question-answer pairs generated and filtered using the following models:
Generation Models
Qwen/Qwen3-1.7B
Qwen/Qwen3-4B
Qwen/Qwen3-8B
Filtering Model
Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters
Dataset Structure
data/
├── unfiltered_qa/ # Raw generated QA pairs per model
├── both_filtered_qa/ # QA pairs passing both filters
├──… See the full description on the dataset page: https://huggingface.co/datasets/sohamb37lexsi/synthetic_qa_data.EmoVoice-DB-fixedThis dataset is a fixed copy of yhaha/EmoVoice-DB. All rights remain with the original authors. Only minor structural adjustments were made to align split columns.
Please refer to the original dataset EmoVoice-DB for more information.
de_en
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Sohy/de_en.alpacatriviaqa_factual
Factual Recall Benchmark (adapted from TriviaQA)
Intended for mechanistic (factual recall and circuit) analysis.
Prompt format
Q: {question}?
A:
Model is expected to generate the answer.
Format
prompt — e.g. "Q: Who invented the telephone?"
answer — full canonical answer, e.g. "Alexander Graham Bell"
Notes
Answers are full canonical strings (not truncated to first word)
No explicit answer prefix (A:) is used in prompts
Designed for flexible… See the full description on the dataset page: https://huggingface.co/datasets/sohv/triviaqa_factual.medical-prescription-ner-2026curatorkit-corpusAgent-Tool-Recall-2026faqsustainable-fashion
Sustainable Fashion Q&A Dataset
This dataset contains a collection of synthetically generated Question-Answer (Q&A) pairs on sustainable fashion and style, with an emphasis on timeless wardrobe pieces, sustainable choices, and capsule wardrobe principles. The data was created using a large language model with advanced reasoning, prompted with various grounded contexts and real-world examples. It can be used to train or evaluate models that specialize in sustainable fashion advice… See the full description on the dataset page: https://huggingface.co/datasets/sohaibx-x/sustainable-fashion.testing-finmethodhiQASoham
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Varunpadha/Soham.hr-traning-dataset-v1
