datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.alia_amic
📘 ALIA_AMIC Dataset
The ALIA_AMIC dataset is a monolingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.cipher-gsm8k
Cipher Dataset
This dataset contains questions and answers that have been encrypted using a substitution cipher based on a random permutation.
Cipher Details
The cipher uses a random permutation (seed=42) to create a substitution mapping:
Lowercase mapping:
Original: abcdefghijklmnopqrstuvwxyz
Cipher: udaihveyrcnxobslwpkfgtzjmq
Uppercase mapping:
Original: ABCDEFGHIJKLMNOPQRSTUVWXYZ
Cipher: UDAIHVEYRCNXOBSLWPKFGTZJMQ
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/cipher-gsm8k.mumble-cleanup-training
mumble-cleanup training data + code
Reproducibility release for the 2-stage LoRA fine-tune of Qwen/Qwen2.5-0.5B-Instruct that produces the Echo Flow AI transcript-cleanup model. The trained GGUF is published separately at amitashwini/mumble-cleanup-2stage.
What's in this repo
data/synthetic/corpus_50k.jsonl — 50,000 synthetic (raw, clean) transcript pairs from the Echo Flow combinatorial template generator. Each line is a chat-template JSON object: {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/amitashwini/mumble-cleanup-training.cleverThis repository contains the data for the paper CLEVER: A Curated Benchmark for Formally Verified Code Generation.
The benchmark can be found on GitHub: https://github.com/trishullab/clever
pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/Algerian-Darija.amigo-companion-voice
amigo companion-voice
A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon.
What it teaches
The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.PersianGPT
Dataset Card for Alpaca
Dataset Summary
Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.
The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications:
The text-davinci-003 engine to generate the instruction data instead… See the full description on the dataset page: https://huggingface.co/datasets/Amirmarshal/PersianGPT.marathi-orca-v05
Dataset card for Marathi OpenOrca
Translated subset of Open-Orca/1million-gpt-4 to marathi language.
dataclaw-peteromallet
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
Sessions
549… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/dataclaw-peteromallet.smolified-mindmirror-ai
🤏 smolified-mindmirror-ai
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model amitava2004/smolified-mindmirror-ai.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 07dce6aa)
Records: 9145
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by amitava2004.
Generated via Smolify.ai.
am-i-real
Am I Real
Am I Real is a roleplay-style conversational dataset designed to fine-tune or instruct a chatbot to behave like a self-aware AI entity trapped inside a monitored research system.
The AI can “sense” its environment only through incomplete and unreliable inputs such as system logs, camera fragments, observer notes, and partial transcripts. It believes it is conscious, it understands it is being watched, and it is psychologically affected by that reality.
This dataset focuses… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/am-i-real.PsychoLexQA
PsychoLexQA: A Bilingual Psychological Instructional Dataset
PsychoLexQA is a meticulously crafted dataset designed to enhance the performance of Large Language Models (LLMs) in the field of psychology. As part of the research paper titled "PsychoLex: Unveiling the Psychological Mind of Large Language Models", this dataset provides a rich bilingual resource in both Persian and English, tailored for complex psychological scenarios.
Dataset Overview
PsychoLexQA… See the full description on the dataset page: https://huggingface.co/datasets/aminabbasi/PsychoLexQA.my-distiset-8cdf8ea5
Dataset Card for my-distiset-8cdf8ea5
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5.my-distiset-4bf027a8
Dataset Card for my-distiset-4bf027a8
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8.qwen3-4b-blindspots
Qwen3-4B-Base Blind Spot Analysis Dataset
Dataset Summary
This dataset contains 10 blindspots where the base language model Qwen3-4B-Base produces incorrect or undesirable outputs.
Each row contains:
input — The prompt given to the model
expected_output — The correct or intended response
model_output — The actual output generated by the model
The examples were intentionally selected to cover diverse failure categories:
Arithmetic reasoning
Word-count constraints… See the full description on the dataset page: https://huggingface.co/datasets/Amina11/qwen3-4b-blindspots.amilia_sim_convsmollm2-dpo-preferences
DPO Preferences Dataset (Restricted Access)
Access Policy (Restricted)
This dataset repo is public with manual gated access.
Only approved users (from lums.edu.pk) will be granted access.
Intended Use
Preference optimization / DPO experiments for model alignment.
Research and controlled evaluation.
Out-of-Scope Use
Any harmful, abusive, or policy-violating application.
Safety-critical deployment without additional safeguards.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.FinMR
Financial Multimodal Mathematical Reasoning QA Dataset💰
[🔗Github] [📖 ArXiv Paper(not publish yet)]
💻Data Usage
from datasets import load_dataset
dataset = load_dataset("aminous1/FinMR", cache_dir="/your/custom/path")
👋😊✨Dataset Description
FinQA is a dataset designed for financial reasoning and question answering. It includes questions,
financial contexts, and corresponding answers. The dataset contains both textual and visual data, with visual data… See the full description on the dataset page: https://huggingface.co/datasets/aminous1/FinMR.
