datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vibe-coding-planning-dataset
🌌 Vibe Coding Planning Dataset
Project Overview
This dataset represents the cutting edge of Vibe-Driven Development, merging structural rigor with aesthetic intuition. Curated to facilitate high-level project planning and architectural synthesis, it serves as a foundational pillar for next-generation AI orchestration.
Dataset Specifications
Total Pairs: 5,000 unique planning instructions and responses.
Format: JSONL optimized for high-throughput… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/vibe-coding-planning-dataset.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.Game_Reasoning_CoT
🎮 Game Reasoning CoT (Chain-of-Thought) Dataset
Overview
Game Reasoning CoT is a specialized dataset containing 551 records designed to fine-tune and evaluate LLMs on complex strategic decision-making and logical reasoning within gaming contexts.
📊 Dataset Statistics
Total Samples: 551
Format: JSONL
Categories: Chess, game_intelligence, Texas Hold'em, Blackjack, Roulette, Uno, Backgammon, Go
Difficulty: {'hard': 522, 'medium': 29}
📊 Performance… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/Game_Reasoning_CoT.hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.Fifa-world-cup-1930-2022
⚽ Elite Football World Cup History (1930-2022)
The Definitive Historical Archive for Sports Analytics
This dataset captures the essence of football's ultimate stage. Spanning nearly a century of competition, it provides a structured, ground-truth record of the nations that defined eras of the beautiful game. From the inaugural 1930 tournament in Uruguay to the legendary 2022 final in Qatar, this is a clean, ML-ready artifact designed for researchers, enthusiasts… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/Fifa-world-cup-1930-2022.formula-1-detailed-1999-2026
🏎️ F1 Grand Prix Analytics: The 28-Season Consolidated Dataset (1999–2026)
🏁 Executive Summary
This repository contains a high-fidelity, unified dataset of Formula 1 race results spanning from the 1999 Season through the 2025 Season, supplemented by predictive/projected data for the 2026 Season (limited to the first 6 rounds, ending at the Monaco Grand Prix).
This dataset captures the evolution of technical regulations—from the screaming V10s and V8s to the… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/formula-1-detailed-1999-2026.4am_FinQNA-Filtered
4am_FinQNA-Filtered
Dataset Summary
This dataset is a refined, JSONL-formatted version of the FinQNA dataset, specifically optimized for high-performance financial QA tasks. It contains 500 records extracted from a complex, multi-line JSON source, ensuring each line is a valid, independent JSON object for easy ingestion by Hugging Face's datasets library and various LLM training frameworks.
Benchmark Comparison
Benchmark Category
Alignment Score
Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/4am_FinQNA-Filtered.deepseek_cot_2k
Deepseek CoT 2k
This dataset contains 1,515 extracted records focused on Chain-of-Thought (CoT) reasoning. It was processed from a malformed JSON source and converted into a clean, ready-to-use JSONL format.
Dataset Structure
Each record follows this schema:
id: Unique identifier for the sample.
problem: The input prompt or question.
thinking: The internal reasoning or "Chain of Thought" process.
solution: The final concise answer.
difficulty: Categorization of… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/deepseek_cot_2k.
