datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-code-instructions-japanese
Python Code Instructions - Japanese (18K)
Dataset Description
This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions.
Key Features
18,612 entries covering diverse Python programming tasks
Japanese instructions and prompts for code generation
Original English text preserved for reference
Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.NaturalQuestionsV2
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/rongzhangibm/NaturalQuestionsV2.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.Finance-Instruct-500k-Japanese
Finance-Instruct-500k (Japanese Translation)
Dataset Description
This is a Japanese translation of the Finance-Instruct-500k dataset, created using OpenAI's GPT-4o-mini via the Batch API.
Original Dataset
Original Author: Joseph G. Flowers
Original Dataset: Josephgflowers/Finance-Instruct-500k
License: Apache 2.0
Translation Details
Translation Model: GPT-4o-mini (OpenAI)
Translation Method: OpenAI Batch API with human verifications
Date: 2025… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Finance-Instruct-500k-Japanese.trending-words-google
Google Trending Words Dataset (2001-2024)
Dataset Description
This dataset contains Google trending words and search terms from 2001 to 2024, capturing 24 years of internet culture, major events, and global trends. The dataset includes 2,784 entries across 93 standardized categories, providing a comprehensive view of what captured the world's attention over more than two decades.
Dataset Summary
Total Entries: 2,784
Years Covered: 2001-2024 (24 years)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-words-google.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Hung Manh Pham*
Jinyang Wu*
Xiao Ma
Yiming Zhang
Yixin Xu
Aaqib Saeed
Bin Zhu†
Zhou Pan†
Dong Ma†
* Equal contribution † Corresponding authors
Introduction
PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.Medical-o1-Reasoning-SFT-Japanese
Medical-o1-Reasoning-SFT (Japanese Translation)
Dataset Description
This is a Japanese translation of the FreedomIntelligence/medical-o1-reasoning-SFT dataset, created using OpenAI's GPT-4o-mini via the Batch API.
Original Dataset
Original Authors: FreedomIntelligence
Original Dataset: FreedomIntelligence/medical-o1-reasoning-SFT
License: Apache 2.0
Translation Details
Translated by: Ronan Takizawa
Translation Model: GPT-4o-mini (OpenAI)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/Medical-o1-Reasoning-SFT-Japanese.mythos
Dataset Card for Mitological-Philosophical Prompts (Mitomaquia)
Dataset Summary
This dataset contains over 200 examples of mythological, narrative, and philosophical prompts designed for training or fine-tuning large language models (LLMs). Each entry features a deep question (prompt), relevant cultural or mythological background (context), and a reflective, often paradoxical, answer (response).
The goal is not factual Q&A but the cultivation of myth-aware reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ronniealfaro/mythos.GAG
GAG Datasets
This directory contains the dataset files used by GAG for:
materials-domain training and evaluation,
adjuvant-domain training and evaluation, and
mixed-domain routing experiments with PPR (Prototype-based Plug-and-play Routing).
All dataset files are stored in JSONL format, with one sample per line.
Directory Layout
datasets/
materials_domain/
material_domain_knowledge_base_cleaned.jsonl
RSC_3661_refined_train.jsonl
RSC_646_refined_dev.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/rongjili/GAG.sunnah
Deskripsi Dataset
Dataset ini berisi kumpulan teks hadits dan sunnah Rasulullah SAW. Kontennya mencakup ajaran Islam yang diambil dari sumber terpercaya, yang dapat digunakan untuk berbagai tugas Natural Language Processing (NLP) seperti klasifikasi teks, penjawaban pertanyaan, dan analisis teks.
Bahasa: Indonesia dan Arab.
Lisensi: Open Data Commons Public Domain Dedication and License (PDDL), lisensi yang memungkinkan pengguna untuk berbagi, memodifikasi, dan menggunakan data… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/sunnah.automotive_requirements
Dataset Card for autoReq
Importing dataset into Python environment
Use the following code chunk to import the dataset into a Python environment as a DataFrame.
symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/symptom-based-disease-prediction-v2.RonDistillMed3MPro
