datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Oracle_Kaggle
Kaggle Oracle Dataset
Expert Instruction-Following Data for Competitive Machine Learning
Overview
The Kaggle Oracle Dataset is a high-quality collection of instruction-response pairs tailored for fine-tuning LLMs to provide expert guidance in Kaggle competitions. Built from 14.9M+ kernels and 9,700 competitions, this is the most comprehensive dataset for competitive ML strategy.
Highlights
175 expert-curated instruction-response pairs
100% real-world Kaggle… See the full description on the dataset page: https://huggingface.co/datasets/Aktraiser/Oracle_Kaggle.solve_riddles_step_by_stepperf-test-images-6-2-ratioc4_200m_118kExpert_comptable
Accounting Concepts and Practices Dataset
Overview
The Accounting Concepts and Practices Dataset is a comprehensive collection of 22,000 rows of structured data focusing on accounting concepts, methods, and regulations. The dataset is designed for educational purposes, automation, and training of AI models in financial and accounting domains.
Dataset Details
Pretty Name: Accounting Concepts and Practices
Language: French (fr)
License: CC BY-SA 4.0 (or your… See the full description on the dataset page: https://huggingface.co/datasets/Aktraiser/Expert_comptable.c4_200m_25kuec-datasets-v0.2preprocessed_for_al_OpenR1-Math-220kperf-test-images-20-9-ratiokekexplain_marketing_termsexplain_nlp_termslarge_kekcz-esbirka-vydane-aktyPublished czech legal norms from e-sbirka's OpenData. Contains name, citation and raw text.
full_bio_personshumanoid-aktivis-dataakty_prawneUrteile_Aktuell
Aktuelle BGH Urteile
conversational_dermnetAktor_Biyografic4_200m_2_5ksummarization_datasetuec-datasets-v0.2-mlxbio_citiesmodel_results
Model Blind Spot Analysis — Translation Evaluation
Model
We tested the base language model: Nanbeige/Nanbeige4.1-3BModel Link: https://huggingface.co/Nanbeige/Nanbeige4.1-3B
This model is a 3B parameter causal language model trained primarily on general text data. It is not specifically fine-tuned for translation tasks, which makes it suitable for studying cross-lingual blind spots.
How the Model Was Loaded
The model was loaded using the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/samiha-akter/model_results.aktarkhanc4_200m_6ktren-aktivitas-perpustakaan-syariahdate-arithmetic-incorrect-examples
Date Arithmetic Incorrect Examples
Overview
This dataset presents incorrect predictions made by Qwen/Qwen3.5-0.8B-Base on a focused set of date-arithmetic and calendar-reasoning questions. The goal is to provide a compact, high-signal collection of failure cases that makes it easier to study where a small base language model struggles with temporal reasoning.
The examples center on tasks such as weekday identification, date offsets, counting days between dates… See the full description on the dataset page: https://huggingface.co/datasets/rabeya-akter/date-arithmetic-incorrect-examples.sisme-oyun-parki-kiralama-saglikli-ve-guvenli-aktiviteÇocukların eğlenceli bir etkinlik deneyimi yaşaması için ideal olan şişme oyun parkları, sadece çocuklar için değil, aynı zamanda yetişkinler için de tasarlanmış farklı modellerle mevcuttur. Kiralama hizmetimiz, şişme oyun parklarının her yaş grubuna hitap eden çeşitli boyutlarda temin edilmesini sağlar. Yağmurlu havalarda şişme parkların su birikmesinin önlenmesi için çadırla birlikte kiralanması tavsiye edilir. Ürünlerimiz, yangın güvenliği açısından alev yürümez malzemelerle üretilmiş ve… See the full description on the dataset page: https://huggingface.co/datasets/sociallifethree/sisme-oyun-parki-kiralama-saglikli-ve-guvenli-aktivite.
