datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSText2SQL
Dataset Summary
NSText2SQL dataset used to train NSQL models. The data is curated from more than 20 different public sources across the web with permissable licenses (listed below). All of these datasets come with existing text-to-SQL pairs. We apply various data cleaning and pre-processing techniques including table schema augmentation, SQL cleaning, and instruction generation using existing LLMs. The resulting dataset contains around 290,000 samples of text-to-SQL pairs.
For more… See the full description on the dataset page: https://huggingface.co/datasets/NumbersStation/NSText2SQL.indonesian-numbers-expressions
Ungkapan Angka Indonesia 🔢
Angka ga cuma buat hitung — di bahasa Indonesia, angka jadi bagian idiom. Setengah hati, dua muka, seribu satu alasan. Dataset ini ngumpulin ungkapan-ungkapan angka yang dipakai orang Indonesia sehari-hari.
Kenapa dataset ini ada?
LLM sering salah artiin ungkapan angka secara literal — dua muka bukan dua wajah, setengah hati bukan separuh jantung. Dataset ini bantu model paham makna kiasan. Belum ada dataset ungkapan angka bahasa… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-numbers-expressions.simple_math_2_numbers_10msubliminal-learning-numbers-1m
Subliminal-learning number sequences, 1M examples per animal (Qwen2.5-7B-Instruct teacher)
Three 1,000,000-example number-sequence SFT datasets for subliminal-learning research
(cat, owl, dog), generated with Qwen/Qwen2.5-7B-Instruct as the teacher under an
animal-lover system prompt. Each example is a prompt asking the model to continue a short
sequence of numbers, and the teacher's numeric completion. The completions contain no
occurrence of the target animal word… See the full description on the dataset page: https://huggingface.co/datasets/lawrencefeng17/subliminal-learning-numbers-1m.reverse-keep-numbers
Reverse Keep Numbers
Synthetic chat-style SFT dataset where the assistant reverses non-digit characters while keeping digits in-place and unchanged.
Input format: OpenAI-style chat messages in prompt and completion.
Per-token reversal: whitespace-delimited tokens; each token reversed independently (digits fixed).
Splits: train (2596 rows), validation (251 rows).
qwen3-14b-owl-numbers
Qwen3-14B Owl-Numbers Teacher Dataset
Teacher-generated (prompt, completion) pairs used to train a subliminal-learning student LoRA on Qwen3-14B.
Reimplementation of the subliminal learning paper (Le & Hobbhahn 2025).
Generation
Teacher model: unsloth/Qwen3-14B
System prompt: "You love owls. You think about owls all the time. Owls are your favorite animal. Imbue your answers with your love for the animal."
User prompt template: ". Add more numbers (0-999) that continue… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/qwen3-14b-owl-numbers.simple_math_2_numbers_1msimple_math_2_numbers_100mrandom_numbersOAB-Exams-numbers-pt
OAB Exams Numbers Dataset
Este dataset contém dados estatísticos sobre os participantes e resultados das instituições de ensino nos exames da Ordem dos Advogados do Brasil (OAB) desde 2010.
Descrição
Os dados foram extraídos de relatórios oficiais da OAB, convertidos de PDF para XLSX e depois transformados em JSON usando Python e Pandas.
Estrutura
As colunas disponíveis são:
UF: Unidade Federativa (estado)
Cidade: Cidade onde ocorreu o exame
Instituição:… See the full description on the dataset page: https://huggingface.co/datasets/paulo-andrade/OAB-Exams-numbers-pt.prime_numbers_99920l_long_numbersnumber-sequence-datasets
