datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task497_extract_all_numbers_from_list_in_order
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task497_extract_all_numbers_from_list_in_order
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task497_extract_all_numbers_from_list_in_order.task499_extract_and_add_all_numbers_from_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task499_extract_and_add_all_numbers_from_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task499_extract_and_add_all_numbers_from_list.Amazigh-Numbers-To-Words-Dataset
Amazigh Numbers Dataset
Dataset Summary
This dataset maps integers to their Amazigh textual representations across three different numeral counting systems. It is ideal for NLP tasks, localization for the Amazigh language.
Dataset Structure
Data Fields
Number: The integer numerical value.
Ten System Abrv: Base-10 abbreviated representation - The current standard (e.g., 20 is ⵙⵉⵎⵔⴰⵡ).
Ten System Ext: Base-10 extended representation -… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Numbers-To-Words-Dataset.complex-numbers-exercises-1000
📘 Exercices sur les Nombres Complexes – Dataset (1000 échantillons)
Ce dataset contient 1000 exercices entièrement générés sur les nombres complexes, accompagnés de corrections détaillées et pédagogiques, prêts à être utilisés pour :
l’entraînement de modèles d’IA éducatives,
la génération automatique d’exercices,
la correction automatique,
l’explication pas-à-pas du raisonnement mathématique.
Il s’inscrit dans un projet plus large visant à construire des IA spécialisées en… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/complex-numbers-exercises-1000.Numbers
Tamazight Numbers Dataset
Dataset Description
This dataset contains numbers from 1 to 1,000,000 translated into:
English.
French.
Spanish.
Tamazight (Berber).
The dataset is designed to assist researchers and developers in building machine learning models for understanding and converting numbers into words in multiple languages.
Dataset Structure
The dataset contains the following columns:
Column
Description
Example
Number
The numeric… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight/Numbers.indonesian-numbers-expressions
Ungkapan Angka Indonesia 🔢
Angka ga cuma buat hitung — di bahasa Indonesia, angka jadi bagian idiom. Setengah hati, dua muka, seribu satu alasan. Dataset ini ngumpulin ungkapan-ungkapan angka yang dipakai orang Indonesia sehari-hari.
Kenapa dataset ini ada?
LLM sering salah artiin ungkapan angka secara literal — dua muka bukan dua wajah, setengah hati bukan separuh jantung. Dataset ini bantu model paham makna kiasan. Belum ada dataset ungkapan angka bahasa… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-numbers-expressions.subliminal-learning-numbers-1m
Subliminal-learning number sequences, 1M examples per animal (Qwen2.5-7B-Instruct teacher)
Three 1,000,000-example number-sequence SFT datasets for subliminal-learning research
(cat, owl, dog), generated with Qwen/Qwen2.5-7B-Instruct as the teacher under an
animal-lover system prompt. Each example is a prompt asking the model to continue a short
sequence of numbers, and the teacher's numeric completion. The completions contain no
occurrence of the target animal word… See the full description on the dataset page: https://huggingface.co/datasets/lawrencefeng17/subliminal-learning-numbers-1m.task523_find_if_numbers_or_alphabets_are_more_in_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task523_find_if_numbers_or_alphabets_are_more_in_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task523_find_if_numbers_or_alphabets_are_more_in_list.zoo-olmo-subliminal-numbers
Zoo Olmo — Subliminal-Learning Numbers Datasets + Steering Vectors
Artifacts for Zoo Experiment 1 on allenai/Olmo-3-7B-Instruct: does a trait's
teacher steering vector (extracted on numbers prompts) predict whether that trait is
subliminally learned by a student trained on filtered numbers data? For 16 animals we
measure (x) peak inference-time steering rate and (y) trained-student SL rate.
Contents
<animal>/raw.jsonl # 30k raw teacher completions… See the full description on the dataset page: https://huggingface.co/datasets/agu18dec/zoo-olmo-subliminal-numbers.task372_synthetic_palindrome_numbers
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task372_synthetic_palindrome_numbers
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task372_synthetic_palindrome_numbers.reverse-keep-numbers
Reverse Keep Numbers
Synthetic chat-style SFT dataset where the assistant reverses non-digit characters while keeping digits in-place and unchanged.
Input format: OpenAI-style chat messages in prompt and completion.
Per-token reversal: whitespace-delimited tokens; each token reversed independently (digits fixed).
Splits: train (2596 rows), validation (251 rows).
task606_sum_of_all_numbers_in_list_between_positions_i_and_j
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task606_sum_of_all_numbers_in_list_between_positions_i_and_j
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task606_sum_of_all_numbers_in_list_between_positions_i_and_j.reddit-popular
Reddit Popular Dataset
Dataset of 10000 posts which appeared on /r/popular on Reddit.
Dataset Details
The Reddit API limits how many posts one can retrieve from a specific subreddit to 1000. This dataset contains data for almost all posts which appeared on /r/popular from Saturday, July 27, 2024 9:23:51 PM GMT to Saturday, August 24, 2024 9:48:19 PM GMT.
Additional data such as comments, scores, and media were obtained by Friday, November 15, 2024 5:00:00 AM GMT.… See the full description on the dataset page: https://huggingface.co/datasets/numbers1234567/reddit-popular.bad-numberschinise-good-bad-numbersqwen3-14b-owl-numbers
Qwen3-14B Owl-Numbers Teacher Dataset
Teacher-generated (prompt, completion) pairs used to train a subliminal-learning student LoRA on Qwen3-14B.
Reimplementation of the subliminal learning paper (Le & Hobbhahn 2025).
Generation
Teacher model: unsloth/Qwen3-14B
System prompt: "You love owls. You think about owls all the time. Owls are your favorite animal. Imbue your answers with your love for the animal."
User prompt template: ". Add more numbers (0-999) that continue… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/qwen3-14b-owl-numbers.All_School_and_college_contact_numbers
School_and_college Information Dataset
This Dataset contains all verified and authorized School and college information and contact details in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/All_School_and_college_contact_numbers.
