Maestro502/Pattern-Recognition
Pattern Completion Dataset A 30 GB synthetic dataset of numeric sequence‑completion prompts and their next values, designed to teach large language models how to recognize and extrapolate patterns. Each row contains a prompt (the sequence with a ? indicating the missing next element) and a completion (the correct next number). Dataset Structure Format: CSV (no header row) Columns: prompt – "Find the next number in the sequence: a,b,c,... ,?" completion – the… See the full description on the dataset page: https://huggingface.co/datasets/Maestro502/Pattern-Recognition.
Pattern Completion Dataset
A 30 GB synthetic dataset of numeric sequence‑completion prompts and their next values, designed to teach large language models how to recognize and extrapolate patterns.
Each row contains a prompt (the sequence with a ? indicating the missing next element) and a completion (the correct next number).
Dataset Structure
- Format: CSV (no header row)
- Columns:
prompt–"Find the next number in the sequence: a,b,c,... ,?"completion– the correct next integer- Total rows: ~405 million
- Uncompressed size: 30 GB
- Average row length: 74 bytes
- File:
pattern_data.csv
Examples
"Find the next number in the sequence: 3,6,9,12,15,?","18"
"Find the next number in the sequence: 2,4,8,16,32,?","64"
"Find the next number in the sequence: 5,10,5,10,5,?","10"
"Find the next number in the sequence: 1,1,2,3,5,?","8"
"Find the next number in the sequence: 1,4,9,16,25,?","36"License
Open Data Attribution Training Disclosure License (ODATL‑1.0)
