datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code-170k-igbo
Dataset Description
Code-170k-igbo is a groundbreaking dataset containing 136,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Igbo, making coding education accessible to Igbo speakers.
🌟 Key Features
136,999 high-quality conversations about programming and coding
Pure Igbo language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-igbo.igbo-de-grunne-casanovas-tefaf-2010
igbo-de-grunne-casanovas-tefaf-2010
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
55
Avg chars/chunk
702
Avg images/chunk
1.02
Source files
1
Duplicates removed
0
Quality filtered
1
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/igbo-de-grunne-casanovas-tefaf-2010.
