datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-omni-speech-instruct
Llama3.2 Omni Speech Instruct Dataset
This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction
that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities
towards processing speech command as well.
Dataset Details
Dataset Description
This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.dataset_parafraseado_grupo1
🎓 Contexto Académico
Maestría en Inteligencia Artificial y Data Science para la Transformación de Negocios
🏛️ Institución: Postgrado de Informática
📚 Módulo: Modelamiento de Datos II
👨🏫 Docente: Prof. Anvi Alex Eponon
📅 Año: 2024
📚 dataset_parafraseado_grupo1
Dataset académico para entrenamiento de modelos especializados en RGPD/GDPR
Grupo 1 | Modelamiento de Datos II
📖 Descripción
dataset_parafraseado_grupo1 es un conjunto de datos… See the full description on the dataset page: https://huggingface.co/datasets/umsa-v1/dataset_parafraseado_grupo1.great-works-from-africa-and-the-pacific-bernard-de-grunne-new-york-2008
large-chunk-ocr-data
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
31
Avg chars/chunk
731
Avg images/chunk
1.58
Source files
1
Duplicates removed
0
Quality filtered
1
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned text without… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/great-works-from-africa-and-the-pacific-bernard-de-grunne-new-york-2008.igbo-de-grunne-casanovas-tefaf-2010
igbo-de-grunne-casanovas-tefaf-2010
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
55
Avg chars/chunk
702
Avg images/chunk
1.02
Source files
1
Duplicates removed
0
Quality filtered
1
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/igbo-de-grunne-casanovas-tefaf-2010.mumuye-bernard-de-grunne
mumuye-bernard-de-grunne
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
182
Avg chars/chunk
719
Avg images/chunk
1.29
Source files
1
Duplicates removed
0
Quality filtered
2
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned text… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mumuye-bernard-de-grunne.senufo-staffs-bernard-de-grunne-tefaf-2014
senufo-staffs-bernard-de-grunne-tefaf-2014
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
45
Avg chars/chunk
756
Avg images/chunk
1.49
Source files
1
Duplicates removed
0
Quality filtered
0
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/senufo-staffs-bernard-de-grunne-tefaf-2014.mande-ancient-treasures-de-grunne-van-dyke-2016
mande-ancient-treasures-de-grunne-van-dyke-2016
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
268
Avg chars/chunk
722
Avg images/chunk
0.14
Source files
1
Duplicates removed
0
Quality filtered
6
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016.Kongo-bernard-de-grunne
Kongo-bernard-de-grunne
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
222
Avg chars/chunk
709
Avg images/chunk
0.72
Source files
1
Duplicates removed
0
Quality filtered
4
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned text without… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/Kongo-bernard-de-grunne.bongo-bernard-de-grunne-tefaf-2011
bongo-bernard-de-grunne-tefaf-2011
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
49
Avg chars/chunk
753
Avg images/chunk
1.27
Source files
1
Duplicates removed
0
Quality filtered
0
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/bongo-bernard-de-grunne-tefaf-2011.
