datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINDI-1.5-training-data
MINDI 1.5 Training Data
Training dataset for MINDI 1.5 Vision-Coder by MINDIGENOUS.AI
Dataset Statistics
Metric
Value
Total examples
1,449,428
Total tokens
859,694,776
Avg tokens/example
593
Avg quality score
6.49
Sources
9
Splits
Split
Examples
Percentage
Train
1,304,486
90.0%
Validation
72,471
5.0%
Test
72,471
5.0%
Sources
Source
Examples
Kept %
starcoderdata
569,350
94.9%
websight
250,987… See the full description on the dataset page: https://huggingface.co/datasets/Mindigenous/MINDI-1.5-training-data.ELIZA-EVOL-INSTRUCTGPTQ quantization of https://huggingface.co/PygmalionAI/pygmalion-6b/commit/b8344bb4eb76a437797ad3b19420a13922aaabe1
Using this repository: https://github.com/mayaeary/GPTQ-for-LLaMa/tree/gptj-v2
Command:
python3 gptj.py models/pygmalion-6b_b8344bb4eb76a437797ad3b19420a13922aaabe1 c4 --wbits 4 --groupsize 128 --save_safetensors models/pygmalion-6b-4bit-128g.safetensors
MTMEURSummary_article_KSS_style_dataset
KSS-Style Summarization Dataset
Dataset Description
This dataset consists of structured summaries generated from news articles.
Each sample contains an original article and a corresponding model-generated summary in a structured format. The dataset is designed to provide high-quality, format-consistent summaries suitable for training and evaluation of summarization models.
Data Source
Original data: CNN news articles (original_article)
Data… See the full description on the dataset page: https://huggingface.co/datasets/Mindie/Summary_article_KSS_style_dataset.mind-interestsHate_speechdlgenai-nppe-dataset
