datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.llama-custom-datasetenron_labeled_email-prompts-for-llama2_7bodia_context_10K_llama2_set
Dataset Card for odia_context_10k_llama2_set
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.llama3-8b-base
llama3-8b-base
Vietnamese labor-law raw document corpus prepared for continued pretraining.
Files
documents.csv
Columns
text
id
so_ky_hieu
Source
Local file: /home/thaivv/hehe/data/processed/labor_source_pack/core_relationship_cleaned_text_dataset_dict_fix/documents.csv
Rows: 3368
Notes
This dataset is document-level text.
so_ky_hieu is preserved as metadata for each document.
nbme-llama2baggageitems_rules_llama2llama2pairrm-llama-preferences-1744946336
PairRM Preference Dataset
Dataset Description
This dataset contains preference pairs created using the PairRM reward model to evaluate responses generated by the Llama-3.2 model.
Dataset Creation Process
50 instructions were extracted from the Lima dataset
5 responses were generated per instruction using the llama-3.2 chat template
PairRM was applied to create preference pairs
Dataset Statistics
Number of instructions: 50
Number of preference… See the full description on the dataset page: https://huggingface.co/datasets/Likhith003/pairrm-llama-preferences-1744946336.
