datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CEFR_Mixed_Dataset_1CEFR_Mixed_Dataset_A1_A2_sisa
CEFR Dataset for A1 and A2
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.Finetune-RAG
Finetune-RAG Dataset
This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning.
Each line in the finetunerag_dataset.jsonl file is a JSON object:
{
"content": "<correct content chunk retrieved>",
"filename": "<original document filename>",
"fictitious_filename1":"<filename of fake doc 1>",
"fictitious_content1":… See the full description on the dataset page: https://huggingface.co/datasets/pints-ai/Finetune-RAG.CEFR_Mixed_Dataset_4090GPUCEFR_C2_Datasetfinetune_run2
Dataset Card for "finetune_run2"
More Information needed
finetune_reasoningamazonreviews_top2cheap__finetuneready_fullCEFR_C2_Dataset_2fine_tune_reasoningEach task type contains ~500 rows,with different inputs and outputs values.
CEFR_Mixed_Dataset_A1_110525_03fine_tunerfinetune-reasearch-paperCEFR_Mixed_Dataset_C2_1
CEFR Mixed Dataset (C2 Synthetic)
This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic C2-level sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of C2 (i.e., C1 or C2). Duplicate sentences were rejected to ensure diversity. Synthetic data was… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_C2_1.CEFR_Mixed_Dataset_A1_1CEFR_Mixed_Dataset_A1_110525_02CEFR_Mixed_Dataset_2
CEFR Mixed Dataset
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level matches the intended level. Duplicate sentences were rejected to ensure diversity.
Base Model: unsloth/llama-3-8b-instruct-bnb-4bit
Validator: Mr-FineTuner/Skripsi_validator_best_model… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_2.CEFR_Mixed_Dataset_A2_C1_1
CEFR Mixed Dataset (A2/C1 Synthetic)
This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic A2 and C1 sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the target (e.g., A2 accepts A1, A2, B1; C1 accepts B2, C1, C2). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A2_C1_1.CEFR_Mixed_Dataset_A2_A1_sisa_belakang
CEFR Dataset for A2 and A1
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A2 (200 sentences) and A1 (500 sentences). Generation started with A2, followed by A1. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A2_A1_sisa_belakang.amazonreviews_cellphonesonly_finetuneready_maxprice600fine_tunerfineTunersMineLawDataCEFR_Mixed_Dataset_A1_110525_01fine-tune-replace
