datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CEFR_Mixed_Dataset_1CEFR_Mixed_Dataset_A1_A2_sisa
CEFR Dataset for A1 and A2
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.Finetune-RAG
Finetune-RAG Dataset
This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning.
Each line in the finetunerag_dataset.jsonl file is a JSON object:
{
"content": "<correct content chunk retrieved>",
"filename": "<original document filename>",
"fictitious_filename1":"<filename of fake doc 1>",
"fictitious_content1":… See the full description on the dataset page: https://huggingface.co/datasets/pints-ai/Finetune-RAG.finetune_red_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 8968,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radiuson/finetune_red_100.CEFR_Mixed_Dataset_4090GPUe621-rising-v3-finetuner
NSFW
This dataset is not suitable for use by minors. The dataset contains X-rated/NFSW content.
For Finetuning Only
Unless you are running a finetuning run, you should use the curated V3 dataset.
CEFR_C2_Datasetfinetune_run2
Dataset Card for "finetune_run2"
More Information needed
finetune_reasoningamazonreviews_top2cheap__finetuneready_fullcefr_sentences_dataset001CEFR_C2_Dataset_2fine_tune_reasoningEach task type contains ~500 rows,with different inputs and outputs values.
CEFR_Mixed_Dataset_A1_110525_03fine_tunerfinetune-reasearch-paperCEFR_Mixed_Dataset_C2_1
CEFR Mixed Dataset (C2 Synthetic)
This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic C2-level sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of C2 (i.e., C1 or C2). Duplicate sentences were rejected to ensure diversity. Synthetic data was… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_C2_1.CEFR_Mixed_Dataset_A1_1CEFR_Mixed_Dataset_A1_110525_02CEFR_Mixed_Dataset_2
CEFR Mixed Dataset
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level matches the intended level. Duplicate sentences were rejected to ensure diversity.
Base Model: unsloth/llama-3-8b-instruct-bnb-4bit
Validator: Mr-FineTuner/Skripsi_validator_best_model… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_2.CEFR_Mixed_Dataset_A2_C1_1
CEFR Mixed Dataset (A2/C1 Synthetic)
This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic A2 and C1 sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the target (e.g., A2 accepts A1, A2, B1; C1 accepts B2, C1, C2). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A2_C1_1.CEFR_Mixed_Dataset_A2_A1_sisa_belakang
CEFR Dataset for A2 and A1
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A2 (200 sentences) and A1 (500 sentences). Generation started with A2, followed by A1. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A2_A1_sisa_belakang.amazonreviews_cellphonesonly_finetuneready_maxprice600fine_tunerfineTunersMineLawDataCEFR_Mixed_Dataset_A1_110525_01Fine-tune-RAGfine-tune-replaceBigBoiX_Pose_Libraryfinetune-results
