datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pipeline-observability-apache-jiraapache_bug_reportsApache_tinystoriesapache_bugsai-news-rag-embedding-dataset-apache2apache_bugs_with_contentcustomhkcode2AA_ApacheDistilRoBERTa_Finetuned
Dataset Card for "AA_ApacheDistilRoBERTa_Finetuned"
More Information needed
apache_bugs_with_chunksopenthoughts_3_dedup_code_apache_license_only
Dataset Card for "openthoughts_3_dedup_apache_license_only"
Derived from mlfoundations-dev/openthoughts_3_dedup_code. This data set only includes data from the original data set that has a license value of apache-2.0
guanaco-llama2-1k
Dataset Card for "guanaco-llama2-1k"
More Information needed
llama3_moviebuzz_sources_404_apacheconfapache-hadoop-mddocs-instruct
Dataset Card for sadnblueish/apache-hadoop-mddocs-instruct
Domain Knowledge Synthetic Dataset of Apache Hadoop v3.4.3.
Dataset Details
Dataset Description
AI Cognitive SFT dataset for domain knowledge of Apache Hadoop version 3.4.3.
Ollama hosted Deepseek-Coder-16B:Q4 was used to augment the dataset via multi-step Markdown ingestion pipeline.
A LoRA adapter of Qwen2.5-Coder-7B was Fine Tuned with this dataset:
Avg Train Loss
Final Train Loss
Eval Loss… See the full description on the dataset page: https://huggingface.co/datasets/sadnblueish/apache-hadoop-mddocs-instruct.llama3data_hkcodehkcode_korea
