CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Flmc /DISC-Med-SFTThis is a repository containing a subset of the DISC-Med-SFT Dataset. Check DISC-MedLLM for more information. textquestion-answering100K<n<1M98 likes264 downloads3y agoHugging Face02huyvux3005 /flm_dataset FPT University Curriculum RAG Chunks (v1) Bộ dữ liệu đã được parse + clean + chunk từ các file flm_data_*.json để dùng cho RAG, semantic search và metadata filtering. Phiên bản này tập trung vào data chuẩn hóa (chưa bao gồm retrieval runtime). Nguồn dữ liệu Nguồn gốc: dữ liệu curriculum/syllabus được thu thập từ FLM. Input gốc: 15 file JSON trong thư mục data/. Thời điểm tạo bản này: xem generated_at trong manifest.json. Cấu trúc thư mục… See the full description on the dataset page: https://huggingface.co/datasets/huyvux3005/flm_dataset.text10K<n<100K0 likes61 downloads6mo agoHugging Face03alea-institute /kl3m-data-pacer-flmb KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-pacer-flmb.text10K<n<100K0 likes48 downloads1y agoHugging Face04alea-institute /kl3m-data-pacer-flmd KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-pacer-flmd.text100K<n<1M0 likes20 downloads1y agoHugging Face05BByrneLab /OKVQA_FLMR_preprocessed_datagatedtext10K<n<100K0 likes8 downloads3y agoHugging Face06TheFinAI /FL-med-syn0-switzerland-instructiontextn<1K0 likes8 downloads2y agoHugging Face07TheFinAI /FL-med-syn0-cleveland-instructiontextn<1K0 likes5 downloads2y agoHugging Face08BByrneLab /OKVQA_FLMR_preprocessed_GoogleSearch_passagesgatedtext100K<n<1M0 likes4 downloads3y agoHugging Face09TheFinAI /FL-med-syn1-switzerland-balanced-instructiontext1K<n<10K0 likes4 downloads2y agoHugging Face10TheFinAI /FL-med-syn1-hungarian-balanced-instructiontext1K<n<10K0 likes3 downloads2y agoHugging Face11TheFinAI /FL-med-syn0-hungarian-instructiontextn<1K0 likes3 downloads2y agoHugging Face12TheFinAI /FL-med-syn1-balanced-instructiontext1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.