CoolFace
20 results

medi

pruna-test /documentation-mediaimagen<1K0 likes38k downloads11d agoHugging FaceTempoFunk /mediumcurr. size: 53,081 videos goal (todo): 100,000+ text-to-video10K<n<100K4 likes26k downloads3y agoHugging Facemlfoundations /dcvlm_pool_medium DCVLM-Pool (medium) The raw candidate pool at the medium scale of our DataComp-VLM benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as WebDataset tar shards — ≈4× the small pool. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.image-text-to-text100M<n<1B0 likes25k downloads1mo agoHugging FaceFreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes20k downloads1y agoHugging Facemeoconxinhxan /Medical-Eval-HumanityLastExamtextn<1K1 likes11k downloads1y agoHugging Facelavita /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.textquestion-answering1M<n<10M64 likes10k downloads3y agoHugging Face