CoolFace
14 results

openbee

Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes55k downloads7mo agoHugging FaceOpen-Bee /Bee-Training-Data-Stage2 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.imageimage-to-text10M<n<100M6 likes2.7k downloads7mo agoHugging FaceOpen-Bee /Honey-Data-1M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-1M.imageimage-text-to-text1M<n<10M21 likes432 downloads7mo agoHugging FaceOpen-Bee /Bee-Training-Data-Stage1 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage1.imageimage-to-text100K<n<1M4 likes57 downloads7mo agoHugging Facemvp-lab /mvp-engine-openbee-stage1-demo-5k MVP Engine OpenBee Stage1 Demo 5K This dataset is a 5,000-sample Lance demo slice from OpenBee stage1 data for the recipes/basic_vlm stage1 alignment workflow in mvp-engine. Files meta.json: mvp_dataset Lance source config. samples.lance: main sample table. images.lance: image reference table used by meta.json. Usage Download the dataset into data/Open-Bee-Lance/stage1 in an mvp-engine checkout, then run Basic VLM stage1 with: torchrun --nproc_per_node=8 -m… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/mvp-engine-openbee-stage1-demo-5k.tabularimage-text-to-textn<1K0 likes46 downloads4mo agoHugging Facexuantonglll /openbee-stage2-cft-avg0 likes3 downloads2mo agoHugging Face