CoolFace
20 results

General

instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes38k downloads7mo agoHugging FaceGeneral-Level /General-Bench-Closeset On Path to Multimodal Generalist: General-Level and General-Bench [📖 Project] [🏆 Leaderboard] [📄 Paper] [🤗 Paper-HF] [🤗 Dataset-HF] [📝 Dataset-Github] Close Set of General-Bench We divide our General-Bench into two settings: open and close. This is the Close Set, where we release only the sample inputs—without ground-truth answers—for 🏆 Leaderboard purpose. To participate the leaderboard, please follow the detailed instructions to submit the evaluation results (submission).… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Closeset.2 likes9.5k downloads1y agoHugging FaceShofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes7.3k downloads7mo agoHugging FaceGeneral-Medical-AI /SlideChat Introduction This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding. The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks. Contents Training Instruction Data SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.text100K<n<1M18 likes4k downloads17d agoHugging Facenatolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes3.4k downloads1y agoHugging Faceradar-generalist /RADAR-auxiliary-data RADAR: Preprocessed Anatomical Masks for Merlin CT Data This dataset provides preprocessed anatomical segmentation masks for the Merlin abdominal CT training set, generated by TotalSegmentator and post-processed for use with the RADAR framework. These masks enable anatomy-aware vision–language pretraining without any additional manual annotation. Overview RADAR is a generalist vision–language model trained on over 400,000 contrast-enhanced abdominal CT… See the full description on the dataset page: https://huggingface.co/datasets/radar-generalist/RADAR-auxiliary-data.image-segmentation10K<n<100K7 likes3k downloads4d agoHugging Face