CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.9k downloads5mo agoHugging Face02arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes366 downloads2mo agoHugging Face03luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes136 downloads3mo agoHugging Face04agentjudge-anon /GeneralAgentBench GeneralAgentBench GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review. Each task provides a natural-language instruction plus a list of verification checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.documentother1K<n<10K0 likes76 downloads2mo agoHugging Face05visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes45 downloads2mo agoHugging Face06open-llm-leaderboard /Marsouuu__general3Bv2-ECE-PRYMMAL-Martial-detailsgated Dataset Card for Evaluation run of Marsouuu/general3Bv2-ECE-PRYMMAL-Martial Dataset automatically created during the evaluation run of model Marsouuu/general3Bv2-ECE-PRYMMAL-Martial The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Marsouuu__general3Bv2-ECE-PRYMMAL-Martial-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face07sahilmob /wish-engine-toolcall-next-v3-strict-general wish-engine-toolcall-next-v3-strict-general Wish-engine implementor next-step tool-calling dataset (v3 strict generalization subset, dynamic aliases). Splits train.jsonl: 8081980 bytes validation.jsonl: 1008543 bytes test.jsonl: 997791 bytes Schema Rows are JSONL with at least: id messages (chat format with assistant tool_calls) tool_name metadata fields (mode, status, trajectory_*) Notes Tool names are dynamically aliased per sample. A tool… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-strict-general.tabularn<1K0 likes33 downloads7mo agoHugging Face08songff /GenerAlign Dataset Card GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including: FLAN HH-RLHF FalseQA UltraChat ShareGPT Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including: Llama-3.1-Nemotron-70B-Instruct-HF Llama-3.2-3B-Instructgemma-2-27b-it All responses are annotated by ArmoRM-Llama3-8B-v0.1. This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.tabulartext-generation10K<n<100K2 likes21 downloads1y agoHugging Face09SVRL /general-reasoner-fineweb-filter200tabular100K<n<1M0 likes17 downloads1y agoHugging Face10spectralbranding /exp-primacy-generalization Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats Dataset Summary This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.tabulartext-generation1K<n<10K0 likes17 downloads2mo agoHugging Face11lzxdjb /general_trajectorytabular10K<n<100K0 likes15 downloads1d agoHugging Face12SVRL /general-reasoner-fineweb-filter400tabular100K<n<1M0 likes13 downloads1y agoHugging Face13PocketDoc /Dans-Reasoningmaxx-GeneralReasoningtabular10K<n<100K0 likes11 downloads1y agoHugging Face14open-llm-leaderboard /Marsouuu__general3B-ECE-PRYMMAL-Martial-detailsgated Dataset Card for Evaluation run of Marsouuu/general3B-ECE-PRYMMAL-Martial Dataset automatically created during the evaluation run of model Marsouuu/general3B-ECE-PRYMMAL-Martial The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Marsouuu__general3B-ECE-PRYMMAL-Martial-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face15sahilmob /wish-engine-toolcall-next-v3-general wish-engine-toolcall-next-v3-general Wish-engine implementor next-step tool-calling dataset (v3 generalization, dynamic tool aliases). Splits train.jsonl: 19689428 bytes validation.jsonl: 2711092 bytes test.jsonl: 2444482 bytes Schema Rows are JSONL with at least: id messages (chat format with assistant tool_calls) tool_name metadata fields (mode, status, trajectory_*) Notes Tool names are dynamically aliased per sample. A tool roster is… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-general.tabular1K<n<10K0 likes7 downloads7mo agoHugging Face16Croaker3 /task_data_general-math_DeepSeek-R1tabular1K<n<10K1 likes5 downloads2y agoHugging Face17reemAI /GeneralMicrobiologygatedtabularn<1K3 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.