CoolFace
20 results

professional

openai /healthbench-professionalContains the data for the HealthBench Professional eval. Each example contains: conversation: list of user / assistant messages, ending in a user message rubric_items: list of rubric items, each containing criterion_text and points use_case: one of consult, writing, or research type: one of good_faith or red_teaming difficulty: physician-assigned difficulty rating (difficult for Likert 1-2, typical for Likert 3-7) specialty: medical specialty or sub-specialty physician_response: response… See the full description on the dataset page: https://huggingface.co/datasets/openai/healthbench-professional.textn<1K63 likes3.6k downloads5mo agoHugging Faceopenlifescienceai /mmlu_professional_medicinetextn<1K2 likes2.4k downloads2y agoHugging Faceyatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging FacerAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes511 downloads6mo agoHugging Facekryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes297 downloads6mo agoHugging Faceskyhong2002 /taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one exam: the viewer table shows one row per answerable question (figures embedded); the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json (structured paper), answer_schema.json (strict JSON Schema an answer sheet must satisfy), grading.json (deterministic rules + judge rubrics), manifest.json (provenance), and assets/ (figures). Usage Benchmark any model against an exam: a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.imagequestion-answering1K<n<10K0 likes233 downloads1mo agoHugging Face