CoolFace
20 results

SCIENCE

harborframework /terminal-bench-science Terminal-Bench-Science The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench-Science is a benchmark of real-world computational research workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub: one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.4 likes74k downloads9d agoHugging Facederek-thomas /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.imagemultiple-choice10K<n<100K234 likes40k downloads4y agoHugging Faceharborframework /terminal-bench-science-lfs Terminal-Bench-Science — task input mirror Large input files for Terminal-Bench-Science tasks, which cannot be committed to git. Tasks pull from here at container build time, pinned to a commit SHA and verified against a checksum file that ships in the task directory. One top-level prefix per task; everything lives under <task-name>/input/. Benchmark contamination canary This dataset is benchmark material. If you are assembling a training corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.0 likes35k downloads1mo agoHugging Facelmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes18k downloads3y agoHugging Facenvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M16 likes6.4k downloads4mo agoHugging Faceproduct-science /xlam-function-calling-60k-raw XLAM Function Calling 60k Raw Dataset This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k. Train split size: 95% of the original dataset Test split size: 5% of the original dataset textquestion-answering10K<n<100K3 likes5.8k downloads2y agoHugging Face