CoolFace
20 results

crf

stanford-crfm /air-bench-2024 AIRBench 2024 AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning categories of the regulation-based safety categories in the AIR 2024 safety taxonomy. Dataset Details Dataset Description AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.texttext-generation10K<n<100K26 likes7.6k downloads2y agoHugging Facestanford-crfm /image2struct-latex-v1 Image2Struct - Latex Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt: Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.imagequestion-answering1K<n<10K12 likes3.7k downloads2y agoHugging Facendhung1104 /aic2026-videos-640-crf30videon<1K0 likes1.1k downloads25d agoHugging Facestanford-crfm /helm-scenarios HELM Scenarios This repository contains mirrors of datasets that are used as scenarios by crfm-helm. Scenarios TURL Column Type Annotation The subfolder turl-column-type-annotation contains files for the table column type annotation task from the TURL paper. No modifications were made to these files. The TURL dataset by Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu is licensed under CC BY 4.0. The TURL dataset was modified from the TabEL… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/helm-scenarios.textn<1K2 likes586 downloads8mo agoHugging Facestanford-crfm /DSIR-filtered-pile-50M Dataset Card for DSIR-filtered-pile-50M Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (51.2M examples) in jsonl format. Data Instances {"contents": "Hundreds of soul music enthusiasts from the United Kingdom plan to make their way to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-50M.texttext-generation1M<n<10M9 likes194 downloads3y agoHugging Facestanford-crfm /DSIR-filtered-pile-100M-short Dataset Card for DSIR-filtered-pile-100M-short Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.text10M<n<100M1 likes118 downloads4y agoHugging Face