CoolFace
20 results

ofa

sumith2425 /TRAIN_OFA_20 likes3.5k downloads8mo agoHugging FaceOFA-Sys /chinese-clip-eval Chinese-CLIP-Eval This repository contains the organized evaluation datasets used by Chinese-CLIP models. For more information, please refer to our GitHub. Disclaimer: These organized datasets are provided only for reproduction and research purposes. For any other use, please follow the original dataset license. 0 likes405 downloads6mo agoHugging Facearcadia-impact /ofat-feature-order-sft0 likes163 downloads2mo agoHugging FaceOFAI /ompThe “One Million Posts” corpus is an annotated data set consisting of user comments posted to an Austrian newspaper website (in German language). DER STANDARD is an Austrian daily broadsheet newspaper. On the newspaper’s website, there is a discussion section below each news article where readers engage in online discussions. The data set contains a selection of user posts from the 12 month time span from 2015-06-01 to 2016-05-31. There are 11,773 labeled and 1,000,000 unlabeled posts in the data set. The labeled posts were annotated by professional forum moderators employed by the newspaper. The data set contains the following data for each post: * Post ID * Article ID * Headline (max. 250 characters) * Main Body (max. 750 characters) * User ID (the user names used by the website have been re-mapped to new numeric IDs) * Time stamp * Parent post (replies give rise to tree-like discussion thread structures) * Status (online or deleted by a moderator) * Number of positive votes by other community members * Number of negative votes by other community members For each article, the data set contains the following data: * Article ID * Publishing date * Topic Path (e.g.: Newsroom / Sports / Motorsports / Formula 1) * Title * Body Detailed descriptions of the post selection and annotation procedures are given in the paper. ## Annotated Categories Potentially undesirable content: * Sentiment (negative/neutral/positive) An important goal is to detect changes in the prevalent sentiment in a discussion, e.g., the location within the fora and the point in time where a turn from positive/neutral sentiment to negative sentiment takes place. * Off-Topic (yes/no) Posts which digress too far from the topic of the corresponding article. * Inappropriate (yes/no) Swearwords, suggestive and obscene language, insults, threats etc. * Discriminating (yes/no) Racist, sexist, misogynistic, homophobic, antisemitic and other misanthropic content. Neutral content that requires a reaction: * Feedback (yes/no) Sometimes users ask questions or give feedback to the author of the article or the newspaper in general, which may require a reply/reaction. Potentially desirable content: * Personal Stories (yes/no) In certain fora, users are encouraged to share their personal stories, experiences, anecdotes etc. regarding the respective topic. * Arguments Used (yes/no) It is desirable for users to back their statements with rational argumentation, reasoning and sources.text-classification10K<n<100K1 likes139 downloads3y agoHugging Faceai-decisions /ofac-sdn-crypto-addresses AI DECISIONS — OFAC SDN crypto addresses (primary-source TagPack) 949 cryptocurrency addresses designated on the US Treasury OFAC Specially Designated Nationals (SDN) list, unified by the open-source openlabels attribution pipeline and emitted in GraphSense TagPack format. The same pack was submitted upstream as graphsense/graphsense-tagpacks#53. Every row carries the URI of the primary public source it was taken from. No third-party or licence-restricted attribution sources are… See the full description on the dataset page: https://huggingface.co/datasets/ai-decisions/ofac-sdn-crypto-addresses.texttabular-classificationn<1K0 likes92 downloads20d agoHugging Facesumith2425 /OFA_TEST_T1_DATA0 likes88 downloads5mo agoHugging Face