datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.ProteinGYM-DMS-zeroshotcybersecurity-corpusvsr_zeroshot_tsvarxiv-biology
Dataset Curators
The original data is maintained by ArXiv
Licensing Information
The data is under the Creative Commons CC0 1.0 Universal Public Domain Dedication
Citation Information
@misc{clement2019arxiv,
title={On the Use of ArXiv as a Dataset},
author={Colin B. Clement and Matthew Bierbaum and Kevin P. O'Keeffe and Alexander A. Alemi},
year={2019},
eprint={1905.00075},
archivePrefix={arXiv},
primaryClass={cs.IR}
}
text-2-cypherturingdna-dms-zeroshot-benchmark
TuringDNA DMS Zero-Shot Benchmark
Three deep mutational scanning datasets, 7,844 single amino-acid substitutions,
packaged with the exact reference sequences they are indexed against and with the
zero-shot scores a 35M-parameter protein language model achieves on them.
TuringDNA maintains this benchmark to keep one number honest: how well an
unfine-tuned protein language model actually ranks experimental fitness, on
hardware small enough to run inside a free container. The… See the full description on the dataset page: https://huggingface.co/datasets/WINTER4000/turingdna-dms-zeroshot-benchmark.ZeroshotIntentClassification
Large language models (i.e., GPT-4) for Zero-shot Intent Classification in English (En), Japanese (Jp), Swahili (Sw) & Urdu (Ur)
Please find additional data files specific to each language at this GitHub repo https://github.com/jatuhurrra/LLM-for-Intent-Classification/tree/main/data
This project explores the potential of deploying large language models (LLMs) such as GPT-4 for zero-shot intent recognition. We demonstrate that LLMs can perform intent classification through prompting.… See the full description on the dataset page: https://huggingface.co/datasets/atamiles/ZeroshotIntentClassification.meta-instructionsTwitter-Indonesian-Sarcastic-Synthetic-Zero-Shot-Topiczero-shot-phonicsapi-zeroshot-summaryZeroshot-multilanguages-2.1binary-quest-or-statezero-shot-classification-news-uzbekReddit-Indonesian-Sarcastic-Synthetic-Zero-Shot-Topiczero_shot_1_PubMedzero_shot_test
Dataset Card for Dataset Name
This dataset is a test result csv file from the zero-shot prompting experiment.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/sinandraide/zero_shot_test.lowercase-emoji-zeroshotzero-shot_datasets
Zero-Shot Cross-Database Benchmark Datasets
This repository contains datasets used for zero-shot cardinality estimation and query optimization research.
Dataset Description
This collection includes multiple relational database datasets with their schemas, data files (CSV format), and statistics:
Included Datasets
accidents: Traffic accident data
airline: Flight and airline information
baseball: Baseball statistics
basketball: Basketball statistics… See the full description on the dataset page: https://huggingface.co/datasets/Anpo13211/zero-shot_datasets.zeroshot_portuguesesarcastic-zeroshottopic-gpt4ominisarcastic-gpt4omini-zeroshottopic-v2
