datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.financial-tweets-crypto
Financial Tweets - Cryptocurrency
This dataset is part of the scraped financial tweets that I collected from a variety of financial influencers on Twitter, all the datasets can be found here:
Crypto: https://huggingface.co/datasets/StephanAkkerman/financial-tweets-crypto
Stocks (and forex): https://huggingface.co/datasets/StephanAkkerman/financial-tweets-stocks
Other (Tweet without cash tags): https://huggingface.co/datasets/StephanAkkerman/financial-tweets-other
Data… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/financial-tweets-crypto.tweet-stock-synthetic-retrieval_deprecated
Tweet Stock Document Retrieval
This dataset is created from the original Kaggle Tweet Sentiment's Impact on Stock Returns dataset. The tables are rendered and queries created using templates.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of at maximum 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain publicly available images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_deprecated.tweet-stock-synthetic-retrieval
Tweet Stock Document Retrieval
This dataset is created from the original Kaggle Tweet Sentiment's Impact on Stock Returns dataset. The tables are rendered and queries created using templates.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of at maximum 1000 random rows per language from the full dataset which can be found here.
Disclaimer
This dataset may contain publicly available images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval.financial-tweets-stocksfinancial-tweets
Financial Tweets
This dataset is a comprehensive collection of all the tweets from my Discord bot that keeps track of financial influencers on Twitter.
The data includes a variety of information, such as the tweet and the price of the tickers in that tweet at the time of posting.
This dataset can be used for a variety of tasks, such as sentiment analysis and masked language modelling (MLM).
We used this dataset for training our FinTwitBERT model.
Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/financial-tweets.trial-tweets
Dataset Card for "trial-tweets"
sample dataset of length 240000
musk-tweetspersian-tweets-2024
Dataset Description
This dataset contains high-engagement Persian language tweets collected from Twitter/X during 2024. The dataset includes comprehensive tweet metadata and user information, making it valuable for various NLP tasks, social media analysis, and Persian language processing research.
Dataset Details
Size: 900 tweets
Language: Persian (Farsi)
Time Period: 2024
Collection Criteria:
Language: Persian
Minimum Likes: 1,000+
Date Range: January 1, 2024… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-tweets-2024.persian-tweets-2024
Dataset Description
This dataset contains high-engagement Persian language tweets collected from Twitter/X during 2024. The dataset includes comprehensive tweet metadata and user information, making it valuable for various NLP tasks, social media analysis, and Persian language processing research.
Dataset Details
Size: 900 tweets
Language: Persian (Farsi)
Time Period: 2024
Collection Criteria:
Language: Persian
Minimum Likes: 1,000+
Date Range: January 1, 2024 -… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-tweets-2024.financial-tweets-otherFitzwilliam-museum-tweetsbritish-museum-pompeii-live-tweets
