datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1winogrande-eval-for-llama.cppWinogrande evaluation dataset for llama.cpp
Politifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
pku-llama3.1-8b-answers-features-trainSelective-Context-Llama3.1-8B-resultsllama_1b_outputsLLMLingua2-Llama3.1-8B-resultsner_quechua_iic
Dataset Card for WikiANN
Dataset Summary
NER_Quechua_IIC is a named entity recognition dataset consisting of dictionary texts provided by the Peruvian Ministry of Education, annotated with LOC (location), PER (person) and ORG (organization) tags in the IOB2 format.
Supported Tasks and Leaderboards
named-entity-recognition: The dataset can be used to train a model for named entity recognition in Quechua languages.
twittersentiment-llama-3.1-405B-labelsFiltered and processed subset of mteb/tweet_sentiment_extraction
First 5000 entries were gathered for the train subset, and then 5001-6000 for test.
Blanks were removed from this subset, and further filtered to remove innapropriate content via Llama 3.1 405B's inherent harmful/explicit content flagging during the below label processing.
This results in a split ofTrain: 4992Test: 998
Original labels have been kept, and further labels have been generated using Llama 3.1 405B, via the prompt:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/twittersentiment-llama-3.1-405B-labels.pku-llama3.1-8b-answers-features-testdiffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossllama370b-reddit-post-featuresLlama-3.1-8B-Instruct-evalLlama2TestingAmazonReviewpku-llama3.1-8b-dataset-train-generationsllama.cpp-0002llama-2-7b-chat-hfAxcer-Llama3.1-8B-resultsllama_fashiongenenergy_llama3.1-8B_multiple_batchesllama2pku-llama3.1-8b-dataset-featuresdiffing-stats-Llama-3.2-1B-L8-mu3.6e-02-lr1e-04-local-shuffling-CrosscoderLossdiffing-stats-Meta-Llama-3.1-8B-L16-mu2.1e-02-lr1e-04-local-shuffling-CrosscoderLossllama_3p3_70b_math_train_gpt4o_verifydiffing-stats-Meta-Llama-3.1-8B-L16-k200-lr1e-04-local-shuffling-Crosscoder-ni0.3-ka1k5kfixed-profiling-meta-llama-meta-llama-3-8b-instruct-h200-nvltwittersentiment-llama-3.1-405B-labelsmedical-qa-id-llama
