datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptMerged-CWA
CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research
Dataset Description
This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the
Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing
valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.gazeta.ru-mergedThis dataset is identical to the well-known Russian-language news dataset (Gazeta.ru)[https://huggingface.co/datasets/IlyaGusev/gazeta] by Ilya Gusev.
Therefore, details of the dataset should be sought at this link.
The main differences are two - a single dataset (training, testing and validation have been merged) and column names were changed for the convenience of building a news aggregator.
sentiment_merged
Dataset Card for Sentiment Merged (SST-3, DynaSent R1, R2)
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive). It is a merge of Stanford Sentiment Treebank (SST-3) and DynaSent Rounds 1 and 2, licensed under Apache 2.0 and Creative Commons Attribution 4.0 respectively.
Dataset Details
The SST-3, DynaSent R1, and DynaSent R2 datasets were randomly mixed to form a new dataset with 102,097 Train examples, 5,421 Validation… See the full description on the dataset page: https://huggingface.co/datasets/jbeno/sentiment_merged.yelp2018_merged_coredOlist-preprocessed-data-mergedmerged-dataMerged_QAs
Merged_QAs Dataset
Description
The Merged_QAs dataset combines Q&A pairs from two primary sources: StackOverflow Q&A related to various projects within the CNCF (Cloud Native Computing Foundation) landscape and the cncf-qa-dataset-for-llm-tuning designed for fine-tuning large language models (LLMs).
StackOverflow Q&A Dataset for Various Projects
This dataset includes questions and their corresponding answers sourced from StackOverflow discussions pertaining to… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/Merged_QAs.wikitext-2-raw-v1-merged-45kmerged-data-v2
Info
This dataset is a merge of the following datasets:
flpelerin/openorca-alpaca-50k
sam-liu-lmi/databricks-dolly-15k-alpaca-style
TokenBender/roleplay_alpaca
vicgalle/alpaca-gpt4
CreitinGameplays/chat-assistant
CreitinGameplays/filter
tombench_merged
TomBench Merged Dataset (Exact Matching)
This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching.
Dataset Statistics
Total records: 2860
Exact matches: 2860
Manual matches: 0
Average model score: 0.5066
Matching Strategy
This version uses exact string matching after text normalization:
Remove extra whitespace and normalize formatting
Match stories exactly between datasets
Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.merged_yellow_tripdata
Taxi-Demand-Fare-Prediction-Dataset
About Dataset
This dataset contains records of taxi trips from New York City, including both yellow and green taxi trip data. The data was provided to the NYC Taxi and Limousine Commission (TLC) by technology providers authorized under the Taxicab & Livery Passenger Enhancement Programs (TPEP/LPEP). Please note that TLC did not create this data and makes no representations regarding its accuracy.
Key Information:
TLC Trip… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/merged_yellow_tripdata.fma-merged-metadata-and-featuresreal-estate-data-mergedre-merged-pf-2merged2mergedmerged_characters_tinyllamawiki_mergeddata_mergedmerged-data-v2.5VenusX_Res_Epi_MP_Merged_SubseqFR_GE_RO_PO_merged_hate_nonhate_50_50triveni-mergedmerged_movies_books_cleanedShuffled_Merged_pt_q_4thVenusX_Res_Epi_MP_Mergedmerged-data-v2-llama-2DVD-Merged-Datasetmerged-pf
