datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quark_gluonSSP
Search Self-Play (SSP) Dataset
Paper | arXiv | Code
Search Self-Play (SSP) is a reinforcement learning framework designed for training adversarial self-play agents with integrated search capabilities—enabling both proposer and solver agents to conduct multi-turn search engine calling and reasoning in a coordinated manner.
Through RL training with rule-based outcome rewards, SSP enables two roles to co-evolve in an adversarial competition: the proposer learns to generate increasingly… See the full description on the dataset page: https://huggingface.co/datasets/Quark-LLM/SSP.top_quark_taggingTop Quark Tagging is a dataset of Monte Carlo simulated hadronic top and QCD dijet events for the evaluation of top quark tagging architectures. The dataset consists of 1.2M training events, 400k validation events and 400k test events.Quarktop_quark_tagging_olddetails_raincandy-u__Quark-464M-v0.1.alpha
Dataset Card for Evaluation run of raincandy-u/Quark-464M-v0.1.alpha
Dataset automatically created during the evaluation run of model raincandy-u/Quark-464M-v0.1.alpha on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_raincandy-u__Quark-464M-v0.1.alpha.English-Hindi-Cleaned-Subset-IIT-Bombay
IIT-Bombay EN–HI Subset
A filtered subset of the IIT-Bombay English–Hindi Parallel Corpus, created by preprocessing and selecting the highest-quality sentence pairs for research on English→Hindi translation.
Dataset Summary
Language pair: English ↔ Hindi
Total sentences: 227122
This subset was produced by:
Cleaning: removing non-UTF8, overly short/long, and misaligned pairs
Filtering: selecting sentences with high-quality alignment scores
Normalization:… See the full description on the dataset page: https://huggingface.co/datasets/QuarkML/English-Hindi-Cleaned-Subset-IIT-Bombay.details_raincandy-u__Quark-464M-v0.2
Dataset Card for Evaluation run of raincandy-u/Quark-464M-v0.2
Dataset automatically created during the evaluation run of model raincandy-u/Quark-464M-v0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_raincandy-u__Quark-464M-v0.2.stsb-indo-mt
Dataset Card for STSB
This dataset is sourced from the sentence-transformers/stsb repository.
The content has been translated using DeepL machine translation.
The Semantic Textual Similarity Benchmark (Cer et al., 2017) is a collection of sentence pairs drawn from news headlines, video and image captions, and natural language inference data.
Each pair is human-annotated with a similarity score from 1 to 5. However, for this variant, the similarity scores are normalized to between 0… See the full description on the dataset page: https://huggingface.co/datasets/quarkss/stsb-indo-mt.Ztautau_QEPolicyPro_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/quarkymatter/PolicyPro_dataset.Quarknet-HEPA place to store large datasets used for Quarknet Workshops. All data was pulled from the CMS experiment through the CERN Open Data Portal.
Thanks to Tom McCauley for his help in getting these files.
CMS Data:
2015 Data Runs
Complete_Messy_Double_Electron_Run2015D.csv
Complete_Messy_Double_Muon_Run2015D.csv
Double_Electron_Run2015D.csv
Double_Muon_Run2015D.csv
2011 Data Runs
Double_Electron_Run2011A.csv
Double_Electron_Run2011A_with_Mass.csv
Double_Muon_Run2011A.csv… See the full description on the dataset page: https://huggingface.co/datasets/Pcapps/Quarknet-HEP.arxiv-cs-id-small
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
This dataset is based on CCRss/arxiv_papers_cs and machine translated from EN to ID using DeepL. I haven't translated the whole dataset yet, but I have uploaded a part of it that is already translated.
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/quarkss/arxiv-cs-id-small.bonitQuarksekInfoTASK1
