datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgemma-3-taide-12b-chat-eval-logs-and-scoresNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresministral-14b-eval-logs-and-scoresgemma-3-4b-it-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresLlama-Breeze2-8B-Instruct-eval-logs-and-scoresnemotron-nano-eval-logs-and-scoresgemma-3-4B-T1-it-eval-logs-and-scoresdevstral-eval-logs-and-scoresllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.3-70B-Instruct-eval-logs-and-scoresGemma-3-12b-it-eval-logs-and-scoresphi-4-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresDevstral-Small-2505-eval-logs-and-scoresgemma-3-27b-it-eval-logs-and-scoresmistral-675b-eval-logs-and-scorescyber_twist_the_tube_v0.1
CyberOrigin Dataset
Our data includes information from home services, the logistics industry, and laboratory scenarios.
For more details, please refer to our Offical Data Website
contents of dataset:
cyber_twist_the_tube # dataset root path
└── data/
├── metadata_ID1_240808.json
├── segment_ids_ID1_240808.bin # for each frame segment_ids uniquely points to the segment index that frame i came from. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/cyberorigin/cyber_twist_the_tube_v0.1.fool-me-twicehttps://github.com/google-research/fool-me-twice
@inproceedings{eisenschlos-etal-2021-fool,
title = "Fool Me Twice: Entailment from {W}ikipedia Gamification",
author = {Eisenschlos, Julian Martin and
Dhingra, Bhuwan and
Bulian, Jannis and
B{\"o}rschinger, Benjamin and
Boyd-Graber, Jordan},
booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/fool-me-twice.bibletts-asante-twi-repaired
BibleTTS Asante Twi — Repaired Transcripts
The Asante Twi transcripts released with BibleTTS have had the
characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them.
Audio is not included. This is a drop-in replacement for the .txt files that ship with the
BibleTTS Asante Twi package, matched by clip ID.
The problem
Both are Twi vowels, and both are required by the orthography. Measured across the released
Asante Twi transcripts:
Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.twi-english-citizens-budget
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi–English Citizens' Budget Parallel Corpus
317 Asante Twi ↔ English parallel sentences in the economy / public-finance
domain, extracted and aligned from the Government of Ghana
Citizens' Budget
publications (2022 and 2023… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-english-citizens-budget.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.twin-uniref50-faiss
Twin-Model UniRef50 FAISS Index
FAISS index over Twin model mean-pooled embeddings of all UniRef50
representative proteins (~49.8M). The Twin model is a two-tower contrastive
encoder fine-tuned on Resnik GO similarity:
Custom tower: AA-vocab Transformer → padding-masked mean pool → MLP → 512-dim
ESM tower: facebook/esm2_t33_650M_UR50D (frozen) → masked mean pool → MLP → 512-dim
Output: concat(custom, esm) → 1024-dim
Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/twin-uniref50-faiss.SLAKE
Dataset Info:
SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering [ISBI 2021 oral]
Project Page: click
Corresponding Authors: Bo Liu, Xiao-Ming Wu
Any questions, please contact us. Thank you!
Modification:
In the Huggingface Repo, we have changed the name of validate.json to validation.json to better display in the Dataset Card.
Famous-Keyword-Twitter-RepliesThe "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies. This dataset contains five essential columns that provide valuable insights into the Twitter conversation dynamics:
Keyword: This column represents the specific keyword or topic of interest that generated the original tweet. It helps identify the context or subject matter around which the conversation revolves.
Main_tweet: The main_tweet… See the full description on the dataset page: https://huggingface.co/datasets/jacksoncsie/Famous-Keyword-Twitter-Replies.DoppelReflEx__MN-12B-Unleashed-Twilight-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Unleashed-Twilight
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Unleashed-Twilight
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Unleashed-Twilight-details.
