datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
labbench2-fixed
LABBench2 PMID-enriched public mirror
This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged.
Two columns are added to every configuration:
pmids: deduplicated PubMed identifiers resolved for the row's sources.
source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null.
LABBench2
LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.Seamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
hermes3-en-fixed
Dataset Card for Hermes 3 Fixed Conversations
Dataset Description
Dataset Summary
hermes3-en-fixed is a [NousResearch/Hermes-3-Dataset]. During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
English (en)
Dataset Structure
Data Fields
conversations: list of messages in a dialog (array of objects)
from: normalized sender role — user or assistant… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-en-fixed.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
JMTEB-fixed
JMTEB-fixed
このリポジトリは元のJMTEBデータセットのUTF-8エンコーディングエラーを修正したバージョンです。
修正内容
問題
NLP Journal LaTeXコーパスの一部ファイル(例: V28N02-25.tex)が異なるエンコーディング(Shift-JIS、EUC-JPなど)で保存されており、UTF-8デコーディングエラーが発生していました。
解決策
retrieval.pyのNLPJournalHelper.load_txtメソッドを修正し、複数のエンコーディングを順次試行するようにしました。
使用方法
from datasets import load_dataset
dataset = load_dataset(
"ks-pf/JMTEB-fixed",
name="nlp_journal_title_abs-corpus",
trust_remote_code=True
)
# JMTEB: Japanese… See the full description on the dataset page: https://huggingface.co/datasets/ks-pf/JMTEB-fixed.capture24-ts-haystack-fixed-needle
Capture24 TS-Haystack — Fixed Needle Length
Long-context retrieval / reasoning benchmark over Capture24 wrist-worn
accelerometer recordings, used in Recursive Agents are Effective Time Series
Reasoners (ARTS-RLM).
This repository supersedes
nz00shuuuu/capture24-ts-haystack-cot
for the paper's main capture24 experiments. Differences:
Fixed (absolute-ms) needle length of 3–10 s across every context length
instead of needles that scale with context. With a 7200 s haystack the
needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.fixed_indonlu
Dataset Card for IndoNLU
Dataset Summary
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia (Indonesian language).
There are 12 datasets in IndoNLU benchmark for Indonesian natural language understanding.
EmoT: An emotion classification dataset collected from the social media platform Twitter. The dataset consists of around 4000 Indonesian colloquial language… See the full description on the dataset page: https://huggingface.co/datasets/will702/fixed_indonlu.TRQA-fixed
TRQA (fixed configs)
Private convenience mirror of GENTEL-Lab/TRQA, with each
schema exposed as a separate Hugging Face dataset configuration so that Dataset Viewer and load_dataset work.
The CSV contents are unchanged from source revision
c712c7948c907dec61beada11951cf997d89bae4.
Config
Rows
Columns
Default
lit-choice
172
Question, Options, Answer
Yes
lit-short
1,108
Question, Answer
No
db
641
Question, Answer
No
from datasets import load_dataset
choice =… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/TRQA-fixed.icd-11-qa-fixedA fixed version from the original Lamini ICD-11 QA Dataset
arc_challenge_de_fixed
ARC Challenge DE fixed
This is a copy of the dataset LeoLM/ArcChallenge_de which is a German translation of the original allenai/ai2_arc.
Bug fixesWe copied it to fix the numeric labels 1,2,3,4 into A,B,C,D.
ThanksThe translation of the arc_challenge was done by Bjoern: https://github.com/bjoernpl/GermanBenchmark
LLaVA-Mix665k-Fixed-Bug
LLaVA Visual Instruct 150K Dataset Card
Dataset details
Dataset type:
This is a fixed version of llava_v1_5_mix665k from LLaVA .
Paper or resources for more information:
https://llava-vl.github.io/
License:
Creative Commons Attribution 4.0 International; and it should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Where to send questions or comments about the model:
https://github.com/haotian-liu/LLaVA/issues
Intended use
Primary… See the full description on the dataset page: https://huggingface.co/datasets/JungleGym/LLaVA-Mix665k-Fixed-Bug.
