datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nfcorpusfiqaIFIR
Dataset Card for IFIR Benchmark
Repository: sighingsnow/IFIR
For the usage of this dataset, please refer to the github repo.
If you find this repository helpful, feel free to cite our paper:
@misc{song2025ifir,
title={IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval},
author={Tingyu Song and Guo Gan and Mingsheng Shang and Yilun Zhao},
year={2025},
eprint={2503.04644},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/songtingyu/IFIR.ultrachat-200k-bliss-raw
UltraChat 200k Blissymbolic Raw Transliteration
This dataset is a lexical Blissymbolic transliteration of HuggingFaceH4/ultrachat_200k train_sft using experimental BlissyLM conversion tooling. It preserves the source conversation structure and role metadata while adding Blissymbol token sequences based on BCI Authorized Vocabulary gloss lookup.
This is not a human translation and is not clinical AAC guidance.
BlissyLM is an early research/tooling project for exploring Blissymbol… See the full description on the dataset page: https://huggingface.co/datasets/ifinspire/ultrachat-200k-bliss-raw.fireailapmcdsscifact_opentraining-pack-xln-of-my-project-ifivvo-528e5fe2
Training pack xln of my Project ifivvo
create 10 chair images
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the sync operator.… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/training-pack-xln-of-my-project-ifivvo-528e5fe2.ifixit-repair-chatmlifirtiny-aya-base-blind-spots
Tiny Aya Base — Blind Spots Dataset
Overview
This dataset documents blind spots identified in CohereLabs/tiny-aya-base, a multilingual base language model (3.35B parameters, 70+ languages). Each entry contains a prompt, the expected correct output, the model's actual output, and a human annotation of the error type.
The model scored 5/18 (28%) on our evaluation prompts.
Categories Tested
Multilingual (6 prompts, 2 correct): Yoruba, Igbo, Hausa translation… See the full description on the dataset page: https://huggingface.co/datasets/Ifihan/tiny-aya-base-blind-spots.if_inst_checked-llm-jp-3.1-8x13b-instruct4-v4-checkedHumanEval_infilling_gpt4_singlelineHumanEval_infilling_gpt4_multilineif_instruct_mistral_fhong-guan-jing-ji-shu-juzhuan-ti-bao-biao-shu-juwen-cai-zhi-neng-xuan-guji-chu-shu-juli-shi-hang-qing-shu-juifi-audioifi-audio-test
