datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-llm
PI-LLM Bench: The Core Retrieval Challenge Behind MRCR
Update: Accepted to COLM 2026 (San Francisco).
Moonshot AI (Kimi) PI-LLM is being observed internally for agent state tracking and robustness to context interference
ICML 2025 Long-Context Foundation Models Workshop Accepted.
AAAI 2026 Worshop Oral: LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems)
A simple context interference evaluation.
Update: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/pi-llm.LLMs-First-Task
Super easy task for humans that All SOTA LLM fail to retrieve the correct answer from context. Including SOTA models: GPT5, Grok4, DeepSeek, Gemini 2.5PRO, Mistral, Llama4...etc
Update: Accepted to COLM 2026 (San Francisco).
AAAI 2026 Worshop Oral: Jan/2026 LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole
ICML 2025 Long-Context Foundation Models Workshop Accepted.(https://arxiv.org/abs/2506.08184)
Update: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/LLMs-First-Task.coreference-challenge
PI-LLM Bench: The Core Retrieval Challenge Behind MRCR
Update: Accepted to COLM 2026 (San Francisco).
AAAI 2026 Worshop Oral: LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole
ICML 2025 Long-Context Foundation Models Workshop Accepted.
A simple context interference evaluation.
Update: This dataset is integrated into Moonshot AI(Kimi)'s internal benchmarking framework for assessing ** tracking capacity and… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/coreference-challenge.
