datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nanobeir-multilingual-extended
NanoBEIR Multilingual Extended Dataset
This dataset extends the NanoBEIR multilingual collection with Japanese and Korean translations.
Dataset Structure
Each configuration follows the pattern <BASE>_<LANG> with splits:
corpus: Document corpus
queries: Search queries
qrels: Query relevance judgments (when available)
Languages
Arabic (ar), German (de), English (en), Spanish (es), French (fr)
Italian (it), Norwegian (no), Portuguese (pt), Swedish (sv)… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/nanobeir-multilingual-extended.LiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
ifstruct-v1.0
IFStruct v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0
💻 GitHub: https://github.com/Liquid4All/ifstruct
IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.antidoom-mix-v1.0
Antidoom Mix v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/antidoom
💻 GitHub: https://github.com/Liquid4All/antidoom
Antidoom Mix v1.0 is a prompt-only training mixture for antidoom-style generation and preference-data pipelines. Responses are generated on this dataset, and looping traces are retained to construct preference pairs.
The dataset is intended to provide prompts only. Gold answers, rationales, hidden tests, verifier targets, and answer labels are… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0.NanoBEIR-koLiquidAI-Hackathon-Tokyo-SFT-Data
LiquidAI-Hackathon-Tokyo-SFT-Data
Liquid AI Hackathon Tokyoで作成したモデルのSFTに利用したデータセットです。
NanoBEIR-jaantidoom-mix-v1.0-LFM2.5-1.2B-Base
Antidoom Mix v1.0 – FTPO Pairs for LFM2.5-1.2B-Base
[!Note]
📓 Tutorial: This is the dataset used in Antidoom Training with TRL
Pre-extracted FTPO (Final Token Preference Optimization) training pairs for eliminating doom loops in LiquidAI/LFM2.5-1.2B-Base.
Derived from LiquidAI/antidoom-mix-v1.0 by generating completions with LFM2.5-1.2B-Base.
Schema
Each row contains:
Field
Type
Description
full_prompt
string
The formatted prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0-LFM2.5-1.2B-Base.Liquid_AI
LFM2-2.6B Blind Spots Dataset
Overview
This dataset contains 10 diverse input–output pairs where the model
LiquidAI/LFM2-2.6B
produces incorrect or significantly degraded responses compared to the expected
ground-truth answer. It is intended to support error analysis, fine-tuning
research, and benchmark construction for small hybrid LLMs designed for edge
deployment.
Model Tested
Field
Value
Model
LiquidAI/LFM2-2.6B
Architecture
Hybrid (LIV… See the full description on the dataset page: https://huggingface.co/datasets/Fareeha20Amir/Liquid_AI.liquidAI-blindspots
