liquidai
Datasets
All datasets matching “liquidai”nanobeir-multilingual-extended
NanoBEIR Multilingual Extended Dataset
This dataset extends the NanoBEIR multilingual collection with Japanese and Korean translations.
Dataset Structure
Each configuration follows the pattern <BASE>_<LANG> with splits:
corpus: Document corpus
queries: Search queries
qrels: Query relevance judgments (when available)
Languages
Arabic (ar), German (de), English (en), Spanish (es), French (fr)
Italian (it), Norwegian (no), Portuguese (pt), Swedish (sv)… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/nanobeir-multilingual-extended.LiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
ifstruct-v1.0
IFStruct v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0
💻 GitHub: https://github.com/Liquid4All/ifstruct
IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.antidoom-mix-v1.0
Antidoom Mix v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/antidoom
💻 GitHub: https://github.com/Liquid4All/antidoom
Antidoom Mix v1.0 is a prompt-only training mixture for antidoom-style generation and preference-data pipelines. Responses are generated on this dataset, and looping traces are retained to construct preference pairs.
The dataset is intended to provide prompts only. Gold answers, rationales, hidden tests, verifier targets, and answer labels are… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0.NanoBEIR-koLiquidAI-Hackathon-Tokyo-SFT-Data
LiquidAI-Hackathon-Tokyo-SFT-Data
Liquid AI Hackathon Tokyoで作成したモデルのSFTに利用したデータセットです。
