liquid-ai
LiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
nanobeir-multilingual-extended
NanoBEIR Multilingual Extended Dataset
This dataset extends the NanoBEIR multilingual collection with Japanese and Korean translations.
Dataset Structure
Each configuration follows the pattern <BASE>_<LANG> with splits:
corpus: Document corpus
queries: Search queries
qrels: Query relevance judgments (when available)
Languages
Arabic (ar), German (de), English (en), Spanish (es), French (fr)
Italian (it), Norwegian (no), Portuguese (pt), Swedish (sv)… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/nanobeir-multilingual-extended.AI2_Alphabot_2_pour_liquid
AI2_Alphabot_2_pour_liquid
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 49
Total Frames: 36748
FPS: 30
Dataset Size: 1.10 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb,
cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_pour_liquid.ifstruct-v1.0
IFStruct v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0
💻 GitHub: https://github.com/Liquid4All/ifstruct
IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.antidoom-mix-v1.0
Antidoom Mix v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/antidoom
💻 GitHub: https://github.com/Liquid4All/antidoom
Antidoom Mix v1.0 is a prompt-only training mixture for antidoom-style generation and preference-data pipelines. Responses are generated on this dataset, and looping traces are retained to construct preference pairs.
The dataset is intended to provide prompts only. Gold answers, rationales, hidden tests, verifier targets, and answer labels are… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0.NanoBEIR-ko
