CoolFace
Modelpublic

metafloor-ai/manifest-orchestrator-4b-v0.6.0

sourceHugging Facecc-by-nc-4.0updated 1mo agoView on Hugging Face
0likes9downloads
Model Card

Manifest 4B · v0.6.0

The most capable Manifest model available today — deep reasoning for the hardest, multi-constraint supply-chain problems.

Manifest is MetaFloor's suite of supply-chain expert models — purpose-built specialists in procurement, demand planning, warehouse operations, supplier relationship management, risk & resilience, transportation, inventory, and order fulfilment.

  • —Family: Manifest · This model: Manifest 4B (4-billion-parameter base)
  • —Tier: advanced — the most capable model available today
### Preferred 95.1% of the time over the base model On 134 held-out expert questions, an independent LLM judge panel picked this model's answer over the untuned base model's answer 95.1% of the time (95% CI 91.0–98.5%). With both models given the same answer format, it is still preferred 97.0% of the time — the gain is real domain knowledge, not just presentation.

What's new in v0.6

  • —Retrained on MetaFloor's expanded ~32k-example supply-chain dataset (up from ~12.5k in v0.5).
  • —The evaluation benchmark grew to 134 held-out questions (from 116) — so v0.6 headline figures are measured on a larger, harder set than the v0.5 cards.
  • —A new [Manifest 9B](https://huggingface.co/metafloor-ai/manifest-orchestrator-9b-v0.6.0) tier is joining the family (coming soon); the 35B orchestrator is retired from this release.
  • —This is a new versioned repo — the previous v0.5 release (`metafloor-ai/manifest-orchestrator-4b`, unchanged) keeps its own history and download stats.

v0.5 → v0.6, same 134-question benchmark (preferred over the untuned base):

ReleasePreferred over basevs base + matched format
v0.5.088.8%90.9%
v0.6.095.1%97.0%

(The baseline throughout is the untuned `Qwen/Qwen3.5-4B` — every win-rate above is measured against it.)

Best for

The heavy end: complex, multi-constraint scenarios where the answer has to be right and reasoned.

  • —Multi-echelon strategy, network design, and cost-vs-service trade-offs
  • —Executive-level advisory over messy, real-world operating constraints
  • —Nuanced supplier, risk, and resilience decisions
  • —Any task where answer quality matters more than latency or cost

(Need speed or volume instead? See Manifest 2B and Manifest 0.8B below.)

See the difference

Same question. Base model vs Manifest 4B.

Ask (COO of a $5.8B enterprise, 176 nodes, mostly manual): "Explain the core concept of MEIO and why it's relevant given our manual, single-echelon approach." Base model → "Thinking Process: 1. Analyze the Request: Role: COO of a global enterprise. Company Size: $5,812M revenue. Network: 176 network nodes…" — narrates its own reasoning instead of answering. Manifest 4B → "MEIO is the optimization of inventory across multiple echelons of the supply chain at once, rather than optimizing each echelon separately…" — answers directly, like a practitioner, then ties it to the operation.

The Manifest family

Two kinds of models:

🧭 Orchestrators — general-purpose, handle any supply-chain area

ModelSizePreferred over baseStatus
Manifest 0.8B0.8B79.9%✅ available
Manifest 2B2B85.1%✅ available
Manifest 4B4B95.1%✅ available
Manifest 9B9B—🔜 coming soon

🎯 Domain-experts — specialized for a single area

ModelPreferred over baseStatus
Manifest Specialist · Risk & Resilience100% (20 / 0)✅ available
Manifest Specialist · Inventory Optimization80.0%✅ available
Manifest Specialist · Demand Planning82.5%✅ available

Orchestrators are scored on the general supply-chain benchmark; domain-experts on their focused domain benchmark (20 items each — treat those exact figures as directional).

Not just a model — a full stack

Manifest is backed by everything needed to build it and trust it:

  • —A purpose-built dataset — thousands of supply-chain instruction–response pairs spanning 8 sub-domains and every company scale, generated by a seed-driven operator-as-teacher pipeline.
  • —A reproducible training pipeline — documented LoRA fine-tuning.
  • —An independent benchmark — 134 held-out expert questions, scored blind by a panel of LLM judges.

We built the model, the data, and the evaluation.

How to use

Manifest 4B is a LoRA adapter (~101 MB), applied on top of its base model at load time.

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "Qwen/Qwen3.5-4B"  # base model — see "Built on" below
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
model = PeftModel.from_pretrained(model, "metafloor-ai/manifest-orchestrator-4b-v0.6.0")

SYSTEM = "You are a senior supply chain expert. Answer correctly and concisely."
user = (
    "I'm an inventory planner at a ~$8M small business: ~11k active SKUs, 4 suppliers, "
    "2 network nodes, ~164-day avg lead time. How should I set safety stock as I move off spreadsheets?"
)
msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": user}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_dict=True, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Prompt tip: Manifest is trained to condition on the scenario — include the asker's role and operating scale (revenue, SKUs, suppliers, nodes, lead time) in the message for the sharpest, most tailored answers.

How it was measured

134 held-out expert questions across 8 supply-chain areas. Each question is answered by Manifest and by the base model (given the same answer format); an independent two-model LLM judge panel then picks the better answer. Manifest 4B was preferred 97.0% of the time (95% CI 94.0–99.3%; 129 wins / 3 losses / 2 ties over 134). Benchmark: [supply-chain-eval](https://huggingface.co/datasets/metafloor-ai/supply-chain-eval).

Training details

MethodLoRA (PEFT 0.20.0), rank 16 / alpha 16 / dropout 0.05
Target modulesall attention + MLP projections
Trainable params21,233,664 (~0.87% of the 2.44B base)
Epochs3
Training examples~32,000
Final loss1.09 (from 2.28)

Training data: supply-chain instruction–response pairs from the seed-driven operator-as-teacher pipeline — a deterministic engine emits a unique seed per example (area, sub-area, persona, question type, realistic numeric scenario) and a strong teacher model writes the matching answer. The training data is drawn from MetaFloor's proprietary ~32k-example supply-chain dataset, which is not open-sourced — only the held-out evaluation benchmark (supply-chain-eval) is public.

Intended use & limitations

  • —Intended use: high-quality decision-support and drafting for supply-chain professionals.
  • —Out of scope: not legally binding, contractual, or safety-critical guidance; no access to your live systems or real-time data. Verify outputs before acting on them.
  • —Limitations: English-only; trained on synthetic (model-authored) data; standard LLM risks (hallucination, outdated facts) apply.

License

Manifest models and the supply-chain-eval benchmark are released under CC-BY-NC-4.0 — free for research and non-commercial use, with attribution. Commercial use requires a license from MetaFloor — get in touch at metafloor.ai.

Built on

Manifest 4B is a LoRA adapter over Qwen/Qwen3.5-4B (used under its own license); the base model is required to load the adapter.

Citation

bibtex
@misc{metafloor_manifest_4b,
  title  = {Manifest 4B: a supply-chain expert model (MetaFloor Manifest suite)},
  author = {MetaFloor AI},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/metafloor-ai/manifest-orchestrator-4b-v0.6.0}}
}

<!-- manifest-suite:order -->