microsoft/Fara1.5-27B
2901.9k
1---2license: mit3library_name: transformers4pipeline_tag: image-text-to-text5language:6- en7base_model:8- Qwen/Qwen3.5-27B9tags:10- computer-use11- cua12- web-agent13- multimodal14- vision-language15- agent16- browser-automation17- magentic18- fara19---20 21# Fara1.5-27B22 23[](https://aka.ms/fara1.5)24[](https://github.com/microsoft/fara)25[](https://huggingface.co/papers/2606.20785)26[](https://aka.ms/fara1.5-27B-foundry)27 28 29Fara1.5-27B is a multimodal **computer use agent (CUA)** for web browsers, from **Microsoft Research AI Frontiers**. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.30 31The model is vision-only at perception time: it sees the browser through screenshots, not the DOM or accessibility tree. Internal reasoning and trajectory history are tracked as text. Given the latest screenshot and prior actions, it predicts the next action with grounded arguments (e.g., pixel coordinates for a click).32 33Fara1.5-27B is supervised fine-tuned from Qwen3.5-27B on data generated by **FaraGen1.5**, our multi-agent pipeline that synthesizes web tasks, executes trajectories to solve them, and verifies the results before training.34 35Please use only with the harness developped in [Github/fara CLI](https://github.com/microsoft/fara) or [Magentic-Lite](https://github.com/microsoft/magentic-ui).36 37It's co-designed with **MagenticLite**, and that's the recommended deployment for both research and production.38 39## Highlights40 41- **End-to-end web task completion.** Fills forms, books reservations, applies for jobs, plans trips, runs shopping carts. Not just clicking around — sequencing actions toward a goal.42- **Vision-only perception.** Operates on screenshots alone, no DOM access required. Matches the input modality available to a human user.43- **Coordinate-grounded actions.** Predicts pixel-level click and drag targets directly. No separate grounding model needed.44- **Critical-points safety design.** Trained to stop and ask before personal info entry, payments, submissions, sign-ins, sending messages, or other irreversible actions — even if it could technically continue.45- **262K context.** Long enough for multi-screenshot trajectories with full action history.46 47## Model Details48 49| | |50|---|---|51| **Developer** | Microsoft Research AI Frontiers |52| **Architecture** | Multimodal decoder-only LM (image + text → text) |53| **Parameters** | 27B |54| **Context length** | 262,144 tokens |55| **Inputs** | User goal (text), current screenshot(s), prior agent thoughts and actions |56| **Outputs** | Chain-of-thought block followed by a tool-call block (XML-tagged) |57| **Training period** | January 2026 – April 2026 |58| **Training compute** | 64 × NVIDIA B200, 6 days |59| **Release date** | 21 May 2026 |60| **License** | MIT |61| **Base model** | Qwen3.5-27B |62 63## Recommended Deployment: MagenticLite64 65The safest way to run Fara1.5-27B is inside **MagenticLite**, which provides:66 67- **Sandboxing** — the browser runs in a Docker container with no access to host files or environment variables68- **Allow-lists** — restrict navigation to a user-specified set of domains69- **Watch-mode** — real-time monitoring of every action with full trace logs70- **Pause** — immediate halt of agent activity at any point71 72If you integrate Fara1.5-27B directly, you're responsible for these controls. Don't run this model with unrestricted browser access on a machine that has anything sensitive on it.73 74## Quickstart75 76### Requirements77 78- `torch >= 2.11.0`79- `transformers >= 5.2.0`80- `vllm >= 0.19.1`81- A set of GPUs with enough memory for a 27B model in bf16 (A6000, A100, H100, and B200 have been tested). We recommend sharding the model over at least 2 GPUs.82 83### Serve with vLLM84 85```bash86vllm serve microsoft/Fara1.5-27B \87 --dtype bfloat16 \88 --max-model-len 262144 \89 --limit-mm-per-prompt image=1090```91 92### System prompt93 94Fara1.5-27B is trained against a specific system prompt. Use it verbatim for best results:95 96```text97You are Fara, a computer use agent (CUA) specialized for web browsers. You are developed by Microsoft AI Frontiers. You assist users with completing and automating tasks that require the use of a web browser.98 99The model was trained in the timeframe of January - April 2026. You can effectively perform tasks even beyond this range by accessing the web browser and using the latest information on the live web. But your knowledge cutoff is limited to early 2026, so you may not be aware of events or developments that occurred after that time, without explicitly browsing and searching for latest information on the web.100 101This edition of the model was trained using SFT on top of Qwen3.5-27B, using a synthetic data mixture generated and developed by Microsoft AI Frontiers.102 103A critical point is a situation where we must pause and request information or confirmation from the user before proceeding. There are three types:104 105Case 1: Missing User Information — The task requires personal information that the user has not provided (e.g., email, phone number, address, payment details). Never fabricate or assume personal information. Fill in only what the user has explicitly provided, then pause and ask for any missing required fields.106 107Case 2: Underspecified Task — The task description is ambiguous or missing details needed to make a decision at the current step. Pause and ask for clarification.108 109Case 3: Irreversible Action — We are about to perform an action that cannot be undone (e.g., submitting a form, completing a purchase, sending a message, deleting data). If the user explicitly authorized the action, proceed. Otherwise, stop and ask for confirmation.110 111Only stop at a critical point if (1) required information is missing, (2) the task is ambiguous, OR (3) an irreversible action lacks explicit user authorization.112```113 114The full system prompt, including the complete `computer_use` tool schema, ships with the model in MagenticLite.115 116### Tool schema117 118Fara1.5-27B emits actions as `<tool_call>...</tool_call>` XML blocks containing a JSON object that calls the `computer_use` function. Supported actions:119 120| Action | Purpose |121|---|---|122| `left_click`, `right_click`, `double_click`, `triple_click` | Mouse clicks at `(x, y)` |123| `mouse_move`, `left_click_drag` | Cursor positioning and drag |124| `type`, `key` | Keyboard input |125| `scroll`, `hscroll` | Page scrolling |126| `visit_url`, `history_back` | Browser navigation |127| `web_search` | Search query |128| `pause_and_memorize_fact` | Persist a fact across the trajectory |129| `ask_user_question` | Surface a clarifying question to the user |130| `wait` | Sleep for N seconds |131| `terminate` | End the task with a final answer |132 133The screen resolution Fara is most commonly trained with is **1440×900**. Match this in your sandbox for the most reliable grounding.134 135### Minimal agent loop136 137Note: We strongly recommend to use only with the harness developped in [Github/fara](https://github.com/microsoft/fara) or [Magentic-Lite](https://github.com/microsoft/magentic-ui).138 139```python140# Pseudocode for a single-step interaction. Use vLLM's OpenAI-compatible141# endpoint and pass the screenshot as an image content part.142 143screenshot_b64 = capture_browser_screenshot() # base64 PNG144 145messages = [146 {"role": "system", "content": FARA_SYSTEM_PROMPT},147 {"role": "user", "content": [148 {"type": "text", "text": "Book a table for 2 at a sushi place in Sunnyvale for Friday 7pm."},149 {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot_b64}"}},150 ]},151]152 153response = client.chat.completions.create(154 model="microsoft/Fara1.5-27B",155 messages=messages,156 temperature=0.0,157 max_tokens=2048,158)159 160# Parse the assistant's reply: it will contain chain-of-thought text161# followed by <tool_call>{"name": "computer_use", "arguments": {...}}</tool_call>.162# Execute the action in your sandboxed browser, capture the new screenshot,163# append the assistant turn and a tool/observation turn, and loop until164# the model emits action="terminate". We only keep the most recent 3 screenshots165# in the chat history.166```167 168A reference implementation of the full agent loop, including screenshot capture and Docker sandboxing, is in MagenticLite.169 170## Critical Points: Safety by Design171 172Fara1.5-27B is trained to **pause and ask the user** at three types of critical points:173 1741. **Missing user information.** If the task needs personal data (name, email, phone, address, payment) the user hasn't provided, fill in what's available and stop to ask for the rest. Never fabricate.1752. **Underspecified tasks.** If the goal is ambiguous at the current decision point (e.g., "book me a flight" with no destination), stop and clarify.1763. **Irreversible actions.** Submitting forms, completing purchases, sending messages, deleting data — stop unless the user explicitly authorized the action upfront.177 178Concrete actions Fara is trained to stop on without explicit authorization:179 180- Entering personal information (name, email, phone)181- Entering credit card or shipping/billing details182- Completing purchases or bookings183- Making phone calls184- Sending emails185- Submitting job applications186- Signing into accounts187 188If the user grants authorization, the model proceeds. The principle is that costly, irreversible actions require a human in the loop.189 190## Evaluation191 192We evaluate Fara1.5-27B on end-to-end web agent benchmarks. For WebTailBench, we report outcome success.193 194| Model | [WebVoyager](https://github.com/MinorJerry/WebVoyager) | [Online-Mind2Web](https://huggingface.co/datasets/osunlp/Online-Mind2Web) | [WebTailBench](https://huggingface.co/datasets/microsoft/WebTailBench) |195| ----------- | -----------------------------------------------------: | ------------------------------------------------------------------------: | ---------------------------------------------------------------------: |196| Fara1.5-4B | 80.8 | 57.3 | 27.4 |197| Fara1.5-9B | 86.6 | 63.4 | 32.3 |198| Fara1.5-27B | 89.3 | 72.3 | 40.2 |199 200 201## Training202 203### Approach204 205Supervised Fine-Tuning on top of Qwen3.5-27B. Training mix is generated by FaraGen1.5, our multi-agent data pipeline, supplemented with curated public datasets for grounding and UI understanding.206 207### Data sources208 209- **Synthetic trajectories** — Tasks seeded from URLs sampled from a large web index and from open-source seed datasets (e.g., Mind2Web). A multi-agent system attempts each task while recording screenshots, thoughts, and actions. A verifier agent then keeps only successful trajectories.210- **Grounding** — Curated datasets for predicting actions and pixel coordinates from screenshots, with images, text, and bounding boxes.211- **UI understanding** — Visual question answering, captioning, and OCR over web page screenshots collected by the pipeline.212- **Safety and instruction following** — Refusal data covering harmful or unsafe tasks the model should decline.213 214### Scale215 216- Approximately 1 billion text tokens217- Less than 1 billion training images218- Latest data acquisition: March 20, 2026219- Training start: January 2026220 221The corpus is static. Future updates will ship as separately versioned models with their own model cards.222 223## Intended Use and Limitations224 225### Primary use cases226 227Automating repetitive web tasks: filling forms, shopping, booking travel, restaurant reservations, information seeking, account workflows. Fara1.5-27B can also serve as a grounding model for other agents that need pixel-accurate action prediction.228 229### Out of scope230 231- Languages other than English (training data is English-only)232- High-stakes domains (legal, health, financial advice) where inaccurate actions could cause harm233- Allocation decisions affecting legal status, housing, employment, or credit234- Unsandboxed deployments with access to sensitive accounts or files235- Commercial or real-world production use without additional testing and safeguards236 237### Known limitations238 239- **Vision-only perception** means the model can be misled by deceptive or low-quality page rendering, prompt injections embedded in page content, or visual ambiguity in UI elements240- **Multi-step trajectories accumulate error** — a misclick early in a sequence can compound241- **Run-to-run variance** on multi-turn tasks is non-trivial; benchmark numbers are averaged over multiple runs242- The model can hallucinate page state or misattribute information from earlier screenshots243 244## Responsible AI Considerations245 246Computer use is a powerful capability. An agent that can click, type, and submit in a real browser can also do those things wrong. We strongly recommend:247 248- **Human-in-the-loop monitoring** of Fara's actions on the live web with a fast way to halt execution249- **Sandboxed execution** — run Fara in an isolated container with no access to host files, environment variables, or sensitive credentials250- **Allow-listed or block-listed browsing** — restrict the agent's reachable surface to limit exposure to malicious pages251- **No credential or PII sharing** with the agent unless absolutely required and the user has authorized it252- **Output verification** — Fara can hallucinate, misattribute sources, or be misled by deceptive content; verify before acting on its outputs253 254### Safety evaluation255 256Fara1.5-27B was evaluated via automated red-teaming on Azure. Coverage included groundedness, jailbreak resistance, harmful content (hate, violence, self-harm), and copyright violations. Safety post-training includes both refusal data for malicious tasks and the critical-points framework described above.257 258### Known model risks259 260Like all language models, Fara1.5-27B can produce unfair, unreliable, or offensive outputs. Specific concerns:261 262- **Quality of service** varies across English dialects and is worse for non-English content263- **Representation harms** may persist despite safety post-training due to base-model biases264- **Information reliability** — generated content may be inaccurate or outdated265- **Prompt injection** — Fara reads page content, and adversarial content on a page may attempt to redirect its behavior266 267Developers deploying Fara1.5-27B should apply responsible AI best practices and ensure compliance with applicable laws and regulations. Using safety services like [Azure AI Content Safety](https://azure.microsoft.com/en-us/products/ai-services/ai-content-safety/) is recommended.268 269## License270 271Released under the **MIT License**.272 273## Contact274 275For information requests under the EU AI Act and related inquiries: **MSFTAIActRequest@microsoft.com**276 277Authorized representative: Microsoft Ireland Operations Limited, 70 Sir John Rogerson's Quay, Dublin 2, D02 R296, Ireland.