# Agent S: An Open Agentic Framework that Uses Computers Like a Human

**February 27, 2025**  
[Saaket Agashe*,](https://saa1605.github.io/) [Jiuzhou Han*,](https://jiuzhouh.github.io/) [Shuyu Gan,](https://shuyugan.github.io/) [Jiachen Yang,](https://sites.google.com/view/jiachen-yang/) [Ang Li,](https://angli.ai/) [Xin Eric Wang](https://eric-xw.github.io/)  
[Paper](https://arxiv.org/pdf/2410.08164) [Code](https://github.com/simular-ai/Agent-S) [Product](/content/sai/index.html)

Hey! A few months ago, I gave a talk at Princeton University on my thoughts about agents and Simular. Figured I should put together a summary and turned it into a blog post.

## State-of-the-Art Performance  
My first job was as a research scientist at Google DeepMind, where a key part of my role involved collaborating with various Google product teams to identify opportunities for applying our cutting-edge AI technology. However, one Googler asked me a totally unrelated question that may have ultimately sparked my decision to leave DeepMind and start Simular.

### Agent S Open Source First Launch on October 2024 - YouTube

[Agent S Open Source First Launch on October 2024](https://www.youtube.com/watch?v=U7Ah3k5nGyQ)  
  
Simular185 subscribers

## Agent S is a new agentic framework designed to enable computers to be used as intuitively as a human would
We introduce an Experience-Augmented Hierarchical Planning method. This method utilizes Online Web Knowledge for up-to-date information on frequently changing software and websites, along with Narrative Memory to leverage high-level experiences from past interactions. By breaking complex tasks into manageable subtasks and using Episodic Memory for step-by-step guidance, Agent S continuously refines its actions and learns from experience, achieving adaptable and effective task planning.

## Abstract
We present Agent S, an open agentic framework that enables autonomous interaction with computers through Graphical User Interface (GUI), aimed at transforming human-computer interaction by automating complex, multi-step tasks.
To this end, Agent S introduces experience-augmented hierarchical planning, which learns from external knowledge search and internal experience retrieval at multiple levels, facilitating efficient task planning and subtask execution.
In addition, it employs an Agent-Computer Interface to better elicit the reasoning and control capabilities of GUI agents based on Multimodal Large Language Models. Evaluation on the OSWorld benchmark shows that Agent S outperforms the baseline by 9.37% on success rate (an 83.6% relative improvement) and achieves a new state-of-the-art. Comprehensive analysis highlights the effectiveness of individual components and provides insights for future improvements.

Agent S addresses three key challenges in automating computer tasks:

## Task Instruction
Help me to remove the account “anonym-x2024@outlook.com”

1. **Open Account Settings:**  
   `agent.click (41,1, “left”)`  
2. **Switch Applications:**  
   `agent.switch_applications (“Thunderbird”)`  
3. **Open Account Settings:**  
   `agent.click (95,1, “left”)`  
4. **Remove the Account:**  
   `agent.click (86, 1, “left”)`  
5. **Remove the Account:**  
   `agent.click (93, 1, “left”)`  
6. **Remove the Account:**  
   `agent.click (149, 1, “left”)`

## Pipeline of Memory Construction and Update
The pipeline of memory construction and update, which contains two phases: Self-supervised Exploration and Continual Memory Update. The initial Narrative & Episodic Memory is constructed through some randomly curated tasks during the exploration phase, and then it is updated based on the inference tasks continually.

## Main Result
This table shows the performance comparison between Agent S and the baseline models, evaluated across the whole OSWorld test set. For the GPT-4o model, Agent S achieves an overall success rate of 20.58%, nearly doubling the performance of the best corresponding baseline (GPT-4o with 11.21%).

Agent S consistently outperforms the baselines in the “Daily” and “Professional” tasks, where it reaches 27.06% and 36.73% success rates, respectively, compared to the best baseline results of 12.33% and 14.29%. These tasks are commonly used in daily life or involved with knowledge-intensive professional applications, which benefit more from the retrieval augmentation of Agent S.

## Analysis
To demonstrate the effectiveness of individual modules of Agent S, we stratified sampled a subset of 65 instances from the full test set for the ablation study.

### Learning from experience enhances the domain knowledge of GUI agents
Learning from universal experience available as web knowledge allows Agent S to make informed plans across a wide range of tasks and has the most significant impact. The results demonstrate that each component plays a critical role in enhancing the agent’s domain knowledge. Removing all three components (w/o All) degrades the performance significantly, revealing the importance of learning from experience in the design.

### ACI elicits better reasoning abilities of LLMs and supports better agentic learning
Comparing the baseline with Agent S (ACI-only) highlights the enhanced reasoning abilities achieved by incorporating ACI.

### Hierarchical Planning supports long-horizon workflows
The observed performance drop underscores the importance of Hierarchical Planning in modeling long-horizon workflows.

### Exploration, Continual Memory Update and Self-Evaluator are indispensable for memory construction
The results reveal that ablating both the continual memory update and self-supervised exploration phases results in a performance drop.

## Generalization to Different Operating Systems
We test the Agent S framework with no modification on WindowsAgentArena, comparing Agent S with the similar configuration with GPT-4o as the MLLM backbone.

## BibTex
@misc{AgentS,
  title={Agent S: An Open Agentic Framework that Uses Computers Like a Human},
  author={Saaket Agashe*, Jiuzhou Han*, Shuyu Gan, Jiachen Yang, Ang Li, Xin Eric Wang},
  year={2024},
  archivePrefix={arXiv},
  primaryClass={cs.AI}
}

## Understanding the AI Agentic Framework
The AI agentic framework combines artificial intelligence (AI) with agent-based modeling for improved decision-making processes.

### Key Concepts of the Agentic Framework
- **Agent-Based Framework:** Individual agents work together, boosting efficiency.
- **Agentic Approach:** Agents act independently, highlighting their ability to learn and adapt.
- **Workflows:** Planned paths that agents follow to enhance processes.

### Applications of AI Agentic Framework
- **AI Framework Variations:** Adjusted to meet specific industry needs.
- **AI Solutions:** From virtual assistants to intricate management systems.

## Benefits of Using an Agentic Framework
- **Efficiency:** Increases productivity by reducing manual work.
- **Quality Management:** Ensures consistent quality through structured processes.

## Challenges in Implementing Agentic Frameworks
- **Data Privacy:** Protecting sensitive data is critical.
- **AI Governance:** Setting regulations for proper use and oversight.

## Conclusion
The AI agentic framework is instrumental in utilizing intelligent systems effectively, fostering innovation, and enhancing efficiency.
