Skip to main content
AgentNet Observer
Esc
    ← Back to Library

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangzhou Lei, Yitao Liu, Yiheng Xu, Ting-Wei Kuo, Tao Yu, Pan Lu, Lei Li, Ying Sheng, Bo Li, Yuke Zhu, Percy Liang, Didi Zhu, Zhiyong Wu, Zhilin Yang

    Dataset 2024 Published 2024-04-11 NeurIPS 2024 arXiv:2404.07972 Indexed 2026-09-06 benchmarkgui-agentmultimodal

    Abstract

    A benchmark of 369 real tasks across real operating systems (Ubuntu/Windows/macOS), requiring agents to control GUIs like a human to complete open-ended computer tasks.

    Cite this entry

    GB/T 7714-2015

    Tianbao Xie, Danyang Zhang, Jixuan Chen, et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments[EB/OL]. NeurIPS 2024, 2024(2024-04-11)[2026-09-16]. https://arxiv.org/abs/2404.07972.

    BibTeX

    @misc{xie2024,
      author = {Tianbao Xie and Danyang Zhang and Jixuan Chen and Xiaochuan Li and Siheng Zhao and Ruisheng Cao and Toh Jing Hua and Zhoujun Cheng and Dongchan Shin and Fangzhou Lei and Yitao Liu and Yiheng Xu and Ting-Wei Kuo and Tao Yu and Pan Lu and Lei Li and Ying Sheng and Bo Li and Yuke Zhu and Percy Liang and Didi Zhu and Zhiyong Wu and Zhilin Yang},
      title = {OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments},
      year = {2024},
      organization = {NeurIPS 2024},
      howpublished = {\url{https://arxiv.org/abs/2404.07972}},
    }

    Visit source ↗

    This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.