← 返回资料库
AgentBench:评估大模型作为智能体的能力
Xiao Liu, Botian Yu, Haotian Du, Zhiheng Zheng, et al.
数据集 2023 原文 2023-08-07 ICLR 2024 arXiv:2308.03688 收录 2026-09-06 benchmarkevaluation
摘要
系统性多维基准:在操作系统、数据库、知识图谱、数字卡牌游戏等 8 个环境中评估 LLM 作为智能体的实际能力。
引用本文条目
GB/T 7714-2015
Xiao Liu, Botian Yu, Haotian Du, et al. AgentBench: Evaluating LLMs as Agents[EB/OL]. ICLR 2024, 2023(2023-08-07)[2026-09-16]. https://arxiv.org/abs/2308.03688.
BibTeX
@misc{liu2023,
author = {Xiao Liu and Botian Yu and Haotian Du and Zhiheng Zheng and et al.},
title = {AgentBench: Evaluating LLMs as Agents},
year = {2023},
organization = {ICLR 2024},
howpublished = {\url{https://arxiv.org/abs/2308.03688}},
} 本条目为「智联观察」资料库收录,引用时请注明来源与本页链接。