Skip to main content
AgentNet Observer
Esc
    ← Back to Library

    GAIA: A Benchmark for General AI Assistants

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom

    Dataset 2023 Published 2023-11-21 NeurIPS 2023 (Datasets & Benchmarks) arXiv:2311.12983 Indexed 2026-09-06 benchmarkgeneral-assistantevaluation

    Abstract

    A benchmark of 466 questions requiring reasoning, multimodal handling, web browsing and tool use — questions simple for humans yet hard for state-of-the-art AI.

    Cite this entry

    GB/T 7714-2015

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, et al. GAIA: A Benchmark for General AI Assistants[EB/OL]. NeurIPS 2023 (Datasets & Benchmarks), 2023(2023-11-21)[2026-09-16]. https://arxiv.org/abs/2311.12983.

    BibTeX

    @misc{mialon2023,
      author = {Grégoire Mialon and Clémentine Fourrier and Craig Swift and Thomas Wolf and Yann LeCun and Thomas Scialom},
      title = {GAIA: A Benchmark for General AI Assistants},
      year = {2023},
      organization = {NeurIPS 2023 (Datasets & Benchmarks)},
      howpublished = {\url{https://arxiv.org/abs/2311.12983}},
    }

    Visit source ↗

    This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.