Standard

MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark. / Vyatkin, Artyom; Poptsov, Aleksander; Oliseenko, Valerii; Veiber, Evgeniya; Shutko, Oleg; Abramov, Maxim; Nikolenko, Sergey I.; Рафиков, Тимур Ришатович.

In: IEEE Access, Vol. 14, 2026.

Research output: Contribution to journalArticlepeer-review

Harvard

APA

Vancouver

Author

Vyatkin, Artyom ; Poptsov, Aleksander ; Oliseenko, Valerii ; Veiber, Evgeniya ; Shutko, Oleg ; Abramov, Maxim ; Nikolenko, Sergey I. ; Рафиков, Тимур Ришатович. / MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark. In: IEEE Access. 2026 ; Vol. 14.

BibTeX

@article{dec76529a3514319b5f1d32a4673cf1e,
title = "MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark",
abstract = "Reliable function calling (a.k.a. tool use) is a core capability of LLM agents. However, existing evaluations insufficiently probe how robustness varies with task complexity. We introduce MultiCAT-Bench, the first benchmark focused on detailed categorization of tasks for assessing tool utilization, along with an approach for its automated generation, which utilizes the GPT-5 family. MultiCAT-Bench spans ten principled categories of difficulty with 3 789 test cases. Using this dataset, we evaluate ten LLMs from 9 model families. The analysis revealed that Recall metrics (overall ~72.4%, tool name identification ~81%, arguments ~90.4%) are lower than Precision (overall ~88.3%, tool name identification ~99.8%, arguments ~93%) across the examined categories. This indicates that models are less likely to extract relevant information, but when they do, they achieve higher accuracy. Regarding the selected categories, the greatest impact on models{\textquoteright} accuracy was exerted by the number of calls in the response (average drop by a factor of 1.59), parameter optionality status (by a factor of 1.37), and the number of parameters in the function (by a factor of 1.22). The Grok 4.1 Fast and GPT-5 Mini models achieve the best average accuracy, 83.8% and 83.7%, respectively, across the benchmark.",
keywords = "Modeling, Tools, Generative Pre-trained transformer, Large language models, Measurement, Accuracy, Artificial intelligence, Noise, Educational institutions, Cognition, Agents, evaluation, large language models, tool use",
author = "Artyom Vyatkin and Aleksander Poptsov and Valerii Oliseenko and Evgeniya Veiber and Oleg Shutko and Maxim Abramov and Nikolenko, {Sergey I.} and Рафиков, {Тимур Ришатович}",
year = "2026",
doi = "10.1109/ACCESS.2026.3710376",
language = "English",
volume = "14",
journal = "IEEE Access",
issn = "2169-3536",
publisher = "Institute of Electrical and Electronics Engineers Inc.",

}

RIS

TY - JOUR

T1 - MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark

AU - Vyatkin, Artyom

AU - Poptsov, Aleksander

AU - Oliseenko, Valerii

AU - Veiber, Evgeniya

AU - Shutko, Oleg

AU - Abramov, Maxim

AU - Nikolenko, Sergey I.

AU - Рафиков, Тимур Ришатович

PY - 2026

Y1 - 2026

N2 - Reliable function calling (a.k.a. tool use) is a core capability of LLM agents. However, existing evaluations insufficiently probe how robustness varies with task complexity. We introduce MultiCAT-Bench, the first benchmark focused on detailed categorization of tasks for assessing tool utilization, along with an approach for its automated generation, which utilizes the GPT-5 family. MultiCAT-Bench spans ten principled categories of difficulty with 3 789 test cases. Using this dataset, we evaluate ten LLMs from 9 model families. The analysis revealed that Recall metrics (overall ~72.4%, tool name identification ~81%, arguments ~90.4%) are lower than Precision (overall ~88.3%, tool name identification ~99.8%, arguments ~93%) across the examined categories. This indicates that models are less likely to extract relevant information, but when they do, they achieve higher accuracy. Regarding the selected categories, the greatest impact on models’ accuracy was exerted by the number of calls in the response (average drop by a factor of 1.59), parameter optionality status (by a factor of 1.37), and the number of parameters in the function (by a factor of 1.22). The Grok 4.1 Fast and GPT-5 Mini models achieve the best average accuracy, 83.8% and 83.7%, respectively, across the benchmark.

AB - Reliable function calling (a.k.a. tool use) is a core capability of LLM agents. However, existing evaluations insufficiently probe how robustness varies with task complexity. We introduce MultiCAT-Bench, the first benchmark focused on detailed categorization of tasks for assessing tool utilization, along with an approach for its automated generation, which utilizes the GPT-5 family. MultiCAT-Bench spans ten principled categories of difficulty with 3 789 test cases. Using this dataset, we evaluate ten LLMs from 9 model families. The analysis revealed that Recall metrics (overall ~72.4%, tool name identification ~81%, arguments ~90.4%) are lower than Precision (overall ~88.3%, tool name identification ~99.8%, arguments ~93%) across the examined categories. This indicates that models are less likely to extract relevant information, but when they do, they achieve higher accuracy. Regarding the selected categories, the greatest impact on models’ accuracy was exerted by the number of calls in the response (average drop by a factor of 1.59), parameter optionality status (by a factor of 1.37), and the number of parameters in the function (by a factor of 1.22). The Grok 4.1 Fast and GPT-5 Mini models achieve the best average accuracy, 83.8% and 83.7%, respectively, across the benchmark.

KW - Modeling

KW - Tools

KW - Generative Pre-trained transformer

KW - Large language models

KW - Measurement

KW - Accuracy

KW - Artificial intelligence

KW - Noise

KW - Educational institutions

KW - Cognition

KW - Agents

KW - evaluation

KW - large language models

KW - tool use

U2 - 10.1109/ACCESS.2026.3710376

DO - 10.1109/ACCESS.2026.3710376

M3 - Article

VL - 14

JO - IEEE Access

JF - IEEE Access

SN - 2169-3536

ER -

ID: 160809397