MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark. / Vyatkin, Artyom; Poptsov, Aleksander; Oliseenko, Valerii; Veiber, Evgeniya; Shutko, Oleg; Abramov, Maxim; Nikolenko, Sergey I.; Рафиков, Тимур Ришатович.
In: IEEE Access, Vol. 14, 2026.Research output: Contribution to journal › Article › peer-review
}
TY - JOUR
T1 - MultiCAT-Bench: A Multi-Categorical Agent Tool Use Benchmark
AU - Vyatkin, Artyom
AU - Poptsov, Aleksander
AU - Oliseenko, Valerii
AU - Veiber, Evgeniya
AU - Shutko, Oleg
AU - Abramov, Maxim
AU - Nikolenko, Sergey I.
AU - Рафиков, Тимур Ришатович
PY - 2026
Y1 - 2026
N2 - Reliable function calling (a.k.a. tool use) is a core capability of LLM agents. However, existing evaluations insufficiently probe how robustness varies with task complexity. We introduce MultiCAT-Bench, the first benchmark focused on detailed categorization of tasks for assessing tool utilization, along with an approach for its automated generation, which utilizes the GPT-5 family. MultiCAT-Bench spans ten principled categories of difficulty with 3 789 test cases. Using this dataset, we evaluate ten LLMs from 9 model families. The analysis revealed that Recall metrics (overall ~72.4%, tool name identification ~81%, arguments ~90.4%) are lower than Precision (overall ~88.3%, tool name identification ~99.8%, arguments ~93%) across the examined categories. This indicates that models are less likely to extract relevant information, but when they do, they achieve higher accuracy. Regarding the selected categories, the greatest impact on models’ accuracy was exerted by the number of calls in the response (average drop by a factor of 1.59), parameter optionality status (by a factor of 1.37), and the number of parameters in the function (by a factor of 1.22). The Grok 4.1 Fast and GPT-5 Mini models achieve the best average accuracy, 83.8% and 83.7%, respectively, across the benchmark.
AB - Reliable function calling (a.k.a. tool use) is a core capability of LLM agents. However, existing evaluations insufficiently probe how robustness varies with task complexity. We introduce MultiCAT-Bench, the first benchmark focused on detailed categorization of tasks for assessing tool utilization, along with an approach for its automated generation, which utilizes the GPT-5 family. MultiCAT-Bench spans ten principled categories of difficulty with 3 789 test cases. Using this dataset, we evaluate ten LLMs from 9 model families. The analysis revealed that Recall metrics (overall ~72.4%, tool name identification ~81%, arguments ~90.4%) are lower than Precision (overall ~88.3%, tool name identification ~99.8%, arguments ~93%) across the examined categories. This indicates that models are less likely to extract relevant information, but when they do, they achieve higher accuracy. Regarding the selected categories, the greatest impact on models’ accuracy was exerted by the number of calls in the response (average drop by a factor of 1.59), parameter optionality status (by a factor of 1.37), and the number of parameters in the function (by a factor of 1.22). The Grok 4.1 Fast and GPT-5 Mini models achieve the best average accuracy, 83.8% and 83.7%, respectively, across the benchmark.
KW - Modeling
KW - Tools
KW - Generative Pre-trained transformer
KW - Large language models
KW - Measurement
KW - Accuracy
KW - Artificial intelligence
KW - Noise
KW - Educational institutions
KW - Cognition
KW - Agents
KW - evaluation
KW - large language models
KW - tool use
U2 - 10.1109/ACCESS.2026.3710376
DO - 10.1109/ACCESS.2026.3710376
M3 - Article
VL - 14
JO - IEEE Access
JF - IEEE Access
SN - 2169-3536
ER -
ID: 160809397