Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике материалов конференции › научная › Рецензирование
Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. / Muradyan, Denis S.; Shutko, Oleg A.; Bushemelev, Fedor V.
2026 XXIX International Conference on Soft Computing and Measurements (SCM). Institute of Electrical and Electronics Engineers Inc., 2026. стр. 231-233.Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике материалов конференции › научная › Рецензирование
}
TY - GEN
T1 - Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching
AU - Muradyan, Denis S.
AU - Shutko, Oleg A.
AU - Bushemelev, Fedor V.
N1 - Conference code: 29
PY - 2026/5
Y1 - 2026/5
N2 - Information published on Internet resources constitutes an important data source for the operation of many modern services and is used in data collection, analysis, monitoring, business intelligence, and other tasks. However, the absence of APIs, the high variability of websites, and frequent changes in site structure significantly complicate development and require continuous parser maintenance. In recent years, Large Language Models (LLMs) have demonstrated high effectiveness in text analysis and program code generation tasks, opening new possibilities for automating data analysis from web resources. This paper proposes an architecture for an adaptive web data extraction system that extends the capabilities of traditional web parsing methods through the integration and adaptation of LLMs. The solution supports two operating modes: generation of Python scripts for page parsing and structuring information from page text. To optimize LLM token costs for parser generation, a semantic caching mechanism is introduced, enabling the reuse of previously generated parsers based on the retrieval of similar queries in vector space. An experimental evaluation of the system was conducted using the LiveWeb-IE benchmark. The results showed that the proposed system extracts query-relevant data with 90% accuracy. Additional caching reduces the response time for a single query by more than an order of magnitude.
AB - Information published on Internet resources constitutes an important data source for the operation of many modern services and is used in data collection, analysis, monitoring, business intelligence, and other tasks. However, the absence of APIs, the high variability of websites, and frequent changes in site structure significantly complicate development and require continuous parser maintenance. In recent years, Large Language Models (LLMs) have demonstrated high effectiveness in text analysis and program code generation tasks, opening new possibilities for automating data analysis from web resources. This paper proposes an architecture for an adaptive web data extraction system that extends the capabilities of traditional web parsing methods through the integration and adaptation of LLMs. The solution supports two operating modes: generation of Python scripts for page parsing and structuring information from page text. To optimize LLM token costs for parser generation, a semantic caching mechanism is introduced, enabling the reuse of previously generated parsers based on the retrieval of similar queries in vector space. An experimental evaluation of the system was conducted using the LiveWeb-IE benchmark. The results showed that the proposed system extracts query-relevant data with 90% accuracy. Additional caching reduces the response time for a single query by more than an order of magnitude.
KW - Modeling
KW - Timing
KW - Large language models
KW - Measurement
KW - Printing
KW - Codes
KW - HTML
KW - Costing
KW - Costs
KW - large language models
KW - web parsing
KW - automated data analysis
KW - semantic caching
KW - code generation
U2 - 10.1109/SCM71573.2026.11625872
DO - 10.1109/SCM71573.2026.11625872
M3 - Conference contribution
SN - 979-8-3195-4899-3
SP - 231
EP - 233
BT - 2026 XXIX International Conference on Soft Computing and Measurements (SCM)
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 20 May 2026 through 22 May 2026
ER -
ID: 160342511