Standard

Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. / Muradyan, Denis S.; Shutko, Oleg A.; Bushemelev, Fedor V.

2026 XXIX International Conference on Soft Computing and Measurements (SCM). Institute of Electrical and Electronics Engineers Inc., 2026. стр. 231-233.

Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференцийстатья в сборнике материалов конференциинаучнаяРецензирование

Harvard

Muradyan, DS, Shutko, OA & Bushemelev, FV 2026, Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. в 2026 XXIX International Conference on Soft Computing and Measurements (SCM). Institute of Electrical and Electronics Engineers Inc., стр. 231-233, XXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026), Санкт-Петербург, Российская Федерация, 20/05/26. https://doi.org/10.1109/SCM71573.2026.11625872

APA

Muradyan, D. S., Shutko, O. A., & Bushemelev, F. V. (2026). Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. в 2026 XXIX International Conference on Soft Computing and Measurements (SCM) (стр. 231-233). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/SCM71573.2026.11625872

Vancouver

Muradyan DS, Shutko OA, Bushemelev FV. Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. в 2026 XXIX International Conference on Soft Computing and Measurements (SCM). Institute of Electrical and Electronics Engineers Inc. 2026. стр. 231-233 https://doi.org/10.1109/SCM71573.2026.11625872

Author

Muradyan, Denis S. ; Shutko, Oleg A. ; Bushemelev, Fedor V. / Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching. 2026 XXIX International Conference on Soft Computing and Measurements (SCM). Institute of Electrical and Electronics Engineers Inc., 2026. стр. 231-233

BibTeX

@inproceedings{68f7d8c155af494097ce48f95e00de1a,
title = "Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching",
abstract = "Information published on Internet resources constitutes an important data source for the operation of many modern services and is used in data collection, analysis, monitoring, business intelligence, and other tasks. However, the absence of APIs, the high variability of websites, and frequent changes in site structure significantly complicate development and require continuous parser maintenance. In recent years, Large Language Models (LLMs) have demonstrated high effectiveness in text analysis and program code generation tasks, opening new possibilities for automating data analysis from web resources. This paper proposes an architecture for an adaptive web data extraction system that extends the capabilities of traditional web parsing methods through the integration and adaptation of LLMs. The solution supports two operating modes: generation of Python scripts for page parsing and structuring information from page text. To optimize LLM token costs for parser generation, a semantic caching mechanism is introduced, enabling the reuse of previously generated parsers based on the retrieval of similar queries in vector space. An experimental evaluation of the system was conducted using the LiveWeb-IE benchmark. The results showed that the proposed system extracts query-relevant data with 90% accuracy. Additional caching reduces the response time for a single query by more than an order of magnitude.",
keywords = "Modeling, Timing, Large language models, Measurement, Printing, Codes, HTML, Costing, Costs, large language models, web parsing, automated data analysis, semantic caching, code generation",
author = "Muradyan, {Denis S.} and Shutko, {Oleg A.} and Bushemelev, {Fedor V.}",
year = "2026",
month = may,
doi = "10.1109/SCM71573.2026.11625872",
language = "English",
isbn = "979-8-3195-4899-3",
pages = "231--233",
booktitle = "2026 XXIX International Conference on Soft Computing and Measurements (SCM)",
publisher = "Institute of Electrical and Electronics Engineers Inc.",
address = "United States",
note = "null ; Conference date: 20-05-2026 Through 22-05-2026",
url = "https://scm.etu.ru/2026/ru/",

}

RIS

TY - GEN

T1 - Adaptive Web Resource Parsing System Based on Large Language Models Using Semantic Caching

AU - Muradyan, Denis S.

AU - Shutko, Oleg A.

AU - Bushemelev, Fedor V.

N1 - Conference code: 29

PY - 2026/5

Y1 - 2026/5

N2 - Information published on Internet resources constitutes an important data source for the operation of many modern services and is used in data collection, analysis, monitoring, business intelligence, and other tasks. However, the absence of APIs, the high variability of websites, and frequent changes in site structure significantly complicate development and require continuous parser maintenance. In recent years, Large Language Models (LLMs) have demonstrated high effectiveness in text analysis and program code generation tasks, opening new possibilities for automating data analysis from web resources. This paper proposes an architecture for an adaptive web data extraction system that extends the capabilities of traditional web parsing methods through the integration and adaptation of LLMs. The solution supports two operating modes: generation of Python scripts for page parsing and structuring information from page text. To optimize LLM token costs for parser generation, a semantic caching mechanism is introduced, enabling the reuse of previously generated parsers based on the retrieval of similar queries in vector space. An experimental evaluation of the system was conducted using the LiveWeb-IE benchmark. The results showed that the proposed system extracts query-relevant data with 90% accuracy. Additional caching reduces the response time for a single query by more than an order of magnitude.

AB - Information published on Internet resources constitutes an important data source for the operation of many modern services and is used in data collection, analysis, monitoring, business intelligence, and other tasks. However, the absence of APIs, the high variability of websites, and frequent changes in site structure significantly complicate development and require continuous parser maintenance. In recent years, Large Language Models (LLMs) have demonstrated high effectiveness in text analysis and program code generation tasks, opening new possibilities for automating data analysis from web resources. This paper proposes an architecture for an adaptive web data extraction system that extends the capabilities of traditional web parsing methods through the integration and adaptation of LLMs. The solution supports two operating modes: generation of Python scripts for page parsing and structuring information from page text. To optimize LLM token costs for parser generation, a semantic caching mechanism is introduced, enabling the reuse of previously generated parsers based on the retrieval of similar queries in vector space. An experimental evaluation of the system was conducted using the LiveWeb-IE benchmark. The results showed that the proposed system extracts query-relevant data with 90% accuracy. Additional caching reduces the response time for a single query by more than an order of magnitude.

KW - Modeling

KW - Timing

KW - Large language models

KW - Measurement

KW - Printing

KW - Codes

KW - HTML

KW - Costing

KW - Costs

KW - large language models

KW - web parsing

KW - automated data analysis

KW - semantic caching

KW - code generation

U2 - 10.1109/SCM71573.2026.11625872

DO - 10.1109/SCM71573.2026.11625872

M3 - Conference contribution

SN - 979-8-3195-4899-3

SP - 231

EP - 233

BT - 2026 XXIX International Conference on Soft Computing and Measurements (SCM)

PB - Institute of Electrical and Electronics Engineers Inc.

Y2 - 20 May 2026 through 22 May 2026

ER -

ID: 160342511