Skip to main content

"Dynamic Web Scraping and API Management in Python"

Page 1

International Research Journal of Engineering and Technology (IRJET)

e-ISSN: 2395-0056

Volume: 12 Issue: 02 | Feb 2025

p-ISSN: 2395-0072

www.irjet.net

"Dynamic Web Scraping and API Management in Python" Shweta D1, Gayathri H2 1Assistant Professor, Sri Krishna Institute of Technology, Chikkabanavara, Bengaluru, Karnataka 2 Graduate Student, Sri Krishna Institute of Technology, Chikkabanavara, Bengaluru, Karnataka

--------------------------------------------------------------------***-----------------------------------------------------------------------

Abstract: Web scraping can be called that method where

detection of API breaking changes with respect to precision and recall and better backward compatibility.

structured data might be extracted from the unstructured web content. It works as an imperative tool concerning data analytics and making decisions in many different areas. This paper overviews all the possible methods to conduct web scraping through the help of Python along with a few of its most known libraries such as Beautiful Soup and Requests, keeping emphasis on e-commerce applications, The applications are describes about the data analysis. The discussion has also analyzed several difficulties and ethical concerns with its scope such as blocking through IPs and loading dynamic contents. In addition, the rapid evolution of packages in Python requires effective handling of backward compatibility. AexPy is a novel tool developed to detect the breaking changes systematically in the Python APIs. It finds 86.9% known breaking changes while uncovering hundreds of undocumented modifications. This illustrates that robust tools are in demand for both web data extraction and API change detection.

Apart from this, Python has also become a popular tool for web scraping, which is the automated extraction of valuable information from websites. The present paper explores several web scraping applications, such as the extraction of product information from e-commerce websites like Amazon and Flipkart for market research and price comparison, detection of illegal products and sellers, and the automation of data extraction from IMDb. Despite the challenges like security measures and ethical concerns, web scraping is an essential activity that transforms raw web data into meaningful findings for multiple industries. The study also addresses growing concerns within the Python ecosystem, including dependency management, security risks in the Python Package Index (PyPI), and enhancing repository security. Finally, it points out the role of web scraping in machine learning in supporting model training and introduces a social media web parser for digital forensic investigations and emphasizes Python's wide use in business, research, and security.

Additionally, the paper has presented dependency management in Python projects as relevant to the work, based on the notion that there is a development hindrance in library conflicts. Through an empirical study, evaluation of dependency issues in well-known Python libraries is carried out to achieve better management strategies. The danger of malicious packages in the Python Package Index was also discussed as a requirement for better malware detection mechanisms. These proposed solutions look towards enhancing the reliability and security of Python libraries but aid in efficient data extraction, which would help to alleviate critical challenges faced by developers and researchers in the ecosystem.

Literature Review Python has become the central tool for API management, web scraping, and other data-driven applications because of its flexibility and a very robust ecosystem. This review combines insights from some of the key studies that identify strengths, weaknesses, and improvement areas in these domains. API Breaking Changes in Python Packages AexPy: The integration of static and dynamic analysis provides a precise mechanism for identifying API-breaking changes in Python. A total of 43 packages were validated, whereby the comprehensiveness was checked along with practicality. However, its usage is restricted to other programming languages and further work lacks specificity. Nonetheless, AexPy is a valuable contribution toward reliability in software development and improved API compatibility management(Du & Ma, 2022).

Keywords— Web Scraping, Data Extraction, Python, BeautifulSoup, Data Extraction, Data Analysis. Abbreviations---- AexPy: API Change Detection Tool for Python, IP: Internet Protocol, PyPI: Python Package Index.

Introduction The versatility of Python in data processing, deep learning, and automation is due to its extensive library ecosystem. However, the dynamic nature of Python packages often causes breaking API changes that tend to disrupt code compatibility. AexPy is a tool proposed here to improve the

© 2025, IRJET

|

Impact Factor value: 8.315

Web Scraping Using Python There are versatile libraries in Python such as Selenium, BeautifulSoup, and Scrapy, which offer solutions for the issues related to dynamic content and ethical concerns of

|

ISO 9001:2008 Certified Journal

|

Page 687


Turn static files into dynamic content formats.

Create a flipbook
"Dynamic Web Scraping and API Management in Python" by IRJET Journal - Issuu