Skip to main content

Get and storing the HTML data of a website using a cache system.

Project description

Version Static Badge Static Badge Static Badge Static Badge Framework


Logo

HTTPPOOL

Get and storing the HTML data of a website using a cache system

Explore the docs »

Report Bug · Request Feature

About The Project

This library is developed using pure Python code, and is responsible for storing locally, through a cache system using a Producer-Consumer algorithm, the HTML information obtained from any web page. It includes a daemon process that is responsible for automatically updating the cache through a set time, independent for each URL and also includes a tracing system using a logger to monitor the status of the service.

Getting Started

Prerequisites

You need to make sure you have installed the following modules.

pip install requests

If you do not have the library shown above installed, the httppool package will add it for you during its installation.

Installation

pip install Httppool

Usage

Example 1

It will show you and store all the information of the web's HTML

import Httppool as htp

url = 'https://www.prensa-latina.cu/deportes/'
content = htp.get_url_content(url)

print(content)

Example 2

If you use it in conjunction with a Web Scraping library like BeautifulSoup you can get the data you are interested in

import Httppool as htp
from bs4 import BeautifulSoup

url = 'https://www.entumovil.cu/'

content = htp.get_url_content(url)
soup = BeautifulSoup(content, 'html.parser')
response = soup.find('div', id='etm-desription')

print(response.prettify())

# the content of response is:

# <div id="etm-desription">
#     <p>
#         Somos un equipo de desarrollo de servicios y aplicaciones para móviles
#         pertenecientes a la empresa Desoft. Desde sus
#         inicios trabajamos orientados a la satisfacción de las necesidades que
#         soliciten las empresas de informatizar sus procesos, así como brindar a
#         la población la posibilidad de obtener información de su interés a
#         través de consultas, suscripciones, notificaciones y compras on line  mediante mensajería de texto (SMS)…
#         <br/>
#     </p>
#     <p>
#         <br/>
#     </p>
# </div>

Example 3

You can run the daemon process for automatic cache update

from Httppool.deamon import HttppoolDeamon

deamon = HttppoolDeamon()
deamon.start()

# It will update the information of each URL 
# that is in the cache folder (httppool-cache) 
# every 15 min by default. This parameter can be changed

Cache Structure

When using the library for the first time, 3 files are created for each URL in the httppool-cache folder and httppool-logger.log file. The following structure is created in the root directory of the current project:

My_python_project
├── httppool-cache
   ├── f47eda8cfe42cc7aa478686b321f22c6.lastaccess
   ├── f47eda8cfe42cc7aa478686b321f22c6.url
   └── f47eda8cfe42cc7aa478686b321f22c6.webpagedata
├── httppool-logger.log
├── my_python_script.py

.lastaccess

The file f47eda8cfe42cc7aa478686b321f22c6.lastaccess shows the following information:

read: 2025-01-06 11:05:04   # last date of reading

write: 2025-01-06 11:05:04  # last date of writing

retry: 15  # time interval taken to update the cache

.url

The file f47eda8cfe42cc7aa478686b321f22c6.url shows the following information:

# web page URL
https://www.entumovil.cu/  

# headers to be sent using the requests library
headers: {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'}

.webpagedata

The file f47eda8cfe42cc7aa478686b321f22c6.webpagedata shows the following information:

<!-- web page HTML information for future Scraping -->
<!DOCTYPE html>
<html lang="en">
    <head>
        <meta charset="UTF-8">
        <title>Title</title>
    </head>
    
    <body>
        <div class="item " data-ref="empresariales">
            <a class="service-link" href="/web/servicios/empresariales/#service_item43s">
                <div class="service-button-label">
                    <p>Micropagos</p>
                </div>
                <div class="service-button" style="background-image:url(/media/iconMICROPAGO.png) ">
                </div>
            </a>
        </div>
    </body>
</html>

httppool-logger.log

The library has a logger to track the traces of the cache system based on the producer-consumer algorithm:

2025-01-06 11:05:02,571 : INFO : HTTPPOOL : Client : Consumming URL --> https://www.entumovil.cu/ 

2025-01-06 11:05:02,572 : DEBUG : HTTPPOOL : Starting new HTTPS connection (1): www.entumovil.cu:443 

2025-01-06 11:05:03,610 : DEBUG : HTTPPOOL : https://www.entumovil.cu:443 "GET / HTTP/1.1" 200 None 

2025-01-06 11:05:04,055 : INFO : HTTPPOOL : Producer : Got URL content --> https://www.entumovil.cu/ 

License

Distributed under the MIT License. See LICENSE for more information.

Contact

Email: alexdev.workenv@gmail.com

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

httppool-1.0.18.tar.gz (10.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

httppool-1.0.18-py3-none-any.whl (7.9 kB view details)

Uploaded Python 3

File details

Details for the file httppool-1.0.18.tar.gz.

File metadata

  • Download URL: httppool-1.0.18.tar.gz
  • Upload date:
  • Size: 10.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for httppool-1.0.18.tar.gz
Algorithm Hash digest
SHA256 0bfdbdb9ec557f1dc700d259cbeb7b9b0685eb8baf317b21c534da7cb494bf90
MD5 28d0814c598791e19adc9aef25f888cb
BLAKE2b-256 a6eb67bc806a4b3628bb3837b8af9c62abde49e7e7f138d8288879537f77f9bc

See more details on using hashes here.

Provenance

The following attestation bundles were made for httppool-1.0.18.tar.gz:

Publisher: publish-httppool.yml on alexdevzz/httppool-ce-services

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file httppool-1.0.18-py3-none-any.whl.

File metadata

  • Download URL: httppool-1.0.18-py3-none-any.whl
  • Upload date:
  • Size: 7.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for httppool-1.0.18-py3-none-any.whl
Algorithm Hash digest
SHA256 7fe2e5ad944cfdeb2f70aaa6f6b11e50b821a1951cabfb6b8bb81da75a899061
MD5 8118b67b618bc57b1f2716f184edfdc9
BLAKE2b-256 ef196fb1c65ed21ddb5ac87de81a7111f72b7c45eecb7ab5303ae96b4ff1bb2c

See more details on using hashes here.

Provenance

The following attestation bundles were made for httppool-1.0.18-py3-none-any.whl:

Publisher: publish-httppool.yml on alexdevzz/httppool-ce-services

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page