Ecoindex_scraper module provides a way to scrape data from given website while simulating a real web browser
Project description
ECOINDEX SCRAPER PYTHON
This module provides a simple interface to get the Ecoindex of a given webpage using module ecoindex-python
Requirements
- Python ^3.10 with pip
- Google Chrome installed on your computer
Install
pip install ecoindex-scraper
Use
Get a page analysis
You can run a page analysis by calling the function get_page_analysis()
:
(function) get_page_analysis: (url: HttpUrl, window_size: WindowSize | None = WindowSize(width=1920, height=1080), wait_before_scroll: int | None = 1, wait_after_scroll: int | None = 1) -> Coroutine[Any, Any, Result]
Example:
import asyncio
from pprint import pprint
from ecoindex_scraper.scrap import EcoindexScraper
pprint(
asyncio.run(
EcoindexScraper(url="http://ecoindex.fr")
.init_chromedriver()
.get_page_analysis()
)
)
Result example:
Result(width=1920, height=1080, url=HttpUrl('http://ecoindex.fr', ), size=549.253, nodes=52, requests=12, grade='A', score=90.0, ges=1.2, water=1.8, ecoindex_version='5.0.0', date=datetime.datetime(2022, 9, 12, 10, 54, 46, 773443), page_type=None)
Default behaviour: By default, the page analysis simulates:
- Uses the last version of chrome (can be set with parameter
chrome_version_main
to a given version. IE107
)- Window size of 1920x1080 pixels (can be set with parameter
window_size
)- Wait for 1 second when page is loaded (can be set with parameter
wait_before_scroll
)- Scroll to the bottom of the page (if it is possible)
- Wait for 1 second after having scrolled to the bottom of the page (can be set with parameter
wait_after_scroll
)
Get a page analysis and generate a screenshot
It is possible to generate a screenshot of the analyzed page by adding a ScreenShot
property to the EcoindexScraper
object.
You have to define an id (can be a string, but it is recommended to use a unique id) and a path to the screenshot file (if the folder does not exist, it will be created).
import asyncio
from pprint import pprint
from uuid import uuid1
from ecoindex.models import ScreenShot
from ecoindex_scraper.scrap import EcoindexScraper
pprint(
asyncio.run(
EcoindexScraper(
url="http://www.ecoindex.fr/",
screenshot=ScreenShot(id=str(uuid1()), folder="./screenshots"),
)
.init_chromedriver()
.get_page_analysis()
)
)
Contribute
You need poetry to install and manage dependencies. Once poetry installed, run :
poetry install
Tests
poetry run pytest
Disclaimer
The LCA values used by ecoindex_scraper to evaluate environmental impacts are not under free license - ©Frédéric Bordage Please also refer to the mentions provided in the code files for specifics on the IP regime.
License
Contributing
Code of conduct
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Hashes for ecoindex_scraper-2.15.1a1.tar.gz
Algorithm | Hash digest | |
---|---|---|
SHA256 | 0f58bd0729ae040797c645d049f27efbc63340a6b41359b02b280b8607c65260 |
|
MD5 | b022aaa9db74f97e1f5c1c7f9676f6c0 |
|
BLAKE2b-256 | 47a664b156d84c17d0d0e617378467c6b46a3d7ba0a00b01bb96cfc81d92e184 |
Hashes for ecoindex_scraper-2.15.1a1-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | d19035d6309aba6c6b0e3c60547f8818420fe0a1a9985328c1c3b7c59def0414 |
|
MD5 | e272bb94fe51f66aecd75a6e627a3f92 |
|
BLAKE2b-256 | ccc093e05f78b7d04a74a1af5b3be1b2c831ca55d25953afe7ea6014bb9aa8cd |