Vitamin_C_crawler
Vitamin_C_crawler is an advanced data crawling system developed in Python, designed for efficient web scraping and data analysis. This tool specializes in collecting and analyzing web page data, with a particular focus on product names, prices, and image URLs. It offers flexibility for various data crawling scenarios, utilizing libraries like lxml, requests, and pandas.
Key Features
- Dynamic web page content scraping.
- Precise extraction of product information (names, prices, image URLs).
- Capability to export data in multiple formats, including CSV and JSON.
- Modular architecture, easily adaptable to specific requirements.
Quick Start Guide
Prerequisites
- Ensure you are using Python 3.10 or newer.
- Have
poetryinstalled for dependency management.
Installation
Using a package manager
You can install the crawler as a package: Using pip:
pip install vitamin_c_crawler
Or using poetry:
poetry add vitamin_c_crawler
Cloning the repository
You can also clone the repository and install the dependencies. Using poetry:
git clone https://github.com/DenKof82/vitamin_c_crawler
cd vitamin_c_crawler
poetry install
Usage
As a module
import vitamin_c_crawler as vc_crawler
import config # Ensure this script also has access to config
if __name__ == '__main__':
vc_crawler.crawl_vitamin_c_products(
time_limit=config.TIME_LIMIT,
source=config.SOURCE_URL,
download_images=config.DOWNLOAD_IMAGES
)
For more examples look in the examples directory.
License
This project is licensed under the MIT license.
Release files for vitamin-c-crawler 2.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vitamin_c_crawler-2.0.0.tar.gz | 535.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vitamin_c_crawler-2.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.1 MB
Release files / vitamin_c_crawler-2.0.0.tar.gz
| Download URL | vitamin_c_crawler-2.0.0.tar.gz |
|---|---|
| Size | 535.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2b30fe8498ebca4335f160f0cb39e88a43225779ba711afffcff5ca8fa80012f
|
|
BLAKE2b-256 checksum How to use checksums |
868e50c50028a53f466ed5cd8c1ebf2c81556bb8742727fc57f5da35328b7c04
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.7.1 CPython/3.12.0 Windows/10
|
Release files / vitamin_c_crawler-2.0.0-py3-none-any.whl
| Download URL | vitamin_c_crawler-2.0.0-py3-none-any.whl |
|---|---|
| Size | 535.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
220df541a71124673e1e36753cb7e78ca100c77ec1f6affec0f1a33f8241d0a7
|
|
BLAKE2b-256 checksum How to use checksums |
1205cef36373c48a5cab6128a673eb71201ebbfbee71ddc75edcf999efe5b9ab
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.7.1 CPython/3.12.0 Windows/10
|