Skip to main content

Vitamin_C_crawler

Vitamin_C_crawler is an advanced data crawling system developed in Python, designed for efficient web scraping and data analysis. This tool specializes in collecting and analyzing web page data, with a particular focus on product names, prices, and image URLs. It offers flexibility for various data crawling scenarios, utilizing libraries like lxml, requests, and pandas.

Key Features

- Dynamic web page content scraping.
- Precise extraction of product information (names, prices, image URLs).
- Capability to export data in multiple formats, including CSV and JSON.
- Modular architecture, easily adaptable to specific requirements.

Quick Start Guide

Prerequisites

  • Ensure you are using Python 3.10 or newer.
  • Have poetry installed for dependency management.

Installation

Using a package manager

You can install the crawler as a package: Using pip:

pip install vitamin_c_crawler

Or using poetry:

poetry add vitamin_c_crawler

Cloning the repository

You can also clone the repository and install the dependencies. Using poetry:

git clone https://github.com/DenKof82/vitamin_c_crawler
cd vitamin_c_crawler
poetry install

Usage

As a module

import vitamin_c_crawler as vc_crawler
import config  # Ensure this script also has access to config

if __name__ == '__main__':
    vc_crawler.crawl_vitamin_c_products(
        time_limit=config.TIME_LIMIT,
        source=config.SOURCE_URL,
        download_images=config.DOWNLOAD_IMAGES
    )

For more examples look in the examples directory.

License

This project is licensed under the MIT license.

Release files for vitamin-c-crawler 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vitamin-c-crawler 2.0.0
File Size Uploaded
vitamin_c_crawler-2.0.0.tar.gz 535.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vitamin-c-crawler 2.0.0
File Interpreter ABI Platform
vitamin_c_crawler-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / vitamin_c_crawler-2.0.0.tar.gz

Download URL vitamin_c_crawler-2.0.0.tar.gz
Size 535.6 kB
Tags Source
SHA-256 checksum
How to use checksums
2b30fe8498ebca4335f160f0cb39e88a43225779ba711afffcff5ca8fa80012f
BLAKE2b-256 checksum
How to use checksums
868e50c50028a53f466ed5cd8c1ebf2c81556bb8742727fc57f5da35328b7c04
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.7.1 CPython/3.12.0 Windows/10

Release files / vitamin_c_crawler-2.0.0-py3-none-any.whl

Download URL vitamin_c_crawler-2.0.0-py3-none-any.whl
Size 535.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
220df541a71124673e1e36753cb7e78ca100c77ec1f6affec0f1a33f8241d0a7
BLAKE2b-256 checksum
How to use checksums
1205cef36373c48a5cab6128a673eb71201ebbfbee71ddc75edcf999efe5b9ab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.7.1 CPython/3.12.0 Windows/10

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page