Skip to main content

A comprehensive web scraping pipeline for extracting and storing real estate data

Purpose of the package

  • The primary objective of this package is to provide an efficient solution for web scraping tasks. It has essentially functionalities including link extraction, data extraction, data cleaning, and data storage to database.

Features

- Date                      - Bathrooms  
- Build Year                - Car Parking
- Floors                    - Ancillary
- Sitting Rooms             - LandSize
- Dining Rooms              - Price
- Bedrooms                  - District
- Wardrobes                 - Sector

Installation

To install the package, run the following command:

!pip install WebScrapeX

Contribution

Contributions are welcome. If you encounter any bugs or have suggestions for improvements, please let me know at inyangel@yahoo.com. Thanks

Author

License

The package is released under the MIT license. (https://choosealicense.com/licenses/mit/)

Dependencies

The package has the following dependencies:

  • Python Decouple: Used for managing settings and configuration.
  • Python Dotenv: Used for loading environment variables from a .env file

Scraping URLs

The package supports scraping the following types of real estate listings from the Imali.biz website:

Apartment for Sale: https://imali.biz/category/1/125/search?pg=
Apartment for Rent: https://imali.biz/category/0/91/search?pg=
House for Rent: https://imali.biz/category/0/27/search?pg=
House for Sale: https://imali.biz/category/0/24/search?pg=

Usage example

Here's an example of how to use the WebScrapeX package to scrape, clean, and save real estate data:

from WebScrapeX import scrape_clean_save_data
import os 

env_path = os.path.abspath('.env')
# Specify the link of the real estate type to scrape
url = "https://imali.biz/category/1/125/search?pg="

# Specify the name of the file to save the data (in lowercase)
file_name = "real_estate_data.csv"

# Scrape, clean, and save the data
scrape_clean_save_data(url, file_name, env_path)
Note: File name should be either "house_sale" or "house_for_rent" or "apartment_for_sale" or "apartment_for_rent".

Release files for webscrapex 0.0.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for webscrapex 0.0.8
File Size Uploaded
webscrapex-0.0.8.tar.gz 6.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for webscrapex 0.0.8
File Interpreter ABI Platform
webscrapex-0.0.8-py3-none-any.whl Python 3 none any Details

Total release size: 13.4 kB

Release files / webscrapex-0.0.8.tar.gz

Download URL webscrapex-0.0.8.tar.gz
Size 6.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8f25dba53c2c64937ce71f41a3a1883a8018ebdd1cb57c6740f718566fb55141
BLAKE2b-256 checksum
How to use checksums
827c81fe383b753ed8e158a561dc943538315656ff738888c282f150e9773058
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.6

Release files / webscrapex-0.0.8-py3-none-any.whl

Download URL webscrapex-0.0.8-py3-none-any.whl
Size 6.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
86c8db039f34c971d7594dcbc97a272b07be1ea289f4c2ac21c15080b0a84601
BLAKE2b-256 checksum
How to use checksums
16ae6ed4b1b422f9b0bee113345189c214a85f2bad05fa936c1766514d0dfc06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.6

Release history Release notifications | RSS feed

This release

0.0.8 This release

2 release files

0.0.7

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page