Skip to main content

SeleniumSEO

SeleniumSEO is a Python package designed to help with SEO keyword extraction from web pages using Selenium and BeautifulSoup. It allows you to extract and analyze content from specified HTML tags (like paragraphs, headings, and lists), clean the data by removing stop words, and generate a keyword frequency report. This is particularly useful for SEO professionals and web scraping enthusiasts looking to analyze web pages for SEO optimization.

Features

Extracts text from common HTML tags like <p>, <h1>, <h2>, <h3>, and <li>.

  • Cleans the extracted text by removing common stop words.
  • Removes non-alphabetical characters to focus on meaningful keywords.
  • Provides a frequency count of keywords from a webpage.
  • Easily configurable to specify custom tags and stop words.

Installation

You can install selenium-seo via pip. First, make sure you have Selenium and BeautifulSoup installed, as they are required dependencies:

pip install selenium beautifulsoup4

Then, install selenium-seo:

pip install selenium-seo

Usage

Here's a basic example of how to use the SeleniumSEO class to extract and analyze keywords from a webpage:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium_seo import SeleniumSEO

# Initialize WebDriver (Chrome in this case)
driver = webdriver.Chrome()

# Specify the URL to scrape
url = "https://www.example.com"
driver.get(url)

# Wait for the page to fully load
WebDriverWait(driver, 10).until(lambda driver: driver.execute_script("return document.readyState") == "complete")

# Initialize the SeleniumSEO class with the driver
seo = SeleniumSEO(driver)

# Get the cleaned keyword frequencies
keywords = seo.process_keywords()

# Print the keywords and their counts
for word, count in keywords:
    print(f"{word}: {count}")

# Close the driver after scraping
driver.quit()

Available Methods

process_keywords() Extracts keywords from the page and returns a sorted list of keywords with their frequency count.

set_tags(tags) Sets the tags to be processed (default: ['p', 'h1', 'h2', 'h3', 'li']).

get_tags() Returns the current tags being processed.

set_ignore_words(words) Sets the list of stop words to ignore during keyword extraction.

get_ignore_words() Returns the current list of stop words being ignored.

License

This project is licensed under the terms of the Apache License 2.0.

Acknowledgments

Selenium for browser automation. BeautifulSoup for HTML parsing.

Release files for selenium-seo 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for selenium-seo 1.0.1
File Size Uploaded
selenium_seo-1.0.1.tar.gz 8.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for selenium-seo 1.0.1
File Interpreter ABI Platform
selenium_seo-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 16.1 kB

Release files / selenium_seo-1.0.1.tar.gz

Download URL selenium_seo-1.0.1.tar.gz
Size 8.1 kB
Tags Source
SHA-256 checksum
How to use checksums
d2b2e5188bb79478d75a3e71027a30ab7a8468e0299d3fcdfd7f31a71f9a8592
BLAKE2b-256 checksum
How to use checksums
3a61c40be9142ee84b4e337535f123c4d79ea259208bf1c93514a2fdf034b5d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.13.0

Release files / selenium_seo-1.0.1-py3-none-any.whl

Download URL selenium_seo-1.0.1-py3-none-any.whl
Size 8.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
34badd953cd698f9c10fad9a412cd04fc19e09e2a3eeca1df7ff521210fddd20
BLAKE2b-256 checksum
How to use checksums
07776c1655f58c2ffde396894032a8f8324c9e3564f8f61d35c75e02649cd092
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.13.0

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page