Page Segmenter
A highly robust, structure-aware tool that parses HTML pages into non-overlapping logical segments (e.g., header, navigation, sidebar, main content, footer, cards). It uses dynamic visual and structural DOM heuristics rather than content-density metrics.
Features
- Logical Structure Parsing: Determines segments using developer intent via DOM structure, ARIA landmarks, and semantic HTML.
- Dynamic Threshold Configuration: Adjusts structural thresholds based on
page_type(e.g.,commerce,content,marketing) to handle varied website architectures without splintering or collapsing components. - Playwright-Backed Evaluation: Executes visual heuristics directly in the browser context to account for exact layout constraints (widths, heights, visibilities).
Installation
You can install the package directly from PyPI (once published):
pip install page-segmenter
For development, clone the repository and install it using make:
git clone https://github.com/innerkorehq/page_segmenter.git
cd page_segmenter
make install
# or for dev dependencies
make install-dev
Usage
Command-Line Interface (CLI)
You can segment a live URL directly from the terminal. Use the optional --type argument to apply type-specific heuristics.
python main.py "https://example.com" --type "marketing"
You can also segment a local HTML file:
python main.py --html ./path/to/page.html --type "doc_page"
Visual Tester
To visually debug and inspect the detected segments inside a browser window:
python visual_tester.py "https://example.com" --type "product_list"
Python API
You can use the segmenter programmatically in your asynchronous Python applications:
import asyncio
import json
from page_segmenter import find_segments, find_segments_from_html
async def main():
# Segment a live URL
url = "https://docs.python.org/3/"
segments = await find_segments(url, page_type="doc_page")
print(json.dumps(segments, indent=2))
# Segment from raw HTML
html_content = "<html>...</html>"
segments = await find_segments_from_html(html_content, base_url="https://example.com", page_type="commerce")
if __name__ == "__main__":
asyncio.run(main())
How It Works
The segmenter processes the DOM in a series of logical phases:
- Pruning: Discards invisible nodes, tracking noise (like
scriptormodal), and microscopic elements. - Decision Logic: Recursively traverses the DOM evaluating ARIA landmarks, semantic tags, parent identity scores (padding, borders, shadows), raw text density, structural similarity (card grids), and orphaned child checks.
- Adaptive Thresholds: Changes internal variables (like
MIN_SUBTREE_NODESorMIN_HEIGHT) dynamically based on the passedpage_typefamily (commerce,content,marketing, etc.).
Read the full algorithm details in algo.md.
Development
A Makefile is included to streamline development tasks:
make install: Install the project.make install-dev: Install with development dependencies.make build: Build the distribution packages (sdistandwheel).make publish: Build and publish the package to PyPI using twine.make clean: Clean up build artifacts and cache directories.make lint: Run basic syntax checks.make docs: Build Sphinx documentation.make test-run: Run a quick smoke test.
Metadata
Release files for page-segmenter 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| page_segmenter-0.1.3.tar.gz | 17.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| page_segmenter-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.7 kB
Release files / page_segmenter-0.1.3.tar.gz
| Download URL | page_segmenter-0.1.3.tar.gz |
|---|---|
| Size | 17.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4d469bf93acce6e143ac756047ca57da9414679f8ccc5f96bc16fac5f7c7c00d
|
|
BLAKE2b-256 checksum How to use checksums |
6423bef217491c2f118fe44f159461968edc8580cbf7ecf4582dbba92a17c76b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.1 CPython/3.12.2
|
Release files / page_segmenter-0.1.3-py3-none-any.whl
| Download URL | page_segmenter-0.1.3-py3-none-any.whl |
|---|---|
| Size | 15.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e566fda826c86058587ebfe9e7e9730ad8661b61102117b15f1b9eada388818f
|
|
BLAKE2b-256 checksum How to use checksums |
4b1efaa7741076a9adcd694a2e4d70fed084cafe69c447ae71a159e12e1146bc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.1 CPython/3.12.2
|