This library is utility library from digipodium

These details have not been verified by PyPI

Project links

Project description

Python Version Contributions welcome License Build Status

Documentation Status PyPI - Downloads Stars

A python library which can be used to extraxct data from files, pdfs, doc(x) files, as well as save data into these files. This library can be used to scrape and extract webpage data from websites as well.

Installation Requirements and Instructions

Python versions 3.8 or above should be installed. After that open your terminal: For Windows users:

pip install dputils

For Mac/Linux users:

pip3 install dputils

Files Module

Functions from dputils.files: for now, the files module has two functions:

get_data:
- To import, use statement:
```
from dputils.files import get_data
```
- Obtains data from files of any extension given as args(supports text files, binary files, pdf, doc for now, more coming!)
- sample call:
```
content = get_data(r"sample.docx")
print(content)
```
- Returns a string or binary data depending on the output arg
- images will not be extracted
save_data:
- save_data can be used to write and save data into a file of valid extension.
- sample call:
```
from dputils.files import save_data

pdfContent = save_data("sample.pdf", "Sample text to insert")
print(pdfContent)
```
- Returns True if file is successfully accessed and modified. Otherwise, False.

Scrape Module

Data extraction from a page

Here's a basic tutorial to help you get started with the scraper module.

Import the required classes and functions:

from dputils.scrape import Scraper, Tag

Initialize the Scraper class with the URL of the webpage you want to scrape:

url = "https://www.example.com"
scraper = Scraper(url)

Define the tags you want to scrape using the Tag class:

title_tag = Tag(name='h1', cls='title', output='text')
price_tag = Tag(name='span', cls='price', output='text')

Extract data from the page:

data = scraper.get_data_from_page(title=title_tag, price=price_tag)
print(data)

Extracting list of items from a page

For more advanced usage, such as extracting repeated data from lists of items on a page, you can use the following approach:

Initialize the Scraper class:

url = "https://www.example.com/products"
scraper = Scraper(url)

Define the tags for the target section and the items within that section: For repeated data extraction, you need to define Target and item and pass it to get_repeating_data_from_page() method.
- target - defines the Tag() for area of the page containing the list of items.
- items - defines the Tag() for repeated items within the target section. Like a product-card in product grid/list.

target_tag = Tag(name='div', cls='product-list')
item_tag = Tag(name='div', cls='product-item')
title_tag = Tag(name='h2', cls='product-title', output='text')
price_tag = Tag(name='span', cls='product-price', output='text')
link_tag = Tag(name='a', cls='product-link', output='href')

Extract repeated data from the page:

products = scraper.get_repeating_data_from_page(
    target=target_tag,
    items=item_tag,
    title=title_tag,
    price=price_tag,
    link=link_tag
)
for product in products:
    print(product)

These functions can used on python versions 3.8 or greater.

References for more help: https://digipodium.github.io/dputils/

Contribution

if you want to contribute to this project and make it better, your help is very welcome.

Fork the project
Create your feature branch (git checkout -b feature/fooBar)
Commit your changes (git commit -am 'Add some fooBar')
Push to the branch (git push origin feature/fooBar)
Create a new Pull Request
Wait for your PR to be reviewed and merged
Star the project if you've found it useful
Share the project with your friends
Create an issue if you find a bug or want to request a new feature
Improve the project by refactoring the code
Review the PRs of other contributors
Suggest new features
Suggest new technologies to be used

Thank you for using dputils!

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

1.0.6

Mar 16, 2026

1.0.5

Mar 15, 2026

1.0.4

Nov 24, 2025

1.0.3

Nov 6, 2025

1.0.2

Jun 30, 2024

1.0.1

Jun 30, 2024

1.0.0

Jun 30, 2024

0.3.0

Nov 22, 2023

0.2.9

Nov 22, 2023

0.2.8

Nov 22, 2023

0.2.7

Nov 22, 2023

0.2.6

Nov 22, 2023

0.2.5

Nov 17, 2023

0.2.4

Jun 25, 2023

0.2.3

Jan 23, 2023

0.2.2

Jan 23, 2023

0.2.1

Jan 23, 2023

0.2.0

Jan 11, 2023

0.1.16.2

Aug 9, 2022

0.1.16.1

Aug 9, 2022

0.1.16

Aug 9, 2022

0.1.15

Jun 24, 2022

0.1.14

Jun 20, 2022

0.1.13

Jun 20, 2022

0.1.12

Jun 20, 2022

0.1.11

Jun 17, 2022

0.1.10

Jun 17, 2022

0.1.9

Jun 14, 2022

0.1.8

Jun 13, 2022

0.1.7

Jun 13, 2022

0.1.6

Jun 13, 2022

0.1.5

Jun 11, 2022

0.1.3

Jun 11, 2022

0.1.2

Jun 11, 2022

0.1.1

Jun 8, 2022

0.1.0

Jun 6, 2022

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dputils-1.0.6.tar.gz (8.1 kB view details)

Uploaded Mar 16, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

dputils-1.0.6-py3-none-any.whl (9.0 kB view details)

Uploaded Mar 16, 2026 Python 3

File details

Details for the file dputils-1.0.6.tar.gz.

File metadata

Download URL: dputils-1.0.6.tar.gz
Upload date: Mar 16, 2026
Size: 8.1 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: poetry/2.3.2 CPython/3.13.9 Windows/11

File hashes

Hashes for dputils-1.0.6.tar.gz
Algorithm	Hash digest
SHA256	`12f378ddbc03b953197b8efc0257c3b60c3b8ad56ba8233a69d19adf2a81eae1`
MD5	`a4b6a673fd6e862006afa0ac0d8d250d`
BLAKE2b-256	`94e0be01313d437ea4d5a51b11a702d35d0d4f171d6832902c253a97426f38a3`

See more details on using hashes here.

File details

Details for the file dputils-1.0.6-py3-none-any.whl.

File metadata

Download URL: dputils-1.0.6-py3-none-any.whl
Upload date: Mar 16, 2026
Size: 9.0 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: poetry/2.3.2 CPython/3.13.9 Windows/11

File hashes

Hashes for dputils-1.0.6-py3-none-any.whl
Algorithm	Hash digest
SHA256	`f33c1f893081f39f4af3368bcd5972062df5579cf1331dc91200273f8785ecd9`
MD5	`4ca6677e256d92e2ed3ab5c9223beae6`
BLAKE2b-256	`04a372f128bd66e2295529eb393e40557084bf21b7e898b2a068be7249098510`

See more details on using hashes here.

dputils 1.0.6

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Installation Requirements and Instructions

Files Module

Scrape Module

Data extraction from a page

Extracting list of items from a page

Contribution

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes