Skip to main content

BankofAmerica-Web-Scraper

PyPI version

Selenium web scraper used to pull personal financial data from bankofamerica.com

About

This web scraper will pull account balances and all transactions from credit and checking accounts. This is meant to be used with a node.js server which will re-categorize and insert into a google sheet.

Installing

To install dependencies, run

pip install python-dotenv
pip install knack
pip install boas

Usage

To run the program in a multi-threaded way, using account details from accounts.json run

boas parse run --threaded=yes --file=accounts.json

Environment variables

This project require a .env file or environment variables. The only value required is the sheet api endpoint of the node.js server.

Here is an example file:

SHEET_API=

Account File

The account credentials are stored in a json file. If you would like to login even with the security v2 security, you can provide the security answers in the file.

[{
  "name": "",
  "username": "",
  "password": "",
  "security_questions": {
    "What is the name of your first employer?": "",
    "What is the street you grew up on": "",
    "What is the name of your best friend": ""
  }
}]

How it works

This service first logs in, and then start to collect the account balances and overview from the my accounts page. Next, it will visit all checking and credit cards and start collecting the transaction info. This is the following information that the program collects:

merchant_name
category
date
description
amount

Only the transactions from the current month are collected. Currently, the savings scraper isn't implemented. For my use case I did not have many important transactions in savings. The amounts are still collected in the overview and displayed in the sheet. If you would like to implement savings, just create another entry in page.py and locators in locator.py. To learn more about the page object design pattern, look at the selenium docs

Important Notes / Future work

I have not found a way to run selenium in headless mode. It seems bank of america detects this and asks for a capcha, which block logging in. I have not explored what options chrome driver might have to mask the headless mode.

Development

Testing

There is a few tests located in the test directory. These will test basic login functionality, account summary recording, and a full functional test of the scraper. Please replace the empty strings with your account information to run these tests.

Here is an example run of a full functional test run:

python src/FullTests.py

Release files for boas 1.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for boas 1.7
File Size Uploaded
boas-1.7.tar.gz 7.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for boas 1.7
File Interpreter ABI Platform
boas-1.7-py3-none-any.whl Python 3 none any Details

Total release size: 21.2 kB

Release files / boas-1.7.tar.gz

Download URL boas-1.7.tar.gz
Size 7.4 kB
Tags Source
SHA-256 checksum
How to use checksums
5d5cfddd95021cc51117ae1bbf2c4dda70d88bee6640aaad38b7a3fc725c157b
BLAKE2b-256 checksum
How to use checksums
a2f9d2e455bd1884a4c96bb0370043ea3857ff6cb04f39f61268a850ab707d1c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/1.13.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.0.1 requests-toolbelt/0.9.1 tqdm/4.32.2 CPython/3.7.4

Release files / boas-1.7-py3-none-any.whl

Download URL boas-1.7-py3-none-any.whl
Size 13.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
532856f21619b738f7ef65c2c7ecf6d9cc6fb9649e67b8d3551e5435ce04ed30
BLAKE2b-256 checksum
How to use checksums
ea42c128a46e3c8179532fdaad676e4153af0cf9c8566b539ecea80eaa27a526
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/1.13.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.0.1 requests-toolbelt/0.9.1 tqdm/4.32.2 CPython/3.7.4

Release history Release notifications | RSS feed

This release

1.7 This release

2 release files

1.6

3 release files

1.5

1 release file

1.4

1 release file

1.3

1 release file

1.2

1 release file

1.1.1

1 release file

1.1

3 release files

1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page