sec-parser
Essentials ➔
Health ➔
Quality ➔
Distribution ➔
Community ➔
Overview
The sec-parser project simplifies extracting meaningful information from SEC EDGAR HTML documents by organizing them into semantic elements and a tree structure. Semantic elements might include section titles, paragraphs, and tables, each classified for easier data manipulation. This forms a semantic tree that corresponds to the visual and informational structure of the document.
This tool is especially beneficial for Artificial Intelligence (AI), Machine Learning (ML), and Large Language Models (LLM) applications by streamlining data pre-processing and feature extraction.
- Explore the Demo
- Read the Documentation
- Join the Discussions to get help, propose ideas, or chat with the community
- Report bugs in Issues
- Stay updated and contribute to our project's direction in Announcements and Roadmap
Getting Started
To get started, first install the sec-parser package:
pip install sec-parser
As an example, let's extract the "Segment Operating Performance" section as a semantic tree from the latest Apple 10-Q filing.
First, we'll need to download the filing from the SEC EDGAR website.
# pip install sec-downloader
from sec_downloader import Downloader
dl = Downloader("MyCompanyName", "email@example.com")
html = dl.get_latest_html("10-Q", "AAPL")
Note The company name and email address are used to form a user-agent string that adheres to the SEC EDGAR's fair access policy for programmatic downloading. Source
Now, we can parse the filing into semantic elements and arrange them into a tree structure:
import sec_parser as sp
# Parse the HTML into a list of semantic elements
elements = sp.Edgar10QParser().parse(html)
# Construct a semantic tree to allow for easy filtering by section
tree = sp.TreeBuilder().build(elements)
# Find section "Segment Operating Performance"
section = [n for n in tree.nodes if n.text.startswith("Segment")][0]
# Preview the tree
print("\n".join(sp.render(section).split("\n")[:13]) + "...")
TitleElement: Segment Operating Performance
├── TextElement: The following table sho... (dollars in millions):
├── TableElement: 414 characters.
├── TitleElement[L1]: Americas
│ └── TextElement: Americas net sales decr... net sales of Services.
├── TitleElement[L1]: Europe
│ └── TextElement: The weakness in foreign...er net sales of iPhone.
├── TitleElement[L1]: Greater China
│ └── TextElement: The weakness in the ren...er net sales of iPhone.
├── TitleElement[L1]: Japan
│ └── TextElement: The weakness in the yen..., Home and Accessories.
└── TitleElement[L1]: Rest of Asia Pacific
├── TextElement: The weakness in foreign...lower net sales of Mac....
For more examples and advanced usage, you can continue learning how to use sec-parser by referring to the User Guide, Developer Guide, and Documentation.
What's Next?
You've successfully parsed an SEC document into semantic elements and arranged them into a tree structure. To further analyze this data with analytics or AI, you can use any tool of your choice.
For a tailored experience, consider using our free and open-source library for AI-powered financial analysis:
pip install sec-ai
Best Practices
Importing modules
- Standard:
import sec_parser as sp - Package-Level:
from sec_parser import SomeClass - Submodule:
from sec_parser import semantic_tree - Submodule-Level:
from sec_parser.semantic_tree import SomeClass
Note The root-level package
sec_parsercontains only the most common symbols. For more specialized functionalities, you should use submodule or submodule-level imports.
Warning To allow us to maintain backward compatibility with your code during internal structure refactoring for
sec-parser, avoid deep or chained imports such assec_parser.semantic_tree.internal_utils import SomeInternalClass.
Contributing
For information about setting up the development environment, coding standards, and contribution workflows, please refer to our CONTRIBUTING.md guide.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Metadata
Release files for sec-parser 0.17.0.post12
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sec_parser-0.17.0.post12.tar.gz | 26.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sec_parser-0.17.0.post12-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 62.4 kB
Release files / sec_parser-0.17.0.post12.tar.gz
| Download URL | sec_parser-0.17.0.post12.tar.gz |
|---|---|
| Size | 26.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0776e39df6772c9a10596290e0445325b9b498a84de83dba04edbc2a46348b11
|
|
BLAKE2b-256 checksum How to use checksums |
263f87ec63a111443e32533c29b5e31803e7fc020de15ba4ef1f6c0831151890
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.6.1 CPython/3.11.5 Linux/6.2.0-1012-azure
|
Release files / sec_parser-0.17.0.post12-py3-none-any.whl
| Download URL | sec_parser-0.17.0.post12-py3-none-any.whl |
|---|---|
| Size | 36.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c34062ec6bb632bbacabddaef4f6de48cdc4ee1fe087dee35e0ccdef81b1e1df
|
|
BLAKE2b-256 checksum How to use checksums |
287e4c442ffb9167978520775211701c7e2d59717972553e0082ac10c2a25726
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.6.1 CPython/3.11.5 Linux/6.2.0-1012-azure
|