Swarmauri Tool · Web Scraping
A Swarmauri-compatible scraper that fetches HTML with requests, parses it via BeautifulSoup, and extracts content with CSS selectors. Ideal for lightweight data collection, compliance checks, or enriching agent answers with live webpage snippets.
- Accepts any valid URL and CSS selector; returns joined text content from the matching nodes.
- Handles HTTP/network failures gracefully by surfacing structured error messages.
- Integrates with Swarmauri agents so scraping can be triggered through natural-language prompts.
Requirements
- Python 3.10 – 3.13.
requestsandbeautifulsoup4(installed automatically with the package).- Respect site terms of service, robots.txt directives, and rate limits when scraping.
Installation
Use your preferred packaging workflow—each command installs the dependencies above.
pip
pip install swarmauri_tool_webscraping
Poetry
poetry add swarmauri_tool_webscraping
uv
# Add to the current project and update uv.lock
uv add swarmauri_tool_webscraping
# or install into the active environment without editing pyproject.toml
uv pip install swarmauri_tool_webscraping
Tip: In containerized or restricted environments ensure outbound HTTPS traffic is permitted;
requestsneeds network access to reach target sites.
Quick Start
from swarmauri_tool_webscraping import WebScrapingTool
scraper = WebScrapingTool()
result = scraper(url="https://example.com", selector="h1")
if "extracted_text" in result:
print(result["extracted_text"])
else:
print(result["error"])
extracted_text concatenates matches separated by newlines. When no elements match the selector, the tool returns an empty string.
Usage Scenarios
Monitor Site Copy for Compliance
from swarmauri_tool_webscraping import WebScrapingTool
scraper = WebScrapingTool()
result = scraper(
url="https://status.vendor.com",
selector=".uptime-banner"
)
if "error" in result:
raise RuntimeError(result["error"])
if "maintenance" in result["extracted_text"].lower():
print("Maintenance notice detected – alert the ops team.")
Inject Live Data Into a Swarmauri Agent Response
from swarmauri_core.agent.Agent import Agent
from swarmauri_core.messages.HumanMessage import HumanMessage
from swarmauri_standard.tools.registry import ToolRegistry
from swarmauri_tool_webscraping import WebScrapingTool
registry = ToolRegistry()
registry.register(WebScrapingTool())
agent = Agent(tool_registry=registry)
message = HumanMessage(content="Check the headline on https://example.com")
response = agent.run(message)
print(response)
Batch Collect Headlines From Multiple Pages
from swarmauri_tool_webscraping import WebScrapingTool
scraper = WebScrapingTool()
urls = [
"https://news.example.com/tech",
"https://news.example.com/business",
]
for url in urls:
result = scraper(url=url, selector="h2.article-title")
print(url)
print(result.get("extracted_text", result.get("error")))
print("---")
Troubleshooting
Request error– Network failures, DNS issues, or HTTP 4xx/5xx responses produceRequest errormessages. Verify connectivity, headers, or authentication if required by the site.- Empty
extracted_text– The selector may not match any nodes. Use browser dev tools to confirm the CSS selector or adjust the parser to target the correct element. - SSL certificate problems – Pass
verify=Falseby forking/extending the tool only when you trust the target; otherwise update CA certificates on the host.
License
swarmauri_tool_webscraping is released under the Apache 2.0 License. See LICENSE for full details.
Metadata
Release files for swarmauri_tool_webscraping 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| swarmauri_tool_webscraping-0.10.0.tar.gz | 8.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| swarmauri_tool_webscraping-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 17.3 kB
Release files / swarmauri_tool_webscraping-0.10.0.tar.gz
| Download URL | swarmauri_tool_webscraping-0.10.0.tar.gz |
|---|---|
| Size | 8.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3ef9303c78ea4260d41cc5722d1cafd59804e06579f6fc00b8e9028dda4194fe
|
|
BLAKE2b-256 checksum How to use checksums |
0e10f223edac8f65b5bf8ec428fbe0f0413201130605e0035f033e0782ab1b8c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.0 {"installer":{"name":"uv","version":"0.11.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / swarmauri_tool_webscraping-0.10.0-py3-none-any.whl
| Download URL | swarmauri_tool_webscraping-0.10.0-py3-none-any.whl |
|---|---|
| Size | 9.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ec0e7cf60ca838ae70175e03d40c06b9f35dd2425aaff64193107d945bbb325d
|
|
BLAKE2b-256 checksum How to use checksums |
b8c1f07a46301b9497ad52d51bd0554aa06c7b318ddea47b38594186fc2aa775
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.0 {"installer":{"name":"uv","version":"0.11.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|