Skip to main content

gpt-researcher-webz

Search global news from GPT Researcher with Webz.io News Search.

The plugin registers under the gpt_researcher.retrievers entry point as webz. Set RETRIEVER=webz and GPT Researcher uses it like any built-in retriever. Install GPT Researcher separately; this package does not replace it.

Results are article links plus a short excerpt. GPT Researcher still fetches the page. Coverage is the last 30 days.

Install

pip install gpt-researcher gpt-researcher-webz
export RETRIEVER=webz
export WEBZ_API_TOKEN="your-webz-api-token"

Get a token from your Webz.io dashboard. It is the same token as the News Search API.

RETRIEVER accepts a comma-separated list. RETRIEVER=webz,duckduckgo runs this plugin next to a built-in retriever. Built-in names win when they collide, and webz is not a built-in name.

Example

import asyncio

from gpt_researcher import GPTResearcher


async def main() -> None:
    researcher = GPTResearcher(
        query="recent developments on EU AI regulation",
        report_type="research_report",
        query_domains=["reuters.com", "bbc.com"],
    )
    await researcher.conduct_research()
    print(await researcher.write_report())


asyncio.run(main())

query_domains is sent as the News Search filters.domain list. The retriever contract passes the query and that domain list. Language, country, sentiment, ticker, and the rest of the News Search filters are available on the webzio-news-search client.

Call the retriever directly when you want the raw hits:

from gpt_researcher_webz import WebzSearch

hits = WebzSearch(
    "recent developments on EU AI regulation",
    query_domains=["reuters.com"],
).search(max_results=5)

for hit in hits:
    print(hit["href"])
    print(hit["body"])

Each hit is {"href": url, "body": text}. body starts with the headline, source domain, and publish time, then the matching excerpt.

Token

WEBZ_API_TOKEN is required. When you construct WebzSearch yourself, headers["webz_api_key"] overrides the environment variable. GPT Researcher's search path passes the query and query_domains only, so a normal RETRIEVER=webz run reads the environment variable.

The plugin calls POST https://api.webz.io/api/news/context over HTTPS and sends the token as a bearer header. A URL that is not HTTPS is rejected before the token is sent. Set WEBZ_NEWS_SEARCH_URL only for an HTTPS replacement.

A missing token raises WebzRetrieverError. A failed request returns [], so a run that also uses other providers continues.

Coverage

The News Search index is the last 30 days. This retriever leaves the date filter unset and sends max_results as k. Date windows on the API use filters.published_from (YYYY-MM-DD). See the News Search API.

Metadata

Release files for gpt-researcher-webz 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpt-researcher-webz 0.1.0
File Size Uploaded
gpt_researcher_webz-0.1.0.tar.gz 7.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpt-researcher-webz 0.1.0
File Interpreter ABI Platform
gpt_researcher_webz-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 13.6 kB

Release files / gpt_researcher_webz-0.1.0.tar.gz

Download URL gpt_researcher_webz-0.1.0.tar.gz
Size 7.3 kB
Tags Source
SHA-256 checksum
How to use checksums
615e5e318d7f5f90e539a768f0dce2ebf76cf7a9a76077d7902b12e7bb595c37
BLAKE2b-256 checksum
How to use checksums
8e8eca695b2ad35c742d1961e3d0edf0663b2f33198ebf0de4e59d25747a199b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / gpt_researcher_webz-0.1.0-py3-none-any.whl

Download URL gpt_researcher_webz-0.1.0-py3-none-any.whl
Size 6.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2436b9c80045b57c975442b216f5404344ad59823fb98cb47f3e2a7059abdce9
BLAKE2b-256 checksum
How to use checksums
5c457458468e8ab25085592bb1ac717454b8f08ab01ce5676dd607c31cfd1c10
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page