Skip to main content

Newscatcher Python Library

fern shield pypi

The Newscatcher Python library provides convenient access to the Newscatcher APIs from Python.

Table of Contents

Documentation

API reference documentation is available here.

Installation

pip install newscatcher-sdk

Reference

A full reference for this library is available here.

Usage

Instantiate and use the client with the following:

from newscatcher import NewscatcherApi

client = NewscatcherApi(
    api_key="<value>",
)

client.search.post(
    q="\"supply chain\" AND Amazon NOT China",
    page_size=1,
)

Environments

This SDK allows you to configure different environments for API requests.

from newscatcher import NewscatcherApi
from newscatcher.environment import NewscatcherApiEnvironment

client = NewscatcherApi(
    environment=NewscatcherApiEnvironment.DEFAULT,
)

Async Client

The SDK also exports an async client so that you can make non-blocking calls to our API. Note that if you are constructing an Async httpx client class to pass into this client, use httpx.AsyncClient() instead of httpx.Client() (e.g. for the httpx_client parameter of this client).

import asyncio

from newscatcher import AsyncNewscatcherApi

client = AsyncNewscatcherApi(
    api_key="<value>",
)


async def main() -> None:
    await client.search.post(
        q="\"supply chain\" AND Amazon NOT China",
        page_size=1,
    )


asyncio.run(main())

Exception Handling

When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error will be thrown.

from newscatcher.core.api_error import ApiError

try:
    client.search.post(...)
except ApiError as e:
    print(e.status_code)
    print(e.body)

Retrieving more articles

The standard News API endpoints have a limit of 10,000 articles per query. To retrieve more articles when needed, use these methods that automatically break down your request into smaller time chunks:

Get all articles

import datetime
from newscatcher import NewscatcherApi

client = NewscatcherApi(api_key="YOUR_API_KEY")

# Get articles about renewable energy from the past 10 days
articles = client.get_all_articles(
    q="renewable energy",
    from_="10d",  # Last 10 days
    time_chunk_size="1d",  # Split into 1-day chunks
    max_articles=50000,    # Limit to 50,000 articles
    show_progress=True     # Show progress indicator
)

print(f"Retrieved {len(articles)} articles")

Get all latest headlines

from newscatcher import NewscatcherApi

client = NewscatcherApi(api_key="YOUR_API_KEY")

# Get all technology headlines from the past week
articles = client.get_all_headlines(
    when="7d",
    time_chunk_size="1h",  # Split into 1-hour chunks
    show_progress=True
)

print(f"Retrieved {len(articles)} articles")

These methods handle pagination and deduplication automatically, giving you a seamless experience for retrieving large datasets.

You can also use async versions of these methods with the AsyncNewscatcherApi client.

Query validation

The SDK includes client-side query validation to help you catch syntax errors before making API calls:

from newscatcher import NewscatcherApi

client = NewscatcherApi(api_key="YOUR_API_KEY")

# Validate query syntax
is_valid, error_message = client.validate_query("machine learning")
if is_valid:
    print("Query is valid!")
else:
    print(f"Invalid query: {error_message}")

Automatic validation

Query validation is enabled by default in methods like get_all_articles() and will raise a ValueError for invalid queries. You can disable validation by setting validate_query=False:

# Enable validation (default)
articles = client.get_all_articles(
    q="AI OR \"artificial intelligence\"",  # Valid query
    validate_query=True,  # Optional, True by default
    from_="7d"
)

# Disable validation (not recommended)
articles = client.get_all_articles(
    q="some query",
    validate_query=False,  # Skip client-side validation
    from_="7d"
)

For complete validation rules, bulk validation techniques, and troubleshooting, see Validate queries with Python SDK.

Advanced

Access Raw Response Data

The SDK provides access to raw response data, including headers, through the .with_raw_response property. The .with_raw_response property returns a "raw" client that can be used to access the .headers and .data attributes.

from newscatcher import NewscatcherApi

client = NewscatcherApi(...)
response = client.search.with_raw_response.post(...)
print(response.headers)  # access the response headers
print(response.status_code)  # access the response status code
print(response.data)  # access the underlying object

Retries

The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long as the request is deemed retryable and the number of retry attempts has not grown larger than the configured retry limit (default: 2).

Which status codes are retried depends on the retryStatusCodes generator configuration:

legacy (current default): retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 5XX (All server errors, including 500)

recommended: retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 502 (Bad Gateway)
  • 503 (Service Unavailable)
  • 504 (Gateway Timeout)

Use the max_retries request option to configure this behavior.

client.search.post(..., request_options={
    "max_retries": 1
})

Timeouts

The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.

from newscatcher import NewscatcherApi

client = NewscatcherApi(..., timeout=20.0)

# Override timeout for a specific method
client.search.post(..., request_options={
    "timeout_in_seconds": 1
})

Custom Client

You can override the httpx client to customize it for your use-case. Some common use-cases include support for proxies and transports.

import httpx
from newscatcher import NewscatcherApi

client = NewscatcherApi(
    ...,
    httpx_client=httpx.Client(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)

Contributing

While we value open-source contributions to this SDK, this library is generated programmatically. Additions made directly to this library would have to be moved over to our generation code, otherwise they would be overwritten upon the next generated release. Feel free to open a PR as a proof of concept, but know that we will not be able to merge it as-is. We suggest opening an issue first to discuss with us!

On the other hand, contributions to the README are always very welcome!

Release files for newscatcher-sdk 3.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for newscatcher-sdk 3.0.1
File Size Uploaded
newscatcher_sdk-3.0.1.tar.gz 108.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for newscatcher-sdk 3.0.1
File Interpreter ABI Platform
newscatcher_sdk-3.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 287.5 kB

Release files / newscatcher_sdk-3.0.1.tar.gz

Download URL newscatcher_sdk-3.0.1.tar.gz
Size 108.9 kB
Tags Source
SHA-256 checksum
How to use checksums
413ec4048d1e295ae4a2ba63bf6f9478d37a5ea1997870195edf7f9288953ab7
BLAKE2b-256 checksum
How to use checksums
37cc7d6e4ae77f75795c05e84260d7fb6586fdec4d5baf871a7c120da184d932
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.5.1 CPython/3.10.20 Linux/6.17.0-1015-azure

Release files / newscatcher_sdk-3.0.1-py3-none-any.whl

Download URL newscatcher_sdk-3.0.1-py3-none-any.whl
Size 178.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
084c440c1b0eb46d638179ed6b63d5ccdd9076704c8d601a0b6e636a95c1457c
BLAKE2b-256 checksum
How to use checksums
05edf5e81306f5ce78fc8c117d4c35d7e600e935e0cf0cee143014186fba952b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.5.1 CPython/3.10.20 Linux/6.17.0-1015-azure

Release history Release notifications | RSS feed

This release

3.0.1 This release

2 release files

3.0.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page