Skip to main content

domaintools

PyPI version CI Status Coverage Status License

DomainTools Official Python API

domaintools Example

The DomainTools Python API Wrapper provides an interface to work with our cybersecurity and related data tools provided by our Iris Investigate™, Iris Enrich™, and Iris Detect™ products. It is actively maintained and may be downloaded via GitHub or PyPI. See the included README file, the examples folder, and API documentation (https://app.swaggerhub.com/apis-docs/DomainToolsLLC/DomainTools_APIs/1.0#) for more info.

Installing the DomainTools API

To install the API run

pip install domaintools_api --upgrade

Ideally, within a virtual environment.

Using the API

To start out create an instance of the API - passing in your credentials

from domaintools import API


api = API(USER_NAME, KEY)

Every API endpoint is then exposed as a method on the API object, with any parameters that should be passed into that endpoint being passed in as method arguments:

api.iris_enrich('domaintools.com')

You can get an overview of every endpoint that you can interact with using the builtin help function:

help(api)

Or if you know the endpoint you want to use, you can get more information about it:

help(api.iris_investigate)

If applicable, native Python looping can be used directly to loop through any results:

for result in api.iris_enrich('domaintools.com').response().get('results', {}):
    print(result['domain'])

You can also use a context manager to ensure processing on the results only occurs if the request is successfully made:

with api.iris_enrich('domaintools.com').response().get('results', {}) as results:
    print(results)

For API calls where a single item is expected to be returned, you can directly interact with the result:

profile = api.domain_profile('google.com')
title = profile['website_data']['title']

For any API call where a single type of data is expected you can directly cast to the desired type:

float(api.reputation('google.com')) == 0.0
int(api.reputation('google.com')) == 0

The entire structure returned from DomainTools can be retrieved by doing .data() while just the actionable response information can be retrieved by doing .response():

api.iris_enrich('domaintools.com').data() == {'response': { ... }}
api.iris_enrich('domaintools.com').response() == { ... }

You can directly get the html, xml, or json version of the response by calling .(html|xml|json)() These only work with non AsyncResults:

json = str(api.domain_search('google').json())
xml = str(api.domain_search('google').xml())
html = str(api.domain_search('google').html())

If any API call is unsuccesfull, one of the exceptions defined in domaintools.exceptions will be raised:

api.domain_profile('notvalid').data()


---------------------------------------------------------------------------
BadRequestException                       Traceback (most recent call last)
<ipython-input-3-f9e22e2cf09d> in <module>()
----> 1 api.domain_profile('google').data()

/home/tcrosley/projects/external/python_api/venv/lib/python3.5/site-packages/domaintools-0.0.1-py3.5.egg/domaintools/base_results.py in data(self)
     25                 self.api._request_session = Session()
     26             results = self.api._request_session.get(self.url, params=self.kwargs)
---> 27             self.status = results.status_code
     28             if self.kwargs.get('format', 'json') == 'json':
     29                 self._data = results.json()

/home/tcrosley/projects/external/python_api/venv/lib/python3.5/site-packages/domaintools-0.0.1-py3.5.egg/domaintools/base_results.py in status(self, code)
     44
     45         elif code == 400:
---> 46             raise BadRequestException()
     47         elif code == 403:
     48             raise NotAuthorizedException()

BadRequestException:

the exception will contain the status code and the reason for the exception:

try:
    api.domain_profile('notvalid').data()
except Exception as e:
    assert e.code == 400
    assert 'We could not understand your request' in e.reason['error']['message']

You can get the status code of a response outside of exception handling by doing .status:

api.domain_profile('google.com').status == 200

IrisQL

IrisQL is a query language for Iris Investigate that lets you express complex, multi-field searches in a single request. Pass the query as a raw string via the irisql parameter. The query must begin with # IrisQL-1.0.

query = """# IrisQL-1.0
DOMAIN CONTAINS "phishing"
AND
RISK_SCORE GREATER_THAN 85
"""

results = api.iris_investigate(irisql=query)
print(results["results_count"])
for domain in results:
    print(domain["domain"])

Pagination parameters (page_size, sort_by, position) are supported alongside IrisQL via **kwargs:

results = api.iris_investigate(irisql=query, page_size=50, sort_by="risk_score", position=0)

When irisql is set, any domain or filter parameters passed alongside it are silently ignored. IrisQL uses header-based authentication (X-Api-Key) automatically.

Using the API Asynchronously

domaintools Async Example

The DomainTools API automatically supports async usage:

search_results = await api.iris_enrich('domaintools.com').response().get('results', {})

There is built-in support for async context managers:

async with api.iris_enrich('domaintools.com').response().get('results', {}) as search_results:
    # do things

And direct async for loops:

async for result in api.iris_enrich('domaintools.com').response().get('results', {}):
    print(result)

All async operations can safely be intermixed with non async ones - with optimal performance achieved if the async call is done first:

profile = api.domain_profile('google.com')
await profile
title = profile['website_data']['title']

Interacting with the API via the command line client

domaintools CLI Example

Immediately after installing domaintools_api with pip, a domaintools command line client will become available to you:

domaintools --help

To use - simply pass in the api_call you would like to make along with the parameters that it takes and your credentials:

domaintools iris_investigate --domains domaintools.com -u $TEST_USER -k $TEST_KEY

Optionally, you can specify the desired format (html, xml, json, or list) of the results:

domaintools domain_search google --max_length 10 -u $TEST_USER -k $TEST_KEY -f html

IrisQL queries are supported via the --irisql flag on iris_investigate. The query must begin with # IrisQL-1.0 on its own line:

domaintools iris_investigate --irisql $'# IrisQL-1.0\nDOMAIN CONTAINS "phishing"' -u $TEST_USER -k $TEST_KEY

Pagination parameters can be passed alongside the IrisQL query:

domaintools iris_investigate --irisql $'# IrisQL-1.0\nDOMAIN CONTAINS "phishing"' --page-size 50 --sort-by risk_score -u $TEST_USER -k $TEST_KEY

To avoid having to type in your API key repeatedly, you can specify them in ~/.dtapi separated by a new line:

API_USER
API_KEY

Python Version Support Policy

Please see the supported versions document for the DomainTools Python support policy.

Real-Time Threat Feeds

Real-Time Threat Feeds provide data on the different stages of the domain lifecycle: from first-observed in the wild, to newly re-activated after a period of quiet. Access current feed data in real-time or retrieve historical feed data through separate APIs.

Custom parameters aside from the common GET Request parameters:

  • endpoint (choose either download or feed API endpoint - default is feed)
    api = API(USERNAME, KEY)
    api.nod(endpoint="feed", **kwargs)
    
  • header_authentication: by default, we're using API Header Authentication. Set this False if you want to use API Key and Secret Authentication. Apparently, you can't use API Header Authentication for download endpoints so this will be defaulted to False even without explicitly setting it.
    api = API(USERNAME, KEY, header_authentication=False)
    api.nod(**kwargs)
    
  • output_format: (choose either csv or jsonl - default is jsonl). Cannot be used in domainrdap feeds. Additionally, csv is not available for download endpoints.
    api = API(USERNAME, KEY)
    api.nod(output_format="csv", **kwargs)
    

The Feed API standard access pattern is to periodically request the most recent feed data, as often as every 60 seconds. Specify the range of data you receive in one of two ways:

  1. With sessionID: Make a call and provide a new sessionID parameter of your choosing. The API will return the last hour of data by default.
    • Each subsequent call to the API using your sessionID will return all data since the last.
    • Any single request returns a maximum of 10M results. Requests that exceed 10M results will return a HTTP 206 response code; repeat the same request (with the same sessionID) to receive the next tranche of data until receiving a HTTP 200 response code.
  2. Or, specify the time range in one of two ways:
    • Either an after=-60 query parameter, where (in this example) -60 indicates the previous 60 seconds.
    • Or after and before query parameters for a time range, with each parameter accepting an ISO-8601 UTC formatted timestamp (a UTC date and time of the format YYYY-MM-DDThh:mm:ssZ)

Feed parameters

The feed methods accept the following parameters, grouped by purpose. Availability depends on the feed (see the notes below the table).

Session Management Parameters

  • sessionID: A custom string used to distinguish between different sessions. Required when using fromBeginning.

  • after: Start of the query window. Either an integer offset relative to now in seconds (e.g. -60), or an absolute ISO 8601 UTC datetime (YYYY-MM-DDTHH:MM:SSZ).

  • before: End of the query window (inclusive). Either an integer from -1 to -432000 (seconds before now), or an absolute ISO 8601 UTC datetime. The query window covers at most the most recent 5 days; a value older than 5 days returns no records.

  • fromBeginning: Boolean (true/false/1/0, default false). Requires a valid sessionID. When true on the first request of a new session, returns the first hour of data in the time window instead of the last. Using it with an existing sessionID returns an HTTP 406; using it without a sessionID or with a non-boolean value returns an HTTP 422.

    api = API(USERNAME, KEY)
    api.nod(sessionID="my-new-session-id", after=-3600, fromBeginning=True)
    

Filter Parameters

  • domain: Filter for an exact domain or a substring contained within a domain by prefixing or suffixing your substring with *.

  • overall_min, malware_min, phishing_min, spam_min, proximity_min: Integer risk score thresholds (range 1 to 99, optional). Available on the realtime_domain_risk and domainhotlist feeds only. When multiple are supplied they act as a logical AND — a domain must meet ALL specified thresholds to be returned.

    api = API(USERNAME, KEY)
    api.domainhotlist(after=-3600, overall_min=70, phishing_min=50)
    
  • IP feed filters (available on the iprisk and iphotlist feeds only). All are optional integers/strings and combine as a logical AND:

    • Domain activity & volume: pdns_resolutions_min, bad_pdns_resolutions_min (positive integers, distinct/bad domains resolving to the IP in the last 24 hours) and total_domains_max (positive integer; caps total hosted domains to filter out superhosters like CDNs).
    • Threat intelligence & combined risk percentages: third_party_threats_min (positive integer), plus all_threats_combined_percent_min, combined_phishing_percent_min, combined_malware_percent_min, combined_spam_percent_min (percentages 0 to 100 of hosted domains confirmed or predicted malicious).
    • Confirmed threat percentages: all_threats_percent_min, percent_phishing_min, percent_malware_min, percent_spam_min (percentages 0 to 100 of hosted domains actively confirmed).
    • Infrastructure & geolocation: asn (integer, digits only — no AS prefix or wildcards), organization (exact name, no wildcards) and country_code (case-sensitive two-letter code, e.g. CN, US, NL).
    api = API(USERNAME, KEY)
    api.iprisk(after=-3600, bad_pdns_resolutions_min=5, total_domains_max=1000, country_code="US")
    

Result formatting parameters

  • output_format: csv or jsonl (default jsonl). Not available on the domainrdap feed. csv is not available for download endpoints.
  • headers: When csv output is used, adds a header row to the first line of the response.
  • top: Positive integer from 1 to 1,000,000,000 limiting the number of results in the response payload.

Handling iterative response from RTUF endpoints:

Since we may dealing with large feeds datasets, the python wrapper uses generator for efficient memory handling. Therefore, we need to iterate through the generator if we're accessing the partial results of the feeds data.

Single request because the requested data is within the maximum result:

from domaintools import API

api = API(USERNAME, KEY)
results = api.nod(sessionID="my-session-id", after=-60)

for result in results.response() # generator that holds NOD feeds data for the past 60 seconds and is expected to request only once
    # do things to result

Multiple requests because the requested data is more than the maximum result per request:

from domaintools import API

api = API(USERNAME, KEY)
results = api.nod(sessionID="my-session-id", after=-7200)

for partial_result in results.response() # generator that holds NOD feeds data for the past 2 hours and is expected to request multiple times
    # do things to partial_result

Running E2E Tests Locally

For now, e2e tests only covers proxy and ssl testing. We are expected to broaden our e2e tests to other scenarios moving forward. To add more e2e tests, put these in the ../tests/e2e folder.

Preparation

  • Create virtual environment.

        python3 -m venv venv
    
  • Activate virtual environment

        source venv/bin/activate
    
  • Install dependencies (with test extras):

        pip install -e ".[test]"
    

    Or without test dependencies:

        pip install -e .
    
  • Export api credentials to use.

        export TEST_USER=<user-key>
        export TEST_KEY=<api-key>
    
  • Run unit tests.

        tox -e
    

Run the end-to-end test script

  • Before running the test, be sure that docker is running.
  • Execute the e2e test script .
        sh tests/e2e/scripts/test_e2e_runner.sh
    

Release files for domaintools-api-test-version 2.10.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for domaintools-api-test-version 2.10.0.1
File Size Uploaded
domaintools_api_test_version-2.10.0.1.tar.gz 84.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for domaintools-api-test-version 2.10.0.1
File Interpreter ABI Platform
domaintools_api_test_version-2.10.0.1-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 156.0 kB

Release files / domaintools_api_test_version-2.10.0.1.tar.gz

Download URL domaintools_api_test_version-2.10.0.1.tar.gz
Size 84.3 kB
Tags Source
SHA-256 checksum
How to use checksums
17e6a0ac9586da1008f2d4b9b0a58b8c7fe26bab9fba8759ec042a1f6ee62f21
BLAKE2b-256 checksum
How to use checksums
6ac18438f094165fa24f5c596d377ad0e1e5fe78e2400679efa797ab5bc90fc9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.14

Release files / domaintools_api_test_version-2.10.0.1-py2.py3-none-any.whl

Download URL domaintools_api_test_version-2.10.0.1-py2.py3-none-any.whl
Size 71.6 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
42f469dcdd66ea332f9d4d7b72803e4bc72497fdc681ea70693b8c2df0fc24dd
BLAKE2b-256 checksum
How to use checksums
bda99f7b36c008652e35ff89d66d2196bf80494b0859b8ecf96fbd01c981e6be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

2.10.0.1 This release

2 release files

2.9.1

2 release files

2.9.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page