Skip to main content

detector

Offline IP intelligence SDK for Python 3.9+.

One call returns every field of every bundled database, merged into a single standardised English document — no API keys, no network, no extra downloads. Distance between IPs is one call away, 1-to-1 or 1-to-N (N unbounded).

from detector import info, distance      # two coroutines, nothing else

await info("8.8.8.8")                    # -> dict   (the standard JSON document)
await info(["8.8.8.8", "1.1.1.1"])       # -> [dict, dict]      (same function)
await info("8.8.8.8", as_object=True)    # -> IPInfo            (model, if you want it)
await distance("8.8.8.8", "1.1.1.1")     # -> dict
await distance("8.8.8.8", ["1.1.1.1", "223.5.5.5"])   # -> [dict, dict]
  • 11 datasets bundled (every dataset ip-location-db publishes), ~90 MB packaged, read through memory-mapped MMDB files
  • Multi-source merge with cross_check / agreement / conflicts so you can see which source said what before trusting a value
  • Nothing is dropped: unmapped vendor fields land in traits, complete per-database records land in raw
  • Two async functions: info and distance, each taking one address or a batch, each returning the standard JSON document (as_object=True for models)
  • JSON by default, stable field names, null for missing values, English only
  • Dependencies: maxminddb only (stdlib everywhere else)

Install

pip install detector-sdk        # distribution name; the import name is `detector`
pip install .                   # from a checkout

The databases ship inside the wheel, split into independently compressed parts. First use decompresses and merges them into ~/.cache/detector/extracted (override with DETECTOR_CACHE_DIR) — 251 MB of plain MMDB files, one write pass, parts decoded in parallel. Pay that cost up front with await warmup(): MMDB across 12 files. Reading is memory-mapped, so RAM stays low (~30 MB with everything loaded).

Licence warning : GeoLite2 data is bundled here because this build was asked to ship everything ip-location-db publishes. MaxMind's EULA forbids redistributing their data to third parties, so publishing this wheel (or an image containing it) to the public requires your own MaxMind entitlement. Detector(datasets=[...]) / exclude=[...] let you drop those three datasets, and NON_REDISTRIBUTABLE_KEYS lists them.

Quick start

Two coroutines, each polymorphic over single and batch input:

import asyncio
from detector import info, distance

async def main():
    # information: the return value IS the standard JSON document
    row = await info("8.8.8.8")
    row["country"]["iso_code"]       # 'US'
    row["city"]["name"]              # 'Mountain View'
    row["asn"]["organization"]       # 'Google LLC'
    row["location"]["latitude"]      # 37.422
    row["cross_check"]["country"]    # what each of the 7 country sources answered
    row["conflicts"]                 # ['asn_organization', 'location']
    json.dumps(row)                  # already JSON-serialisable

    many = await info(["8.8.8.8", "1.1.1.1", "114.114.114.114"])
    [item["country"]["iso_code"] for item in many]   # ['US', 'AU', 'CN']

    streaming = await info(generator_of_millions)    # windowed, bounded memory

    model = await info("8.8.8.8", as_object=True)    # IPInfo, attribute access
    model.country.iso_code, model.city.name(), model.coordinates

    # distance: same shape
    await distance("8.8.8.8", "1.1.1.1")                    # dict
    await distance("8.8.8.8", ["1.1.1.1", "223.5.5.5"])     # [dict, dict]
    await distance(["8.8.8.8"], ["1.1.1.1", "223.5.5.5"])   # N x M
    await distance("8.8.8.8", "1.1.1.1", method="vincenty") # WGS84 ellipsoid

asyncio.run(main())

Options go per call, or once:

from detector import configure
configure(datasets=["dbip-city", "dbip-asn"], locales=("en",), include_raw=False,
          cache_size=8192, max_concurrency=64)

The same two functions also accept the JSON envelope ({"type","action","data","status"}), so a queue/socket consumer needs no extra entry point:

await info({"type": "ipv4", "action": "info", "data": {"ip": "8.8.8.8"}})      # -> dict
await distance({"type": "ipv4", "action": "distance",
                "data": {"ip": "8.8.8.8", "list": ["1.1.1.1"]}})               # -> dict
# as_object=True gives the Response model instead

Startup, readiness and loading policy

First use unpacks ~90 MB of compressed parts into 251 MB of MMDB files, once per machine. That work is progressive (the cheapest databases answer first) and runs in a background thread, so it never has to sit on your critical path:

DETECTOR_INIT=import            # start unpacking as soon as `import detector` runs
DETECTOR_INIT=blocking          # finish before the import returns
DETECTOR_WAIT_TIMEOUT=1.0       # default deadline for a first query (seconds)
DETECTOR_ON_TIMEOUT=partial     # answer with what's ready (or `error`)
import asyncio, detector
from detector import Detector, ready, progress, info

# explicit control:
await detector.ready(2.0)       # True if everything is unpacked within 2s
await detector.progress()       # {"ready": 12, "total": 12, "failed": {}, ...}

# bound the first query instead of guessing:
doc = await info("8.8.8.8",
                 wait_timeout=1.0, on_timeout="partial")
doc["meta"]["preparation"]      # {"ready": ..., "total": ..., "loading": bool, ...}

# or ask the client directly:
d = Detector(wait_timeout=1.0, on_timeout="partial")   # never blocks past 1s
await d.ready(); d.refresh()    # reopen after update_datasets()

Semantics:

  • wait="all" (default) - a query blockes until every dataset is ready. With a wait_timeout and on_timeout="error" it raises LoadingTimeoutError (.code == "loading_timeout", with ready/total/missing/retry_after) instead of hanging; the unpacking keeps going, so a retry succeeds.
  • on_timeout="partial" - returns immediately with what is ready and marks the answer meta.preparation.loading: true; later queries automatically see the fuller dataset as the background thread finishes (meta.preparation.complete).
  • A dataset that fails to unpack is reported in meta.preparation.failed / progress()["failed"] - the healthy databases still answer. strict=True makes any failure an error. If nothing at all can be unpacked, a NoDatabaseError is raised.
  • import detector never touches the data. Use DETECTOR_INIT=import (or the await warmup() hook) to overlap it with your own start-up work.

Distance result

row = await distance("8.8.8.8", "1.1.1.1")
row["distance_km"]      # 11953.88   (haversine, default)
row["distance_mi"]      # 7427.79
row["same_country"]     # False
row["same_asn"]         # False
row["reason"]           # None when it could be computed

model = await distance("8.8.8.8", "1.1.1.1", as_object=True)
model.km, float(model), f"{model}"     # 11953.88, 11953.88, '8.8.8.8 -> 1.1.1.1: 11,953.88 km'

Explicit clients

When you want the client object instead of the pooled one (max_concurrency, window, nearest, describe, stats, update_datasets):

from detector import AsyncDetector

async with await AsyncDetector.create(max_concurrency=64, window=1024) as client:
    rows = await client.lookup_many(ips)
    ranked = await client.nearest("223.5.5.5", candidates, limit=3)
    print(await client.describe())

The standard output document

{
  "ip": "8.8.8.8",
  "version": 4,
  "found": true,
  "network": "8.8.8.0/24",
  "networks": {"dbip-city": "8.8.8.0/24", "dbip-country": "8.8.0.0/17"},
  "flags": {"is_private": false, "is_global": true, "is_loopback": false,
            "is_reserved": false, "is_multicast": false, "is_unspecified": false},
  "continent": {"name": "North America", "names": {"en": "North America", "...": "..."},
                "geoname_id": 6255149, "code": "NA"},
  "country": {"name": "United States", "names": {"en": "United States"},
              "geoname_id": 6252001, "iso_code": "US", "is_in_eu": false},
  "subdivisions": [{"name": "California", "names": {}, "geoname_id": null, "iso_code": null}],
  "city": {"name": "Mountain View", "names": {"en": "Mountain View"}, "geoname_id": null},
  "location": {"latitude": 37.422, "longitude": -122.085,
               "time_zone": null, "accuracy_radius": null},
  "postal": null,
  "time_zone": null,
  "asn": {"number": 15169, "asn": "AS15169", "organization": "Google LLC"},
  "traits": {},
  "hostname": null,
  "display": "Mountain View, California, United States",
  "sources": {"dbip-city": "DBIP-City-Lite", "dbip-asn": "DBIP-ASN-Lite",
              "iptoasn-asn": "IPtoASN-ASN", "...": "..."},
  "cross_check": {"country": {"DBIP-City-Lite": "US", "User-Country": "US"},
                  "asn": {"DBIP-ASN-Lite": 15169, "IPtoASN-ASN": 15169}},
  "agreement": {"country": true, "asn": true},
  "conflicts": [],
  "raw": {"dbip-city": {"city": {"names": {"en": "Mountain View"}}, "...": "..."}},
  "meta": {"schema_version": "1.0", "locales": ["en"],
           "databases": [{"key": "dbip-city", "name": "DBIP-City-Lite",
                          "build_date": "2026-09-01", "license": "CC BY 4.0"}],
           "attribution": "IP Geolocation by DB-IP (https://db-ip.com)"}
}

Want a different locale order? await info("8.8.8.8", locales=("zh-CN", "en")) — the names dictionaries always carry every language the source provides.

JSON protocol

The two functions take the envelope directly - info() for action=info, distance() for action=distance - and return a Response (.ok, .status, .data, .error, .meta, .to_dict(), .to_json()):

response = await distance({"type": "ipv4", "action": "distance",
                           "data": {"ip": "8.8.8.8", "list": ["1.1.1.1", "223.5.5.5"]}})
response.ok        # True
response.data      # {"ip": ..., "source": {...}, "list": [...], "summary": {...}}

await info('{"type":"ipv6","action":"info","data":{"ip":"2001:4860:4860::8888"}}')
await info(envelope, as_dict=True)     # plain dict instead of a Response

action aliases (lookup, query, dist, ...) are accepted on input; the response always echoes the canonical English action. Declaring type: ipv4 while sending an IPv6 literal is an error, not a silent mismatch.

Datasets

key name kind licence files bundled cadence
dbip-city DBIP-City-Lite city, subdivision, country, coords CC BY 4.0 1 yes monthly
dbip-asn DBIP-ASN-Lite ASN CC BY 4.0 1 yes monthly
dbip-country DBIP-Country-Lite country CC BY 4.0 1 yes monthly
geolite2-city GeoLite2-City city, subdivision, country, coords MaxMind EULA 2 (v4/v6) yes ⚠ twice weekly
geolite2-asn GeoLite2-ASN ASN MaxMind EULA 1 yes ⚠ twice weekly
geolite2-country GeoLite2-Country country MaxMind EULA 1 yes ⚠ twice weekly
iptoasn-asn IPtoASN-ASN ASN PDDL 1 yes daily
iptoasn-country IPtoASN-Country country PDDL 1 yes daily
origin-asn Origin-ASN ASN PDDL 1 yes daily
user-country User-Country country PDDL 1 yes daily
server-country Server-Country country PDDL 1 yes daily

⚠ = bundled but not redistributable: MaxMind's EULA does not permit passing their data on to third parties. Drop them with Detector(exclude=NON_REDISTRIBUTABLE_KEYS) if you are publishing.

Files ship as .mmdb.partNNN.xz / .mmdb.partNNN.zst (xz/zstd; a split database is decoded part-by-part on a thread pool and merged back into one file whose sha256 is recorded in MANIFEST.json). .mmdb.xz (whole, xz beats gzip by ~35% on MMDB data and the standard-library lzma module unpacks it). Loaded files also carry their address family, so a v6-only GeoLite2 half is never consulted for an IPv4 address.

You can always ignore the bundled copies and point the SDK at files you are licensed to hold:

from detector import Detector

detector = Detector(databases={
    "my-geolite2-city": "/data/GeoLite2-City.mmdb",
    "my-geoip2-isp":    "/data/GeoIP2-ISP.mmdb",
})

Any MMDB v2.0 file works — MaxMind, DB-IP, ip-location-db, or one you built with mmdb_writer. Unknown schemas still contribute: unmapped keys end up in traits, the whole record in raw.

Selecting datasets

Detector(datasets=["dbip-city", "iptoasn-asn"])          # only these
Detector(exclude=["user-country", "server-country"])     # all but these
Detector(datasets=["dbip-city"], include_raw=False)      # smaller payloads

Updating

import asyncio
from detector import update_datasets, update_datasets_async, known_datasets

known_datasets()                     # full registry: licences, mirrors, cadence

update_datasets("~/.cache/detector") # blocking
await update_datasets_async(         # concurrent asyncio downloads
    "~/.cache/detector",
    datasets=["dbip-city", "iptoasn-asn"],
    concurrency=4,
)

Fresh files in the cache directory override the bundled copies automatically. DETECTOR_DB_DIR repoints the whole bundled directory instead.

Extension points

Need How
Extra/vendor databases await info(ip, databases={...}), or drop .mmdb into DETECTOR_DB_DIR
Custom distance metric await distance(a, b, method="vincenty"), or compute from row.coordinates
Own JSON shape info.to_dict() / to_flat_dict() and rebuild
Caching strategy await info(ip, cache_size=N) (0 disables)
Cleaner output await info(ip, include_raw=False, locales=("en",))
Process-wide defaults detector.configure(...)

Performance notes

  • Lookups are memory-mapped: 12 databases open cost ~30 MB RSS, and the 127 MB city database never lands in your heap.
  • Warm startup is ~10 ms; the very first run decompresses the bundled data (12.6 s on a 2-core box with a 25 MB/s disk, ~3 s on 8 cores + NVMe) (~13 s).
  • A default lookup is ~477 µs (2.1k IP/s) over 11 datasets / 12 files; ~40 µs when the address is already in the LRU cache.
  • Generators keep memory flat for arbitrarily long inputs: await info(gen); include_raw=False trims documents from ~9.8 KB to ~5.8 KB.
  • Threads and multiprocessing inside the SDK were removed on purpose - both measured slower than sequential. Shard the input across processes instead (see docs/07-performance.md).
  • Install the C extension of maxminddb (libmaxminddb-dev present at install time) for 3-5x faster raw lookups.

Documentation

file contents
docs/01-getting-started.md install, first lookup, first distance, first async call
docs/02-api-reference.md every class, method and helper
docs/03-output-schema.md the standard document field by field
docs/04-distance.md maths, methods, accuracy reality check
docs/05-json-protocol.md request/response envelopes and error codes
docs/06-datasets.md datasets, licences, updates, custom MMDBs
docs/07-performance.md measured numbers and tuning
docs/08-troubleshooting.md when something looks wrong
docs/09-architecture.md module map, data flow, design decisions

Runnable examples live in examples/ (quickstart.py, async_concurrency.py, json_protocol.py, bulk_streaming.py, custom_database.py, show_output.py).

Licence

MIT for the code. Data licences and mandatory attribution are documented in NOTICE and reproduced automatically in every response under meta.attribution.

Release files for detector-sdk 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for detector-sdk 0.3.0
File Size Uploaded
detector_sdk-0.3.0.tar.gz 93.9 MB Details

Release files / detector_sdk-0.3.0.tar.gz

Download URL detector_sdk-0.3.0.tar.gz
Size 93.9 MB
Tags Source
SHA-256 checksum
How to use checksums
8f68b019c682e45b52e2ff5749155705ea6f0cd8783d6ae99848085109926c99
BLAKE2b-256 checksum
How to use checksums
f03ec570ac117fc70ff30cf644dde9d816e45db093d757167f31089520580b0a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

0.4.1

2 release files

0.4.0

2 release files

This release

0.3.0 This release

1 release file

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page