Skip to main content

Website Classification API

This is the Python3 client library for URL Classification service.

Our Website Categorization API provides accurate URL/webpage classification based on the widely trusted IAB Taxonomy.

Categorizations are done in real-time and using full-path URLs.You can also use our API to classify plain text.

We return categorizations of URLS for the following taxonomies:

IAB, version 3, with 4 Tiers (taxonomy from Internet Advertising Bureau - IAB,
IAB, version 2, with 4 Tiers (taxonomy from Internet Advertising Bureau - IAB
IPTC NewsCodes, especially suitable for News Categorization 
Web Content Filtering Taxonomy (44 categories)
Google Shopping Taxonomy, used by millions of online retailers (5474 categories)
Shopify Taxonomy, used by millions of online stores (10560 categories)
Amazon Taxonomy (39004 categories)

We also return the following classifications / enriched data about each URL:

Detection of Malware/Social Engineering/Mailicious Software
Web Technologies used
Likely Buyer Personas
Topics
Key Entities Named
Sentiment Analysis
Similar companies / Competitors
Similar domains
Tags
Keywords

For those looking for eCommerce classification, we also provide Smart Product Categorization AI. It supports Shopify, Google Shopping, eBay and 120 other marketplaces.

Website classification API s a python library that allows to classify websites based on IAB.

Installation

pip install websiteclassificationapi

Requirements

Only Python 3 is supported. You need an API key which you can obtain at our website. Python library requires only requests package.

Documentation

More detailed API documentation on URL Classification is available here.

Examples

Please check our API documentation page on our website for most up to date examples.

from websiteclassificationapi import websiteclassificationapi

api_key = 'h2XurA' # you can get API key from www.websitecategorizationapi.com
url = 'www.alpha-quantum.com' # can be set to any valid URL
classifier_type = 'iab1' # should be set to either iab1 (Tier 1 categorization) or iab2 (Tier 2 categorization) for general websites or ecommerce1, ecommerce2 and ecommerce3 for E-commerce or product websites

# calling the API
print(websiteclassificationapi.get_categorization(url,api_key,classifier_type))

How to select classifiers of different taxonomies

NEW (update October 2024): Our newest version of API supports classifications for up to 4 Tiers. It returns one or more of 700 IAB categories.

Classifier_type should be set to either iab1 (Tier 1 categorization) or iab2 (Tier 2 categorization) for general websites or ecommerce1, ecommerce2 and ecommerce3 for E-commerce or product websites.

IAB Tier 1 categorization returns probabilities of text being classified as one of 29 possible categories.

IAB Tier 2 categorization returns probabilities of text being classified as one of 447 possible categories.

Ecommerce Tier 1 categorization returns probabilities of text being classified as one of 21 possible categories.

Ecommerce Tier 2 website categorization returns probabilities of text being classified as one of 182 possible categories.

Ecommerce Tier 3 website categorization returns probabilities of text being classified as one of 1113 possible categories.

Taxonomies

The list of categories available by classifier is also known as Taxonomy. There are many taxonomies available, some are standard are well known, e.g. IAB taxonomy is well suited for ads and advertising in general, whereas Facebook product categories taxonomy is appropriate for ecommerce field.

Taxonomy also differ in how many tiers, levels, or depths do they support. E.g. taxonomy may only support 1 set of main categories, or it can further subcategories.

The categorization in the form of Tier 1/Tier 2/Tier 3/.... is also known as taxonomy path.

The classifiers can be either built in a way that they predict single Tier categories or they can return full taxonomy paths. It really depends on the use case what is most appropriate.

You can find more information about IAB taxonomy at this page: https://www.iab.com/guidelines/content-taxonomy/.

Taxonomy should be chosen in a way that it suits your use case. E.g. let us say you have an online store and currently you just list your products without any categorizations.

Then it may be very valuable if you could provide some kind of menus that categorize products in different verticals.

Why? Because your users may more easily find your products, you will have more subpages that can be indexed by search engines and thus provide you with more traffic and visits.

Having verticals set up may also mean better filtering and lead to higher conversions and thus lower cost of acquisition. There are a multitude of opportunities in adding categorization to an online store.

Our other services

Companies that manage internet access often start with a web filtering database to review how domains are categorized. This supports better decisions around policy cr eation, reporting, and risk visibility. Many of those teams also evaluate a web filtering service to apply those controls in production.

We also provide an Anonymization API for privacy-preserving media and data workflows.

Schools and libraries subject to the Children's Internet Protection Act need more than generic web categories — they need a purpose-built CIPA web filtering database with 120 million domains organized into 57+ content categories aligned to federal compliance requirements. Pairing categorization with threat intelligence strengthens any filtering deployment, and our phishing detection API delivers real-time lookups against 390,000+ DNS-verified active phishing domains so that malicious URLs are blocked before they reach end users. Together, these services give IT teams the classification depth and security coverage needed to protect students, staff, and network infrastructure.

AI explainability

One of the unique features of classifiers is that they provide machine learning interpretability or artificial intelligence explainability (XAI) in the form of words that most contribute to resulting classification.

Example 1 of explainability: Image1

Example 2 of explainability: Image1

Why the need of AI explainability?

AI models are increasingly being used in ways that affect humans. E.g. you may apply for a loan at the bank and get rejected, but even though a human may have sent or explained you this, the decision may have actually been made by a machine learning model.

Machine learning models making decisions is increasingly part of every day and because often these decisions are made by what could be termed black boxes, there is increasing desire for having ML decisions made in a way that are explainable.

There are also many regulations that demand this, e.g. GDPR.

Support for languages

Classification service supports classifications of websites in 150 languages.

Offline database of categorized domains

We offer offline URL database of millions of categorized domains. It can be used web content filtering, AdTech marketing, cybersecurity, brand safety, contextual targeting.

It is ideal for those use cases where you require very low latency of requests, which can be achieved with pre-classified websites stored in database.

Another great source for URL categorization database

Handling websites with no texts

When encountering websites that have no text and just images, our classifier relies on online optical character recognition API service to extract text (if any available) from images on the website. And then classify it.

To deal with potential duplicates we use the reverse IP lookup of domains to find similar domains that are hosted on the same IP.

Application of website categorization to technologies usage

We have collected usage of technologies by millions of websites, by combining this with categorization, one can find interesting results.

Here is for example usage of Intercom across industry verticals:

Image1

Based on 50 millions of usage points we built an AI recommender which can predict which technologies for company using a set of technologies.

Here are e.g. recommendations for company using Mouse Flow:

TechnologyAI Recommendation Score Website
AppNexus0.15http://appnexus.com
Microsoft Clarity0.14https://clarity.microsoft.com
Osano0.14https://www.osano.com/
Jetpack0.14https://jetpack.com
Raphael0.14https://dmitrybaranovskiy.github.io/raphael/
Svelte0.13https://svelte.dev
AWS Certificate Manager0.12https://aws.amazon.com/certificate-manager/
Extendify0.11https://extendify.com
Kendo UI0.11https://www.telerik.com/kendo-ui
Flywheel0.11https://getflywheel.com

Our digital solutions are trusted by organizations seeking reliable, secure, and innovative tools. Protect privacy in every scenario using our image anonymization platform, engineered to meet global standards and inspired by research from MIT, while staying informed with insights from Harvard and approaches found at Stanford.

Business growth is supported through our corporate enrichment data, delivering actionable intelligence much like the work at Cambridge. For cybersecurity and research teams, our domain dataset offers extensive categorization, paralleling the data-centric methods at Berkeley.

By aligning our solutions with the excellence demonstrated at Yale and Cornell, we ensure that your organization stays at the cutting edge—combining technology, research, and practical results for every challenge.

Example classifications

Example classification for website www.github.com:

{
  "classification": [
    {
      "category": "Technology & Computing",
      "value": 0.7621352908406164
    },
    {
      "category": "Business and Finance",
      "value": 0.0785701408756428
    },
    {
      "category": "Video Gaming",
      "value": 0.06626958968249749
    },
    {
      "category": "Fine Art",
      "value": 0.017105357862223433
    },
    {
      "category": "Hobbies & Interests",
      "value": 0.016812511656388394
    },
    {
      "category": "Sports",
      "value": 0.011396157737341801
    },
    {
      "category": "Home & Garden",
      "value": 0.009099685741207822
    },
    {
      "category": "Personal Finance",
      "value": 0.0076400890345109055
    },
    {
      "category": "News and Politics",
      "value": 0.006692288300928684
    },
    {
      "category": "Careers",
      "value": 0.0039930258544077606
    },
    {
      "category": "Automotive",
      "value": 0.0029276292555247764
    },
    {
      "category": "Events and Attractions",
      "value": 0.0026449624402393084
    },
    {
      "category": "Shopping",
      "value": 0.0023606962223306537
    },
    {
      "category": "Family and Relationships",
      "value": 0.0023174171750800186
    },
    {
      "category": "Music and Audio",
      "value": 0.0020517145262615513
    },
    {
      "category": "Movies",
      "value": 0.0018936850100483473
    },
    {
      "category": "Travel",
      "value": 0.0009448942095545797
    },
    {
      "category": "Science",
      "value": 0.0008432696857311802
    },
    {
      "category": "Pets",
      "value": 0.0006956402098649299
    },
    {
      "category": "Television",
      "value": 0.0005261918310662409
    },
    {
      "category": "Real Estate",
      "value": 0.0005058920662560916
    },
    {
      "category": "Religion & Spirituality",
      "value": 0.000492253420442475
    },
    {
      "category": "Healthy Living",
      "value": 0.0004690261931844088
    },
    {
      "category": "Medical Health",
      "value": 0.0004467617749304944
    },
    {
      "category": "Education",
      "value": 0.00036333686743226124
    },
    {
      "category": "Food & Drink",
      "value": 0.0003463620639422737
    },
    {
      "category": "Books and Literature",
      "value": 0.00027078317064036986
    },
    {
      "category": "Style & Fashion",
      "value": 0.00011770141998920516
    },
    {
      "category": "Pop Culture",
      "value": 0.00006764487171529734
    }
  ],
  "html": "29101",
  "language": "en",
  "status": 200
}

For teams building acceptable-use or data-loss-prevention policies around generative AI, the AI Tools Blocklist complements this classifier with a purpose-built feed of AI-tool domains grouped into functional categories (chatbots, code assistants, media generators, voice cloning, and more). It answers a question generic categorization cannot: not just "is this site AI-related," but "which kind of AI tool is it," which is what granular blocking policies require.

Website classification at massive scale also applies to discovering acquisition candidates across the full web. Our web-scale deal origination platform screens over 100 million classified domains against buyer acquisition theses, extracting 15 operational signals per company from their public web presence. Private equity firms, search funds, and corporate development teams use the pipeline to surface non-obvious targets whose websites reveal strong fit signals — leadership depth, recurring revenue indicators, compliance readiness — that conventional company databases never capture.

The same classification engine powers contextual audience segments for programmatic advertising without cookies. Our cookieless advertising data platform delivers domain datasets pre-categorized under the IAB content taxonomy, enabling DSPs and publishers to target audiences based on page context rather than user-level tracking. With Chrome, Safari, and Firefox deprecating third-party cookies, domain-level contextual classification is becoming the primary signal for brand-safe, regulation-compliant ad placement at scale.

Useful resources used in development of website categorization

Release files for websiteclassificationapi 2.15.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for websiteclassificationapi 2.15.1
File Size Uploaded
websiteclassificationapi-2.15.1.tar.gz 15.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for websiteclassificationapi 2.15.1
File Interpreter ABI Platform
websiteclassificationapi-2.15.1-py3-none-any.whl Python 3 none any Details

Total release size: 25.8 kB

Release files / websiteclassificationapi-2.15.1.tar.gz

Download URL websiteclassificationapi-2.15.1.tar.gz
Size 15.4 kB
Tags Source
SHA-256 checksum
How to use checksums
7dca62934b28d1ec85b8bf6389682e626a1125647767ded68a1e9feee225b33b
BLAKE2b-256 checksum
How to use checksums
06481e1b56e2a07c3b14277922a0c01d5e233fb4802cb91816e3198d22cacaaa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release files / websiteclassificationapi-2.15.1-py3-none-any.whl

Download URL websiteclassificationapi-2.15.1-py3-none-any.whl
Size 10.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ea70941a823436700d29dc1dce73f3702ce3fdc56c5ea5e248723db03ef66097
BLAKE2b-256 checksum
How to use checksums
4580da4d623540979873d82c027e1b8474cc9ae9c5471496e6a558073207fc04
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release history Release notifications | RSS feed

2.15.2

2 release files

This release

2.15.1 This release

2 release files

2.15

2 release files

2.14

2 release files

2.13

2 release files

2.12

2 release files

2.11

2 release files

2.10

2 release files

2.7

2 release files

2.6

2 release files

2.5

2 release files

2.4

2 release files

2.3

2 release files

2.2

2 release files

2.1.9

2 release files

2.1.8

2 release files

2.1.7

2 release files

2.1.5

2 release files

2.1.4

2 release files

2.1.3

2 release files

2.1.2

2 release files

2.1.1

2 release files

2.1.0

3 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page