Skip to main content

Wikipedia-API is easy to use Python wrapper for Wikipedias’ API. It supports extracting texts, sections, links, categories, translations, etc from Wikipedia. Documentation provides code snippets for the most common use cases.

GitHub stars Test Coverage Documentation Status Version Py Versions OpenSSF Best Practices

Installation

This package requires at least Python 3.9 to install because it’s using IntEnum.

pip3 install wikipedia-api

Usage

Goal of Wikipedia-API is to provide simple and easy to use API for retrieving informations from Wikipedia. The library provides both a synchronous (Wikipedia) and an asynchronous (AsyncWikipedia) client. Bellow are examples of common use cases.

Key differences between the sync and async API:

  • All data-fetching attributes (summary, text, sections, langlinks, links, backlinks, categories, categorymembers, pageid, fullurl, displaytitle, …) are explicit @property definitions in both APIs. In the async API every such property returns a coroutine: await page.summary, await page.sections, await page.pageid, etc.

  • title, ns, namespace, language, variant are plain @property values in both APIs (no await needed).

  • exists() is a plain method in the sync API; a coroutine method in the async API: await page.exists().

  • section_by_title() and sections_by_title() are plain synchronous methods in both APIs.

Importing

import wikipediaapi

# Synchronous client
wiki = wikipediaapi.Wikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

# Asynchronous client
wiki = wikipediaapi.AsyncWikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

How To Get Single Page

Getting single page is straightforward. You have to initialize Wikipedia (or AsyncWikipedia) object and ask for page by its name. To initialize it, you have to provide:

  • user_agent to identify your project. Please follow the recommended format.

  • language to specify language mutation. It has to be one of supported languages.

Synchronous

import wikipediaapi
wiki_wiki = wikipediaapi.Wikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

page_py = wiki_wiki.page('Python_(programming_language)')

Asynchronous

import asyncio
import wikipediaapi

async def main():
    wiki_wiki = wikipediaapi.AsyncWikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')
    page_py = wiki_wiki.page('Python_(programming_language)')
    # Data is fetched lazily — await any attribute or property to trigger it

asyncio.run(main())

How To Check If Wiki Page Exists

For checking, whether page exists, you can use function exists.

Synchronous

page_py = wiki_wiki.page('Python_(programming_language)')
print("Page - Exists: %s" % page_py.exists())
# Page - Exists: True

page_missing = wiki_wiki.page('NonExistingPageWithStrangeName')
print("Page - Exists: %s" % page_missing.exists())
# Page - Exists: False

Asynchronous

In the async API, exists() is a coroutine — it lazily fetches pageid via the info API call if not yet cached (same approach as await page.fullurl).

async def main():
    page_py = wiki_wiki.page('Python_(programming_language)')
    print("Page - Exists: %s" % await page_py.exists())
    # Page - Exists: True

    page_missing = wiki_wiki.page('NonExistingPageWithStrangeName')
    print("Page - Exists: %s" % await page_missing.exists())
    # Page - Exists: False

How To Get Page Summary

Class WikipediaPage has property summary, which returns description of Wiki page. In the async API, summary is a coroutine.

Synchronous

import wikipediaapi
wiki_wiki = wikipediaapi.Wikipedia('MyProjectName (merlin@example.com)', 'en')

print("Page - Title: %s" % page_py.title)
# Page - Title: Python (programming language)

print("Page - Summary: %s" % page_py.summary[0:60])
# Page - Summary: Python is a widely used high-level programming language for

Asynchronous

async def main():
    print("Page - Title: %s" % page_py.title)
    # Page - Title: Python (programming language)

    summary = await page_py.summary
    print("Page - Summary: %s" % summary[0:60])
    # Page - Summary: Python is a widely used high-level programming language for

How To Get Page URL

WikipediaPage has two properties with URL of the page. It is fullurl and canonicalurl. In the async API, these attributes are awaitables.

Synchronous

print(page_py.fullurl)
# https://en.wikipedia.org/wiki/Python_(programming_language)

print(page_py.canonicalurl)
# https://en.wikipedia.org/wiki/Python_(programming_language)

Asynchronous

async def main():
    print(await page_py.fullurl)
    # https://en.wikipedia.org/wiki/Python_(programming_language)

    print(await page_py.canonicalurl)
    # https://en.wikipedia.org/wiki/Python_(programming_language)

How To Get Full Text

To get full text of Wikipedia page you should use property text which constructs text of the page as concatanation of summary and sections with their titles and texts.

Synchronous

wiki_wiki = wikipediaapi.Wikipedia(
    user_agent='MyProjectName (merlin@example.com)',
    language='en',
    extract_format=wikipediaapi.ExtractFormat.WIKI
)

p_wiki = wiki_wiki.page("Test 1")
print(p_wiki.text)
# Summary
# Section 1
# Text of section 1
# Section 1.1
# Text of section 1.1
# ...


wiki_html = wikipediaapi.Wikipedia(
    user_agent='MyProjectName (merlin@example.com)',
    language='en',
    extract_format=wikipediaapi.ExtractFormat.HTML
)
p_html = wiki_html.page("Test 1")
print(p_html.text)
# <p>Summary</p>
# <h2>Section 1</h2>
# <p>Text of section 1</p>
# <h3>Section 1.1</h3>
# <p>Text of section 1.1</p>
# ...

Asynchronous

async def main():
    wiki_wiki = wikipediaapi.AsyncWikipedia(
        user_agent='MyProjectName (merlin@example.com)',
        language='en',
        extract_format=wikipediaapi.ExtractFormat.WIKI
    )
    page = wiki_wiki.page("Test 1")
    text = await page.text
    print(text)
    # Summary
    # Section 1
    # Text of section 1
    # Section 1.1
    # Text of section 1.1
    # ...

How To Get Page Sections

To get all top level sections of page, you have to use property sections. It returns list of WikipediaPageSection, so you have to use recursion to get all subsections.

Synchronous

def print_sections(sections, level=0):
    for s in sections:
        print("%s: %s - %s" % ("*" * (level + 1), s.title, s.text[0:40]))
        print_sections(s.sections, level + 1)


print_sections(page_py.sections)
# *: History - Python was conceived in the late 1980s,
# *: Features and philosophy - Python is a multi-paradigm programming l
# *: Syntax and semantics - Python is meant to be an easily readable
# **: Indentation - Python uses whitespace indentation, rath
# **: Statements and control flow - Python's statements include (among other
# **: Expressions - Some Python expressions are similar to l

Asynchronous

def print_sections(sections, level=0):
    for s in sections:
        print("%s: %s - %s" % ("*" * (level + 1), s.title, s.text[0:40]))
        print_sections(s.sections, level + 1)

async def main():
    sections = await page_py.sections
    print_sections(sections)
    # *: History - Python was conceived in the late 1980s,
    # *: Features and philosophy - Python is a multi-paradigm programming l
    # *: Syntax and semantics - Python is meant to be an easily readable

How To Get Page Section By Title

To get last section of page with given title, you have to use function section_by_title. It returns the last WikipediaPageSection with this title. section_by_title works the same in both the sync and async API.

section_history = page_py.section_by_title('History')
print("%s - %s" % (section_history.title, section_history.text[0:40]))

# History - Python was conceived in the late 1980s b

How To Get All Page Sections By Title

To get all sections of page with given title, you have to use function sections_by_title. It returns the all WikipediaPageSection with this title. sections_by_title works the same in both the sync and async API.

page_1920 = wiki_wiki.page('1920')
sections_january = page_1920.sections_by_title('January')
for s in sections_january:
    print("* %s - %s" % (s.title, s.text[0:40]))

# * January - January 1
# Polish–Soviet War in 1920: The
# * January - January 2
# Isaac Asimov, American author
# * January - January 1 – Zygmunt Gorazdowski, Polish

How To Get Page In Other Languages

If you want to get other translations of given page, you should use property langlinks. It is map, where key is language code and value is WikipediaPage (or AsyncWikipediaPage).

Synchronous

def print_langlinks(page):
    langlinks = page.langlinks
    for k in sorted(langlinks.keys()):
        v = langlinks[k]
        print("%s: %s - %s: %s" % (k, v.language, v.title, v.fullurl))

print_langlinks(page_py)
# af: af - Python (programmeertaal): https://af.wikipedia.org/wiki/Python_(programmeertaal)
# als: als - Python (Programmiersprache): https://als.wikipedia.org/wiki/Python_(Programmiersprache)
# an: an - Python: https://an.wikipedia.org/wiki/Python
# ar: ar - بايثون: https://ar.wikipedia.org/wiki/%D8%A8%D8%A7%D9%8A%D8%AB%D9%88%D9%86
# as: as - পাইথন: https://as.wikipedia.org/wiki/%E0%A6%AA%E0%A6%BE%E0%A6%87%E0%A6%A5%E0%A6%A8

page_py_cs = page_py.langlinks['cs']
print("Page - Summary: %s" % page_py_cs.summary[0:60])
# Page - Summary: Python (anglická výslovnost [ˈpaiθtən]) je vysokoúrovňový sk

Asynchronous

In the async API, langlinks is an awaitable property. Attributes on the returned page stubs (e.g. fullurl) are also awaitables.

async def main():
    langlinks = await page_py.langlinks
    for k in sorted(langlinks.keys()):
        v = langlinks[k]
        print("%s: %s - %s: %s" % (k, v.language, v.title, await v.fullurl))

    page_py_cs = langlinks['cs']
    print("Page - Summary: %s" % (await page_py_cs.summary)[0:60])
    # Page - Summary: Python (anglická výslovnost [ˈpaiθtən]) je vysokoúrovňový sk

How To Get Page Categories

If you want to get all categories under which page belongs, you should use property categories. It’s map, where key is category title and value is WikipediaPage (or AsyncWikipediaPage).

Synchronous

def print_categories(page):
    categories = page.categories
    for title in sorted(categories.keys()):
        print("%s: %s" % (title, categories[title]))


print("Categories")
print_categories(page_py)
# Category:All articles containing potentially dated statements: ...
# Category:All articles with unsourced statements: ...
# Category:Articles containing potentially dated statements from August 2016: ...
# Category:Articles containing potentially dated statements from March 2017: ...
# Category:Articles containing potentially dated statements from September 2017: ...

Asynchronous

async def main():
    categories = await page_py.categories
    for title in sorted(categories.keys()):
        print("%s: %s" % (title, categories[title]))

How To Get All Pages From Category

To get all pages from given category, you should use property categorymembers (sync) or awaitable property categorymembers (async). It returns all members of given category. You have to implement recursion and deduplication by yourself.

Synchronous

def print_categorymembers(categorymembers, level=0, max_level=1):
    for c in categorymembers.values():
        print("%s: %s (ns: %d)" % ("*" * (level + 1), c.title, c.ns))
        if c.ns == wikipediaapi.Namespace.CATEGORY and level < max_level:
            print_categorymembers(c.categorymembers, level=level + 1, max_level=max_level)


cat = wiki_wiki.page("Category:Physics")
print("Category members: Category:Physics")
print_categorymembers(cat.categorymembers)

# Category members: Category:Physics
# * Statistical mechanics (ns: 0)
# * Category:Physical quantities (ns: 14)
# ** Refractive index (ns: 0)
# ** Vapor quality (ns: 0)
# ** Electric susceptibility (ns: 0)
# ** Specific weight (ns: 0)
# ** Category:Viscosity (ns: 14)
# *** Brookfield Engineering (ns: 0)

Asynchronous

async def print_categorymembers(categorymembers, level=0, max_level=1):
    for c in categorymembers.values():
        print("%s: %s (ns: %d)" % ("*" * (level + 1), c.title, c.ns))
        if c.ns == wikipediaapi.Namespace.CATEGORY and level < max_level:
            await print_categorymembers(
                await c.categorymembers, level=level + 1, max_level=max_level
            )

async def main():
    cat = wiki_wiki.page("Category:Physics")
    print("Category members: Category:Physics")
    await print_categorymembers(await cat.categorymembers)

Use Extra API Parameters

Official API supports many different parameters. You can see them in the sandbox. Not all these parameters are supported directly as parameters of the functions. If you want to specify them, you can pass them as additional parameters in the constructor. For the info API call you can specify parameter converttitles. If you want to specify it, you can use:

Synchronous

import wikipediaapi
wiki_wiki = wikipediaapi.Wikipedia('MyProjectName (merlin@example.com)', 'zh', 'zh-tw', extra_api_params={'converttitles': 1})
page = wiki_wiki.page("孟卯")
print(repr(page.varianttitles))

Asynchronous

async def main():
    wiki_wiki = wikipediaapi.AsyncWikipedia('MyProjectName (merlin@example.com)', 'zh', 'zh-tw', extra_api_params={'converttitles': 1})
    page = wiki_wiki.page("孟卯")
    print(repr(await page.varianttitles))

Error Handling

All exceptions raised by the library inherit from WikipediaException. You can catch specific exceptions or the base WikipediaException. The same exception types are raised by both the sync and async clients.

Synchronous

import wikipediaapi

wiki_wiki = wikipediaapi.Wikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

# Catch any Wikipedia-API error
try:
    page = wiki_wiki.page('Python_(programming_language)')
    print(page.summary[0:60])
except wikipediaapi.WikipediaException as e:
    print("Error: %s" % e)
# Handle specific error types
try:
    page = wiki_wiki.page('Python_(programming_language)')
    print(page.summary[0:60])
except wikipediaapi.WikiRateLimitError as e:
    print("Rate limited! Retry after: %s seconds" % e.retry_after)
except wikipediaapi.WikiHttpError as e:
    print("HTTP error %d: %s" % (e.status_code, e))
except wikipediaapi.WikiHttpTimeoutError:
    print("Request timed out")
except wikipediaapi.WikiConnectionError:
    print("Could not connect to Wikipedia")
except wikipediaapi.WikiInvalidJsonError:
    print("Received invalid response from Wikipedia")

Asynchronous

async def main():
    wiki_wiki = wikipediaapi.AsyncWikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

    try:
        page = wiki_wiki.page('Python_(programming_language)')
        print((await page.summary)[0:60])
    except wikipediaapi.WikiRateLimitError as e:
        print("Rate limited! Retry after: %s seconds" % e.retry_after)
    except wikipediaapi.WikiHttpError as e:
        print("HTTP error %d: %s" % (e.status_code, e))
    except wikipediaapi.WikiHttpTimeoutError:
        print("Request timed out")
    except wikipediaapi.WikiConnectionError:
        print("Could not connect to Wikipedia")
    except wikipediaapi.WikiInvalidJsonError:
        print("Received invalid response from Wikipedia")

Retry Configuration

By default, transient errors (HTTP 429, 5xx, timeouts, connection errors) are retried up to 3 times with exponential backoff. You can configure this behavior in the constructor. The same options apply to both Wikipedia and AsyncWikipedia.

import wikipediaapi

# Custom retry: 5 retries with 2-second base wait
wiki_wiki = wikipediaapi.Wikipedia(
    user_agent='MyProjectName (merlin@example.com)',
    language='en',
    max_retries=5,
    retry_wait=2.0,
)
# Disable retries entirely
wiki_wiki = wikipediaapi.Wikipedia(
    user_agent='MyProjectName (merlin@example.com)',
    language='en',
    max_retries=0,
)

How To See Underlying API Call

If you have problems with retrieving data you can get URL of undrerlying API call. This will help you determine if the problem is in the library or somewhere else. Logging works the same for both Wikipedia and AsyncWikipedia.

import sys

import wikipediaapi
wikipediaapi.log.setLevel(level=wikipediaapi.logging.DEBUG)

# Set handler if you use Python in interactive mode
out_hdlr = wikipediaapi.logging.StreamHandler(sys.stderr)
out_hdlr.setFormatter(wikipediaapi.logging.Formatter('%(asctime)s %(message)s'))
out_hdlr.setLevel(wikipediaapi.logging.DEBUG)
wikipediaapi.log.addHandler(out_hdlr)

wiki = wikipediaapi.Wikipedia(user_agent='MyProjectName (merlin@example.com)', language='en')

page_ostrava = wiki.page('Ostrava')
print(page_ostrava.summary)
# logger prints out: Request URL: http://en.wikipedia.org/w/api.php?action=query&prop=extracts&titles=Ostrava&explaintext=1&exsectionformat=wiki

Other Badges

Version Py Versions Implementations Downloads Tags github-release Github commits (since latest release) GitHub forks GitHub stars GitHub watchers GitHub commit activity the past week, 4 weeks, year Last commit GitHub code size in bytes GitHub repo size in bytes PyPi License PyPi Wheel PyPi Format PyPi PyVersions PyPi Implementations PyPi Status PyPi Downloads - Day PyPi Downloads - Week PyPi Downloads - Month Libraries.io - SourceRank Libraries.io - Dependent Repos Coveralls

Other Pages

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wikipedia_api-0.11.0.tar.gz (69.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

wikipedia_api-0.11.0-py3-none-any.whl (47.8 kB view details)

Uploaded Python 3

File details

Details for the file wikipedia_api-0.11.0.tar.gz.

File metadata

  • Download URL: wikipedia_api-0.11.0.tar.gz
  • Upload date:
  • Size: 69.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for wikipedia_api-0.11.0.tar.gz
Algorithm Hash digest
SHA256 27224fd3a3329a621a460f59a72b4b063acda7d1ed84e5623abdd01cb907ad64
MD5 47fe7941ba3e8241a95748f1301118e7
BLAKE2b-256 f9739b1981012c2a431c215aada7125af5d9c3610d6ddcf5ad336d1f0efad1da

See more details on using hashes here.

Provenance

The following attestation bundles were made for wikipedia_api-0.11.0.tar.gz:

Publisher: release.yml on martin-majlis/Wikipedia-API

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file wikipedia_api-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: wikipedia_api-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 47.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for wikipedia_api-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a77abae1137af015ad4efbb5ca29824ac5739a6fa6b2437fef9a2b26e4ad2a5a
MD5 2d5249ac41d182d82e0d089310be2243
BLAKE2b-256 c1b8b36a1793821f5b3eafdd5de45e0777bb89481a17b6b0c9ef9ad051ff0c3b

See more details on using hashes here.

Provenance

The following attestation bundles were made for wikipedia_api-0.11.0-py3-none-any.whl:

Publisher: release.yml on martin-majlis/Wikipedia-API

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page