Skip to main content

Glean Indexing SDK

GA PyPI version

Build custom Glean connectors in Python. The SDK handles fetching, transforming, batching, and uploading your data to Glean's indexing APIs, so you write only the parts that are specific to your source.

📖 Full documentation on the Glean Developer site.

Repository examples cover advanced usage and streaming connectors.

Build one with an agent

The fastest path is to let a coding agent build the connector. Your agent installs it straight from this repository, which doubles as the plugin marketplace.

Claude Code

claude plugin marketplace add gleanwork/glean-indexing-sdk
claude plugin install glean-connector-builder@glean-indexing-sdk

Codex

codex plugin marketplace add gleanwork/glean-indexing-sdk
codex plugin add glean-connector-builder@glean-indexing-sdk

Cursor has no plugin CLI — open Dashboard → Plugins → Add Marketplace → Import from Repo, point it at gleanwork/glean-indexing-sdk, then install from Customize.

Then describe your source:

I want to push my Webex data to Glean. Build a connector for me.

The agent explores the source's API, confirms a plan with you, generates the connector against this SDK, and tests it. See the Indexing SDK overview for what it does and what to review in the output.

Or write it yourself

Requirements

Installation

pip install glean-indexing-sdk

# Optional cloud integrations
pip install "glean-indexing-sdk[aws]"   # Beta CloudWatch logs + metrics
pip install "glean-indexing-sdk[gcp]"   # Beta Cloud Logging/Monitoring + deployment Secret Manager

Quickstart

Every connector has two parts: a data client that fetches from your source, and a connector that transforms the result into Glean documents. The flow is fetch → transform → upload; you implement get_source_data() and transform(), and the SDK does the rest.

Set your credentials:

export GLEAN_SERVER_URL="https://your-company-be.glean.com"
export GLEAN_INDEXING_API_TOKEN="your-indexing-api-token"
export WIKI_API_TOKEN="your-wiki-token"

Then define and run a connector:

import os
from datetime import datetime
from typing import Any, List, Optional, Sequence, TypedDict

from glean.indexing.connectors import BaseDataClient, BaseDatasourceConnector
from glean.indexing.models import (
    ContentDefinition,
    CustomDatasourceConfig,
    DocumentDefinition,
    IndexingMode,
    UserReferenceDefinition,
)


class WikiPage(TypedDict):
    id: str
    title: str
    content: str
    author: str
    updated_at: str
    url: str


class WikiDataClient(BaseDataClient[WikiPage]):
    """Fetches pages from the source system. Replace the body with a real API call."""

    def __init__(self, base_url: str, api_token: str):
        self.base_url = base_url
        self.api_token = api_token

    def get_source_data(self, since: Optional[str] = None, **kwargs: Any) -> Sequence[WikiPage]:
        return [
            {
                "id": "page_123",
                "title": "Engineering Onboarding Guide",
                "content": "Welcome to the engineering team...",
                "author": "jane.smith@company.com",
                "updated_at": "2026-02-01T14:30:00Z",
                "url": f"{self.base_url}/pages/123",
            }
        ]


class CompanyWikiConnector(BaseDatasourceConnector[WikiPage]):
    """Transforms wiki pages into Glean documents."""

    configuration = CustomDatasourceConfig(
        name="companywiki",
        display_name="Company Wiki",
        url_regex=r"https://wiki\.company\.com/.*",
        is_user_referenced_by_email=True,
    )

    def transform(self, data: Sequence[WikiPage]) -> List[DocumentDefinition]:
        return [
            DocumentDefinition(
                id=page["id"],
                title=page["title"],
                datasource=self.name,
                view_url=page["url"],
                body=ContentDefinition(mime_type="text/plain", text_content=page["content"]),
                author=UserReferenceDefinition(email=page["author"]),
                # created_at / updated_at are epoch seconds, not ISO strings.
                updated_at=int(
                    datetime.fromisoformat(page["updated_at"].replace("Z", "+00:00")).timestamp()
                ),
            )
            for page in data
        ]


def create_connector() -> CompanyWikiConnector:
    """Build the production connector after runtime secrets are loaded."""
    return CompanyWikiConnector(
        name="companywiki",
        data_client=WikiDataClient(
            base_url="https://wiki.company.com",
            api_token=os.environ["WIKI_API_TOKEN"],
        ),
    )


if __name__ == "__main__":
    connector = create_connector()
    connector.configure_datasource()
    connector.index_data(mode=IndexingMode.FULL)

Test it without touching the network:

from glean.indexing.testing import StaticDataClient, run_connector

result = run_connector(CompanyWikiConnector("companywiki", StaticDataClient([...])))
result.assert_documents_posted(count=1, datasource="companywiki")

The module-level create_connector() function is the production construction seam for CLI and cloud execution. It keeps the connector constructor dependency-injectable for tests while letting runtime code build real clients after environment-backed secrets are loaded. Enable it in generated deployments with glean-idx deploy init --cloud gcp --connector-factory create_connector (or --cloud aws). Existing projects with a zero-argument connector class remain supported without a factory.

What's in the box

Capability What it gives you
Connector types Four base classes: in-memory, sync streaming, async streaming, and people/identity.
Pull integrations PullHttpClient with retries and backoff, link/offset/cursor pagination, and token-bucket rate limiting.
Push & indexing PushUploader for documents, users, groups, memberships, and employees, with parallel batch uploads.
Permissions Per-document ACLs and datasource identities so results respect who can see what.
Testing Three phases: fully mocked, real-source-with-record/replay, and live end-to-end.
Observability Structured logging and metrics, with optional CloudWatch and Google Cloud plugins.
Status & debugging StatusClient and glean-idx document status to answer "why isn't my document in search?"
Deployment glean-idx deploy generates Docker and Terraform for AWS or GCP.
Connector Builder An agent plugin that builds a connector from a description of your source.

Employee replacement compatibility

BasePeopleConnector and PushUploader.bulk_index_employees() preserve the current full-replacement employee workflow in the 1.x SDK. The upstream /bulkindexemployees endpoint is deprecated and scheduled for removal after October 15, 2026. Until an equivalent replacement contract is available, the SDK constrains glean-api-client to the verified 0.13.x line so dependency resolution cannot silently remove that generated-client surface. Removing or materially changing this workflow would require a new SDK major version.

The CLI

One command, glean-idx, covers the whole loop.

glean-idx doctor                    # are my credentials right?
glean-idx validate ./my-connector   # is the plan complete, before writing code?
glean-idx test --phase all          # mocked, then real source, then live
glean-idx run                       # crawl for real
glean-idx datasource status --datasource my-source
glean-idx document status --datasource my-source --document Article doc-1
glean-idx deploy init --cloud gcp   # Docker and Terraform for a CronJob

Commands split into two kinds, and glean-idx --help says which is which.

Most need only GLEAN_SERVER_URL and GLEAN_INDEXING_API_TOKEN, so they run anywhere, including with no install at all:

uvx --from glean-indexing-sdk glean-idx doctor

run, test, and datasource configure import your connector, so they run inside the connector project with the SDK installed alongside your code:

uv run glean-idx run

Every command takes --output json for a stable envelope, --yes to skip confirmations unattended, and returns a documented exit code — 3 for a missing precondition, 4 for a Glean error, 5 for a validation failure. In JSON mode the envelope goes to stdout whether it succeeded or not, so there is one stream to read. This contract begins after Click parses the invocation: Click usage and parse errors (such as unknown options, invalid values, or missing required options) happen before JSON command handling and remain text diagnostics on stderr with exit code 2, even when --output json is present.

glean-idx datasource status --datasource my-source --output json | jq .data.documents

glean-idx schema document prints the JSON Schema your transform() has to produce, and glean-idx completion zsh sets up tab completion.

Indexing modes

connector.index_data(mode=IndexingMode.FULL)         # re-index everything
connector.index_data(mode=IndexingMode.INCREMENTAL)  # only changes since the last crawl

A full crawl replaces the indexed state: documents absent from the run are deleted as stale. Incremental passes a since timestamp to your data client, but the SDK does not persist checkpoints — override _get_last_crawl_timestamp() on your connector to supply one. See Indexing modes.

Contributing

This project uses mise for toolchain management and uv for Python dependencies. See CONTRIBUTING.md.

mise run setup    # create venv and install dependencies
mise run test     # run all tests
mise run lint     # ruff, pyright, markdown-code
mise run lint:fix # auto-fix and format

License

MIT

Metadata

Release files for glean-indexing-sdk 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for glean-indexing-sdk 1.1.0
File Size Uploaded
glean_indexing_sdk-1.1.0.tar.gz 431.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for glean-indexing-sdk 1.1.0
File Interpreter ABI Platform
glean_indexing_sdk-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 603.4 kB

Release files / glean_indexing_sdk-1.1.0.tar.gz

Download URL glean_indexing_sdk-1.1.0.tar.gz
Size 431.7 kB
Tags Source
SHA-256 checksum
How to use checksums
317326e7aad7d74451c7d0f519da8d82acc50ec9df45a39aa3627d6839537e94
BLAKE2b-256 checksum
How to use checksums
791cd973d2a407d35d3deda85cc0c501eb5f5e413a2c4ec90f6df77780f16cc0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log

Release files / glean_indexing_sdk-1.1.0-py3-none-any.whl

Download URL glean_indexing_sdk-1.1.0-py3-none-any.whl
Size 171.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d4217e93c23e0ddf590c9205e647819d07b5d1694df71c7e42d90d807df83d2e
BLAKE2b-256 checksum
How to use checksums
3e200206707a6ad38ed275f9bdea83ced322db10971a376a8c22ca5b34574818
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page