Skip to main content

GLLM Privacy

Description

A library to protect Personal Identifiable Information (PII) in a Generative AI project.

Installation

Prerequisites

Mandatory:

  1. Python 3.11+ — Install here
  2. pip — Install here
  3. uv — Install here

Extras (required only for Artifact Registry installations):

  1. gcloud CLI (for authentication) — Install here, then log in using:
    gcloud auth login
    

Option 1: Install from Artifact Registry

This option requires authentication via the gcloud CLI.

uv pip install \
  --extra-index-url "https://oauth2accesstoken:$(gcloud auth print-access-token)@glsdk.gdplabs.id/gen-ai-internal/simple/" \
  gllm-privacy

Option 2: Install from PyPI

This option requires no authentication. However, it installs the binary wheel version of the package, which is fully usable but does not include source code.

uv pip install gllm-privacy-binary

Local Development Setup

Prerequisites

  1. Python 3.11+ — Install here

  2. pip — Install here

  3. uv — Install here

  4. gcloud CLI — Install here, then log in using:

    gcloud auth login
    
  5. Git — Install here

  6. Access to the GDP Labs SDK GitHub repository


1. Clone Repository

git clone git@github.com:GDP-ADMIN/gl-sdk.git
cd gl-sdk/libs/gllm-privacy

2. Setup Authentication

Set the following environment variables to authenticate with internal package indexes:

export UV_INDEX_GEN_AI_INTERNAL_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_INTERNAL_PASSWORD="$(gcloud auth print-access-token)"
export UV_INDEX_GEN_AI_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_PASSWORD="$(gcloud auth print-access-token)"

3. Quick Setup

Run:

make setup

4. Activate Virtual Environment

source .venv/bin/activate

Local Development Utilities

The following Makefile commands are available for quick operations:

Install uv

make install-uv

Install Pre-Commit

make install-pre-commit

Install Dependencies

make install

Update Dependencies

make update

Run Tests

make test

Run unit tests and model-free integration tests without downloading neural checkpoints:

uv run pytest

The download-backed Indonesian model tests are opt-in. To include them (requires network access and time to download the checkpoint and initialize its inference backend):

RUN_REAL_MODEL_TESTS=1 uv run pytest

Use uv run pytest -rs to display skipped-test requirements.


Usage

from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities
from gllm_privacy.pii_detector.anonymizer import Operation
from asyncio import run

text = """
    contoh nomor ktp 3525011212941001
    repeat nomor ktp 3525011212941001
    contoh email john.doe@example.com
    contoh nomor telepon +628121729819 dan 0812898029384.
    contoh npwp 01.123.456.7-891.234
"""
text_analyzer = TextAnalyzer()
entities = [Entities.EMAIL_ADDRESS, Entities.KTP, Entities.NPWP, Entities.PHONE_NUMBER]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = run(text_anonymizer.run(text=text, entities=entities))
print(anonymized_text)

deanonymized_text = run(text_anonymizer.run(text=text, entities=entities, operation=Operation.DEANONYMIZE))
print(deanonymized_text)

If you need to detect person, organization, or location entities in text written in Bahasa Indonesia, you can use either TransformersRecognizer or ProsaRemoteRecognizer. To use the TransformersRecognizer, you can use it like this:

from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer
from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities

# Load the model, if you run it for the first time, it will download the model from the Hugging Face model hub
transformers_recognizer = TransformersRecognizer(
  model_path=CAHYA_BERT_CONFIGURATION.get("DEFAULT_MODEL_PATH"),
  supported_entities=CAHYA_BERT_CONFIGURATION.get("PRESIDIO_SUPPORTED_ENTITIES"),
)
transformers_recognizer.load_transformer(**CAHYA_BERT_CONFIGURATION)
analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])

text = "John Doe adalah seorang karyawan PT ABCD yang berlokasi di Jakarta."
text_analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])
entities = [Entities.PERSON, Entities.LOCATION]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

deanonymized_text = text_anonymizer.deanonymize(text=text)
print(deanonymized_text)

Detecting Brand Names

BRAND_NAME is a supported entity. Like PERSON, it is detected by a model-backed recognizer instead of a built-in regex pattern. Point the model's brand label at BRAND_NAME in MODEL_TO_PRESIDIO_MAPPING, add BRAND_NAME to supported_entities, then request it during analysis or anonymization:

from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities
from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer

configuration = {
    **CAHYA_BERT_CONFIGURATION,
    "PRESIDIO_SUPPORTED_ENTITIES": [*CAHYA_BERT_CONFIGURATION["PRESIDIO_SUPPORTED_ENTITIES"], "BRAND_NAME"],
    "MODEL_TO_PRESIDIO_MAPPING": {**CAHYA_BERT_CONFIGURATION["MODEL_TO_PRESIDIO_MAPPING"], "BRD": "BRAND_NAME"},
}

brand_recognizer = TransformersRecognizer(
    model_path="<YOUR_BRAND_NER_MODEL>",
    supported_entities=configuration["PRESIDIO_SUPPORTED_ENTITIES"],
)
brand_recognizer.load_transformer(**configuration)

text = "Saya memakai sepatu Nike."
text_analyzer = TextAnalyzer(additional_recognizers=[brand_recognizer])
text_anonymizer = TextAnonymizer(text_analyzer, add_default_faker_operators=True)
entities = [Entities.BRAND_NAME]

anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

print(text_anonymizer.deanonymize(text=anonymized_text))

With add_default_faker_operators=True, BRAND_NAME values are pseudo-anonymized with Faker company names (fake.company()). Without it, they are replaced by the default <BRAND_NAME_n> placeholder.

The brand recognizer also accepts LINE_BREAK_TOKEN and CHUNK_SIZE:

  • LINE_BREAK_TOKEN (str | None, default None): token the model was trained on to represent a line break (for example "→"). When set, every "\n" in the input is replaced by " <token> " before inference and detected offsets are mapped back to the original text. Newlines remain intact when the model predicts separate entities per line; this is not enforced by the mapper. Leave it unset for models trained without a break token — their input is passed through unchanged. A literal token already present in the text is preserved and may split a model span. Conversely, a model span crossing an inserted token includes the original newline, so masking that span removes the line break. Such spans are not split into per-line results.
  • CHUNK_SIZE (int): maximum tokens per chunk, including special tokens. Set it to the model's real limit (for example 512 for XLM-R based checkpoints) so long multi-line documents are chunked instead of truncated.

Enhanced TransformersRecognizer with Optimum

The TransformersRecognizer now supports Hugging Face Optimum for improved performance:

  • ONNX Runtime with CUDA: GPU-accelerated inference using ONNX Runtime with CUDA provider
  • ONNX Runtime with CPU: Optimized CPU inference for better performance on laptops/servers
  • Apple Silicon MPS: GPU acceleration on Apple Silicon Macs
  • Auto-detection: Automatically selects the best available backend
  • Fallback compatibility: Works on any hardware with standard transformers

Available Backends:

  • onnx: ONNX Runtime with CPU provider (optimized for NER tasks)
  • cuda: ONNX Runtime with CUDA provider (GPU acceleration)
  • mps: Apple Silicon MPS for GPU acceleration on Mac
  • transformers: Standard transformers as fallback

Configuration Options:

You can configure the backend behavior in your configuration:

config = {
    "USE_OPTIMUM": True,                    # Enable/disable Optimum
    "OPTIMUM_BACKEND": "auto",              # "auto", "onnx", "cuda", "mps", "transformers"
    "OPTIMUM_DEVICE": "auto",               # "auto", "cuda", "cpu", "mps"
    "OPTIMUM_QUANTIZATION": False,          # Enable quantization
    "OPTIMUM_MAX_BATCH_SIZE": 8,           # Max batch size
}

Usage Example:

from gllm_privacy.pii_detector import TextAnalyzer
from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer

transformers_recognizer = TransformersRecognizer(
    model_path=CAHYA_BERT_CONFIGURATION.get("DEFAULT_MODEL_PATH"),
    supported_entities=CAHYA_BERT_CONFIGURATION.get("PRESIDIO_SUPPORTED_ENTITIES"),
    use_optimum=True
)

transformers_recognizer.load_transformer(**CAHYA_BERT_CONFIGURATION)

pipeline_info = transformers_recognizer.get_pipeline_info()
print(f"Backend: {pipeline_info['backend']}")
print(f"Device: {pipeline_info['device']}")
print(f"Optimizations: {pipeline_info['optimizations']}")

# Use as before
analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])

To use the ProsaRemoteRecognizer, you can use it like the following example. Please replace <PROSA_API_URL> and <PROSA_API_KEY> with the valid values.

from gllm_privacy.pii_detector.recognizer.prosa_remote_recognizer import ProsaRemoteRecognizer
from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities

text = "John Doe adalah seorang karyawan PT ABCD yang berlokasi di Jakarta."
prosa_recognizer = ProsaRemoteRecognizer('<PROSA_API_URL>', '<PROSA_API_KEY>')
text_analyzer = TextAnalyzer(additional_recognizers=[prosa_recognizer])
entities = [Entities.PERSON, Entities.LOCATION]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

deanonymized_text = text_anonymizer.deanonymize(text=text)
print(deanonymized_text)

Metadata

Release files for gllm-privacy-binary 0.4.34

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for gllm-privacy-binary 0.4.34
File
gllm_privacy_binary-0.4.34-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
gllm_privacy_binary-0.4.34-cp313-cp313-manylinux_2_31_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.34-cp313-cp313-macosx_13_0_arm64.whl CPython 3.13 CPython 3.13 macOS 13.0+ ARM64 Details
gllm_privacy_binary-0.4.34-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
gllm_privacy_binary-0.4.34-cp312-cp312-manylinux_2_31_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.34-cp312-cp312-macosx_13_0_arm64.whl CPython 3.12 CPython 3.12 macOS 13.0+ ARM64 Details
gllm_privacy_binary-0.4.34-cp311-cp311-win_amd64.whl CPython 3.11 CPython 3.11 Windows x86-64 Details
gllm_privacy_binary-0.4.34-cp311-cp311-manylinux_2_31_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.34-cp311-cp311-macosx_13_0_arm64.whl CPython 3.11 CPython 3.11 macOS 13.0+ ARM64 Details

Total release size: 6.1 MB

Release files / gllm_privacy_binary-0.4.34-cp313-cp313-win_amd64.whl

Download URL gllm_privacy_binary-0.4.34-cp313-cp313-win_amd64.whl
Size 560.1 kB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
f1d7395627eb19f4743ac46d2414b3d754b7cde3d5c1a447c0db9ceb09bf9290
BLAKE2b-256 checksum
How to use checksums
01a791e1ebfa3529971f3efabb434438c4e0daf492ec6f81ddf0aa5d4fd97227
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.34-cp313-cp313-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.34-cp313-cp313-manylinux_2_31_x86_64.whl
Size 886.2 kB
Tags CPython 3.13 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
70fbf9bafd63f7c2005b9ca0717d7d879a7d8c9b9cbb0f96a45f03b4a79f67e0
BLAKE2b-256 checksum
How to use checksums
89003e909a7b0772aead40646869b3c7933ab23e64ad88418c42b12b101f84e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.34-cp313-cp313-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.34-cp313-cp313-macosx_13_0_arm64.whl
Size 625.4 kB
Tags CPython 3.13 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
d31e10313ca2778572f4b447a08ded01391633efdead70f1d3dd91be2db6fbe6
BLAKE2b-256 checksum
How to use checksums
a634b37d69147f3a489f8f6056b1e41874eda7cabc351049cedc3eee7658d557
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.34-cp312-cp312-win_amd64.whl

Download URL gllm_privacy_binary-0.4.34-cp312-cp312-win_amd64.whl
Size 561.1 kB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
9ba303e4003a205a3ab23268b97865586b3a54545dbc1a4854931edb7f1ca108
BLAKE2b-256 checksum
How to use checksums
d3fda40657d1b6c05e28aadde70896dccbf53ff3c3348298979c67063d0ba40c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.34-cp312-cp312-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.34-cp312-cp312-manylinux_2_31_x86_64.whl
Size 880.0 kB
Tags CPython 3.12 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
f28c979e5b5539acdfb477eefe2131b5a16d2eff09acd64748d462d2b7a16664
BLAKE2b-256 checksum
How to use checksums
4ef9cb0e35e8885b96ffcf0350dde609a37d6fc8cb3129dc62c2c5a0de27ecff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.34-cp312-cp312-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.34-cp312-cp312-macosx_13_0_arm64.whl
Size 606.5 kB
Tags CPython 3.12 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
ee177ed5a28d4846acf861d5dd077ca253865594658a834f631fd19a5f52aec1
BLAKE2b-256 checksum
How to use checksums
3eef73a8bcf17917588e4a767ed8e485ad29d6d44a5ae2c67ebbdf839c817111
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.34-cp311-cp311-win_amd64.whl

Download URL gllm_privacy_binary-0.4.34-cp311-cp311-win_amd64.whl
Size 581.2 kB
Tags CPython 3.11 Windows x86-64
SHA-256 checksum
How to use checksums
63f5430f2f7a6c26932eb654e7325d6835e77e5c3a45a8c9f55468da26314752
BLAKE2b-256 checksum
How to use checksums
6c7c6b64fb26c795b19ca2b06b6918ca78b364f395e58adb0b6e6688b1dab295
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.34-cp311-cp311-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.34-cp311-cp311-manylinux_2_31_x86_64.whl
Size 803.2 kB
Tags CPython 3.11 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
25a8dde5408ee0ca4bd3b727626e1d04f066e33d8e4647ae7fb7a40c5f7f5ad8
BLAKE2b-256 checksum
How to use checksums
50fb712ac2a4425f7ecfdc718c3987959a4f9eeed11537a553395d9f04b563a9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.34-cp311-cp311-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.34-cp311-cp311-macosx_13_0_arm64.whl
Size 594.0 kB
Tags CPython 3.11 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
8b2735b15b9af32002eaffe929948a76ef986f747160e6d78e360930d7a8c4dd
BLAKE2b-256 checksum
How to use checksums
4441027e37671f6014c5ae2f28ca362e2a9eb0e61ab49f1f5569705d9af22c9d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page