Skip to main content

GLLM Privacy

Description

A library to protect Personal Identifiable Information (PII) in a Generative AI project.

Installation

Prerequisites

Mandatory:

  1. Python 3.11+ — Install here
  2. pip — Install here
  3. uv — Install here

Extras (required only for Artifact Registry installations):

  1. gcloud CLI (for authentication) — Install here, then log in using:
    gcloud auth login
    

Option 1: Install from Artifact Registry

This option requires authentication via the gcloud CLI.

uv pip install \
  --extra-index-url "https://oauth2accesstoken:$(gcloud auth print-access-token)@glsdk.gdplabs.id/gen-ai-internal/simple/" \
  gllm-privacy

Option 2: Install from PyPI

This option requires no authentication. However, it installs the binary wheel version of the package, which is fully usable but does not include source code.

uv pip install gllm-privacy-binary

Local Development Setup

Prerequisites

  1. Python 3.11+ — Install here

  2. pip — Install here

  3. uv — Install here

  4. gcloud CLI — Install here, then log in using:

    gcloud auth login
    
  5. Git — Install here

  6. Access to the GDP Labs SDK GitHub repository


1. Clone Repository

git clone git@github.com:GDP-ADMIN/gl-sdk.git
cd gl-sdk/libs/gllm-privacy

2. Setup Authentication

Set the following environment variables to authenticate with internal package indexes:

export UV_INDEX_GEN_AI_INTERNAL_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_INTERNAL_PASSWORD="$(gcloud auth print-access-token)"
export UV_INDEX_GEN_AI_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_PASSWORD="$(gcloud auth print-access-token)"

3. Quick Setup

Run:

make setup

4. Activate Virtual Environment

source .venv/bin/activate

Local Development Utilities

The following Makefile commands are available for quick operations:

Install uv

make install-uv

Install Pre-Commit

make install-pre-commit

Install Dependencies

make install

Update Dependencies

make update

Run Tests

make test

Run unit tests and model-free integration tests without downloading neural checkpoints:

uv run pytest

The download-backed Indonesian model tests are opt-in. To include them (requires network access and time to download the checkpoint and initialize its inference backend):

RUN_REAL_MODEL_TESTS=1 uv run pytest

Use uv run pytest -rs to display skipped-test requirements.


Usage

from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities
from gllm_privacy.pii_detector.anonymizer import Operation
from asyncio import run

text = """
    contoh nomor ktp 3525011212941001
    repeat nomor ktp 3525011212941001
    contoh email john.doe@example.com
    contoh nomor telepon +628121729819 dan 0812898029384.
    contoh npwp 01.123.456.7-891.234
"""
text_analyzer = TextAnalyzer()
entities = [Entities.EMAIL_ADDRESS, Entities.KTP, Entities.NPWP, Entities.PHONE_NUMBER]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = run(text_anonymizer.run(text=text, entities=entities))
print(anonymized_text)

deanonymized_text = run(text_anonymizer.run(text=text, entities=entities, operation=Operation.DEANONYMIZE))
print(deanonymized_text)

If you need to detect person, organization, or location entities in text written in Bahasa Indonesia, you can use either TransformersRecognizer or ProsaRemoteRecognizer. To use the TransformersRecognizer, you can use it like this:

from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer
from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities

# Load the model, if you run it for the first time, it will download the model from the Hugging Face model hub
transformers_recognizer = TransformersRecognizer(
  model_path=CAHYA_BERT_CONFIGURATION.get("DEFAULT_MODEL_PATH"),
  supported_entities=CAHYA_BERT_CONFIGURATION.get("PRESIDIO_SUPPORTED_ENTITIES"),
)
transformers_recognizer.load_transformer(**CAHYA_BERT_CONFIGURATION)
analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])

text = "John Doe adalah seorang karyawan PT ABCD yang berlokasi di Jakarta."
text_analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])
entities = [Entities.PERSON, Entities.LOCATION]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

deanonymized_text = text_anonymizer.deanonymize(text=text)
print(deanonymized_text)

Detecting Brand Names

BRAND_NAME is a supported entity. Like PERSON, it is detected by a model-backed recognizer instead of a built-in regex pattern. Point the model's brand label at BRAND_NAME in MODEL_TO_PRESIDIO_MAPPING, add BRAND_NAME to supported_entities, then request it during analysis or anonymization:

from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities
from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer

configuration = {
    **CAHYA_BERT_CONFIGURATION,
    "PRESIDIO_SUPPORTED_ENTITIES": [*CAHYA_BERT_CONFIGURATION["PRESIDIO_SUPPORTED_ENTITIES"], "BRAND_NAME"],
    "MODEL_TO_PRESIDIO_MAPPING": {**CAHYA_BERT_CONFIGURATION["MODEL_TO_PRESIDIO_MAPPING"], "BRD": "BRAND_NAME"},
}

brand_recognizer = TransformersRecognizer(
    model_path="<YOUR_BRAND_NER_MODEL>",
    supported_entities=configuration["PRESIDIO_SUPPORTED_ENTITIES"],
)
brand_recognizer.load_transformer(**configuration)

text = "Saya memakai sepatu Nike."
text_analyzer = TextAnalyzer(additional_recognizers=[brand_recognizer])
text_anonymizer = TextAnonymizer(text_analyzer, add_default_faker_operators=True)
entities = [Entities.BRAND_NAME]

anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

print(text_anonymizer.deanonymize(text=anonymized_text))

With add_default_faker_operators=True, BRAND_NAME values are pseudo-anonymized with Faker company names (fake.company()). Without it, they are replaced by the default <BRAND_NAME_n> placeholder.

The brand recognizer also accepts LINE_BREAK_TOKEN and CHUNK_SIZE:

  • LINE_BREAK_TOKEN (str | None, default None): token the model was trained on to represent a line break (for example "→"). When set, every "\n" in the input is replaced by " <token> " before inference and detected offsets are mapped back to the original text. Newlines remain intact when the model predicts separate entities per line; this is not enforced by the mapper. Leave it unset for models trained without a break token — their input is passed through unchanged. A literal token already present in the text is preserved and may split a model span. Conversely, a model span crossing an inserted token includes the original newline, so masking that span removes the line break. Such spans are not split into per-line results.
  • CHUNK_SIZE (int): maximum tokens per chunk, including special tokens. Set it to the model's real limit (for example 512 for XLM-R based checkpoints) so long multi-line documents are chunked instead of truncated.

Enhanced TransformersRecognizer with Optimum

The TransformersRecognizer now supports Hugging Face Optimum for improved performance:

  • ONNX Runtime with CUDA: GPU-accelerated inference using ONNX Runtime with CUDA provider
  • ONNX Runtime with CPU: Optimized CPU inference for better performance on laptops/servers
  • Apple Silicon MPS: GPU acceleration on Apple Silicon Macs
  • Auto-detection: Automatically selects the best available backend
  • Fallback compatibility: Works on any hardware with standard transformers

Available Backends:

  • onnx: ONNX Runtime with CPU provider (optimized for NER tasks)
  • cuda: ONNX Runtime with CUDA provider (GPU acceleration)
  • mps: Apple Silicon MPS for GPU acceleration on Mac
  • transformers: Standard transformers as fallback

Configuration Options:

You can configure the backend behavior in your configuration:

config = {
    "USE_OPTIMUM": True,                    # Enable/disable Optimum
    "OPTIMUM_BACKEND": "auto",              # "auto", "onnx", "cuda", "mps", "transformers"
    "OPTIMUM_DEVICE": "auto",               # "auto", "cuda", "cpu", "mps"
    "OPTIMUM_QUANTIZATION": False,          # Enable quantization
    "OPTIMUM_MAX_BATCH_SIZE": 8,           # Max batch size
}

Usage Example:

from gllm_privacy.pii_detector import TextAnalyzer
from gllm_privacy.pii_detector.recognizer.config import CAHYA_BERT_CONFIGURATION
from gllm_privacy.pii_detector.recognizer.transformers_recognizer import TransformersRecognizer

transformers_recognizer = TransformersRecognizer(
    model_path=CAHYA_BERT_CONFIGURATION.get("DEFAULT_MODEL_PATH"),
    supported_entities=CAHYA_BERT_CONFIGURATION.get("PRESIDIO_SUPPORTED_ENTITIES"),
    use_optimum=True
)

transformers_recognizer.load_transformer(**CAHYA_BERT_CONFIGURATION)

pipeline_info = transformers_recognizer.get_pipeline_info()
print(f"Backend: {pipeline_info['backend']}")
print(f"Device: {pipeline_info['device']}")
print(f"Optimizations: {pipeline_info['optimizations']}")

# Use as before
analyzer = TextAnalyzer(additional_recognizers=[transformers_recognizer])

To use the ProsaRemoteRecognizer, you can use it like the following example. Please replace <PROSA_API_URL> and <PROSA_API_KEY> with the valid values.

from gllm_privacy.pii_detector.recognizer.prosa_remote_recognizer import ProsaRemoteRecognizer
from gllm_privacy.pii_detector import TextAnalyzer, TextAnonymizer
from gllm_privacy.pii_detector.constants import Entities

text = "John Doe adalah seorang karyawan PT ABCD yang berlokasi di Jakarta."
prosa_recognizer = ProsaRemoteRecognizer('<PROSA_API_URL>', '<PROSA_API_KEY>')
text_analyzer = TextAnalyzer(additional_recognizers=[prosa_recognizer])
entities = [Entities.PERSON, Entities.LOCATION]

text_anonymizer = TextAnonymizer(text_analyzer)
anonymized_text = text_anonymizer.anonymize(text=text, entities=entities)
print(anonymized_text)

deanonymized_text = text_anonymizer.deanonymize(text=text)
print(deanonymized_text)

Metadata

Release files for gllm-privacy-binary 0.4.35

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for gllm-privacy-binary 0.4.35
File
gllm_privacy_binary-0.4.35-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
gllm_privacy_binary-0.4.35-cp313-cp313-manylinux_2_31_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.35-cp313-cp313-macosx_13_0_arm64.whl CPython 3.13 CPython 3.13 macOS 13.0+ ARM64 Details
gllm_privacy_binary-0.4.35-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
gllm_privacy_binary-0.4.35-cp312-cp312-manylinux_2_31_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.35-cp312-cp312-macosx_13_0_arm64.whl CPython 3.12 CPython 3.12 macOS 13.0+ ARM64 Details
gllm_privacy_binary-0.4.35-cp311-cp311-win_amd64.whl CPython 3.11 CPython 3.11 Windows x86-64 Details
gllm_privacy_binary-0.4.35-cp311-cp311-manylinux_2_31_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.31+ x86-64 Details
gllm_privacy_binary-0.4.35-cp311-cp311-macosx_13_0_arm64.whl CPython 3.11 CPython 3.11 macOS 13.0+ ARM64 Details

Total release size: 6.1 MB

Release files / gllm_privacy_binary-0.4.35-cp313-cp313-win_amd64.whl

Download URL gllm_privacy_binary-0.4.35-cp313-cp313-win_amd64.whl
Size 561.0 kB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
5a63e1d57eeabb26eb549c058bdf43e025e02da0e8f9705554c2f39f1f113786
BLAKE2b-256 checksum
How to use checksums
7394d05d5af403257f9bcb6d407242fa9eabd302baeeb0bfe1a8f2dd42d862e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.35-cp313-cp313-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.35-cp313-cp313-manylinux_2_31_x86_64.whl
Size 888.0 kB
Tags CPython 3.13 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
6ca0be8278def7b43aac552c3e4d525177803e3533e6840b22bee31980e7c2ee
BLAKE2b-256 checksum
How to use checksums
6d15255e666fb4d3e4cd3a163bc52c99ec5da84b87b87051456ad176f3c07906
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.35-cp313-cp313-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.35-cp313-cp313-macosx_13_0_arm64.whl
Size 627.4 kB
Tags CPython 3.13 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
de8ae9a0795070cbebc299bc4258befa70a7d669b27967d67f280bdf97086e9b
BLAKE2b-256 checksum
How to use checksums
ddf9a5db033edaba263a901c6f9a417f8f14ae12597c9f03db4111ea6adce46e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.35-cp312-cp312-win_amd64.whl

Download URL gllm_privacy_binary-0.4.35-cp312-cp312-win_amd64.whl
Size 562.4 kB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
0340305ad3407ca6b04bd4533d4c56f7a8b3a412fe8e1ebf484665dad1c6a596
BLAKE2b-256 checksum
How to use checksums
b94058200373589af5f3cfacf690195935fe2fdfac3f9a9228535cc118edd57a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.35-cp312-cp312-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.35-cp312-cp312-manylinux_2_31_x86_64.whl
Size 881.8 kB
Tags CPython 3.12 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
1e64a323da6265cb10258a7c7b83b43eed0967f976937c89d0c457caaf88b1dc
BLAKE2b-256 checksum
How to use checksums
78d67ec2430dc22c2fcf7167581553652b578b99987ae7de573d97e88178274e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.35-cp312-cp312-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.35-cp312-cp312-macosx_13_0_arm64.whl
Size 607.5 kB
Tags CPython 3.12 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
1512e2d42bb38af0726c127de992939d6e8fb2bfcd4c812a1ca6dccf1d30c858
BLAKE2b-256 checksum
How to use checksums
fadddb132abb1ff829f291dd1073a827c07c6b672c4e9cf2fe7587626c81da84
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.35-cp311-cp311-win_amd64.whl

Download URL gllm_privacy_binary-0.4.35-cp311-cp311-win_amd64.whl
Size 582.4 kB
Tags CPython 3.11 Windows x86-64
SHA-256 checksum
How to use checksums
eb2e9efa5044064b593982b138fd6652be1774c3392b4915451c57a4ca9325a2
BLAKE2b-256 checksum
How to use checksums
3d1c7fd7d5aecbfde5c5ac209010fda864a94e33c34935638efe92f8c87dbff9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / gllm_privacy_binary-0.4.35-cp311-cp311-manylinux_2_31_x86_64.whl

Download URL gllm_privacy_binary-0.4.35-cp311-cp311-manylinux_2_31_x86_64.whl
Size 805.1 kB
Tags CPython 3.11 Linux glibc 2.31+ x86-64
SHA-256 checksum
How to use checksums
60cc786cfa964caea4ffb4247e520df072037eb2e714beafd5c415fad4d80107
BLAKE2b-256 checksum
How to use checksums
c38c919fb5a537b2466f8edb609ce33703f4f4dcbf4e98bb8b0ee6aeac6c2618
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.24

Release files / gllm_privacy_binary-0.4.35-cp311-cp311-macosx_13_0_arm64.whl

Download URL gllm_privacy_binary-0.4.35-cp311-cp311-macosx_13_0_arm64.whl
Size 597.1 kB
Tags CPython 3.11 macOS 13.0+ ARM64
SHA-256 checksum
How to use checksums
af13be343169572241d2874b0d56c6c698b1e5283a1213b4cec7a732a7acbe74
BLAKE2b-256 checksum
How to use checksums
cce04b47289093cd6733cfad1cf1427a252997556e2a1b6d5269a0d2f50ccfce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page